prices synced 2026-09-11
I

Granite 4.0 Micro

IBM · Released Oct 2025

Compare

A lightweight language model from IBM's Granite family with a 131K-token context window, designed for efficient inference on resource-constrained environments.

Strengths
Delivers faster inference and lower memory requirements than larger Granite models while maintaining broad language understanding capabilities.
Best for
Applications prioritizing latency and resource efficiency, such as edge deployment, real-time inference, and scenarios with limited computational budgets.
Limitations
Smaller parameter count limits reasoning depth and performance on complex multi-step tasks compared to Granite 4.2 8B and other larger models in the lineup.

Input / 1M

$0.017

Output / 1M

$0.112

Cached input / 1M

Context window

131K

Price history

Price per 1M tokens over timeOutput $0.112; Input $0.017 as of Jun 2026.$0$0.05$0.1Output on 11 Jun 2026: $0.112Out $0.112Input on 11 Jun 2026: $0.017In $0.017Jun 2026

Snapshots

Effective Input Output Cached in Note Source
11 Jun 2026 $0.017 $0.112 Imported from OpenRouter openrouter.ai

More from IBM

Report a problem