A lightweight language model from IBM's Granite family with a 131K-token context window, designed for efficient inference on resource-constrained environments.
- Strengths
- Delivers faster inference and lower memory requirements than larger Granite models while maintaining broad language understanding capabilities.
- Best for
- Applications prioritizing latency and resource efficiency, such as edge deployment, real-time inference, and scenarios with limited computational budgets.
- Limitations
- Smaller parameter count limits reasoning depth and performance on complex multi-step tasks compared to Granite 4.2 8B and other larger models in the lineup.
Input / 1M
$0.017
Output / 1M
$0.112
Cached input / 1M
—
Context window
131K
Price history
Snapshots
| Effective | Input | Output | Cached in | Note | Source |
|---|---|---|---|---|---|
| 11 Jun 2026 | $0.017 | $0.112 | — | Imported from OpenRouter | openrouter.ai |