A 30-billion-parameter mixture-of-experts model from NVIDIA designed for efficient inference and specialized task performance.
- Strengths
- Delivers strong accuracy relative to its size through MoE architecture, with low latency and memory requirements for inference-heavy workloads.
- Best for
- Building specialized agentic systems and multi-step reasoning tasks where you need a balance between capability and compute efficiency.
- Limitations
- As a smaller model, it may struggle with complex reasoning tasks or deep knowledge queries compared to larger generalist models.
Input / 1M
$0.05
Output / 1M
$0.2
Cached input / 1M
—
Context window
262K
Price history
Snapshots
| Effective | Input | Output | Cached in | Note | Source |
|---|---|---|---|---|---|
| 11 Jun 2026 | $0.05 | $0.2 | — | Imported from OpenRouter | openrouter.ai |