A 55-billion-parameter mixture-of-experts model from NVIDIA with a hybrid Transformer-Mamba architecture and 512K token context window.
- Strengths
- Handles very long documents and extended reasoning tasks with its large context window and hybrid architecture that combines Transformer and Mamba components for efficiency.
- Best for
- Batch processing of long-form content, document analysis, and multi-turn reasoning where the expanded context window is critical.
- Limitations
- Superseded by the faster Nemotron 3.5 Lightning released two months later for latency-sensitive workloads, and not optimized for real-time inference like newer models in the lineup.
Input / 1M
$0.6
Output / 1M
$3.60
Cached input / 1M
$0.2
Context window
512K
Price history
Input (solid)Output (dashed)
Price change
- 30d
- in — out —
- 90d
- in increased 100.0% out increased 100.0%
- 1y
- in increased 100.0% out increased 100.0%
- Since launch
- in increased 100.0% out increased 100.0%
Snapshots
| Effective | Input | Output | Cached in | Note | Source |
|---|---|---|---|---|---|
| 19 Aug 2026 | $0.6 | $3.60 | $0.2 | Imported from OpenRouter | openrouter.ai |
| 6 Aug 2026 | $0.3 | $1.80 | $0.1 | Imported from OpenRouter | openrouter.ai |