A 55B-parameter mixture-of-experts model from NVIDIA with a hybrid Transformer-Mamba architecture and 512K token context window for reasoning and task orchestration.
- Strengths
- The combination of Transformer and Mamba layers in a sparse MoE design enables efficient reasoning while maintaining a very large context window for handling lengthy inputs.
- Best for
- Long-context reasoning tasks, multi-step problem-solving, and applications that require both extended context awareness and efficient inference across reasoning workloads.
- Limitations
- As a June 2026 release, this is a newer model in NVIDIA's lineup and may have less production validation than earlier Nemotron releases; the MoE architecture requires compatible hardware and inference infrastructure to realize efficiency gains.
Input / 1M
$0.3
Output / 1M
$1.80
Cached input / 1M
$0.1
Context window
512K
Price history
Snapshots
| Effective | Input | Output | Cached in | Note | Source |
|---|---|---|---|---|---|
| 6 Aug 2026 | $0.3 | $1.80 | $0.1 | Imported from OpenRouter | openrouter.ai |