A 30-billion-parameter mixture-of-experts model from NVIDIA with 3 billion active parameters per token and 262K token context window.
- Strengths
- Runs efficiently on modest hardware due to sparse activation of mixture-of-experts architecture, making it accessible for resource-constrained deployments.
- Best for
- Latency-sensitive applications and edge deployments where computational efficiency matters more than peak reasoning capability.
- Limitations
- Smaller active parameter count than Nemotron 3 Super and earlier, and lacks the reasoning optimizations or architectural innovations (Transformer-Mamba hybrid, extended context) present in newer NVIDIA releases.
Input / 1M
$0.05
Output / 1M
$0.2
Cached input / 1M
$0.03
Context window
262K
Price history
Input (solid)Output (dashed)
Price change
- 30d
- in decreased 0.0% out decreased 0.0%
- 90d
- in decreased 0.0% out decreased 0.0%
- 1y
- in decreased 0.0% out decreased 0.0%
- Since launch
- in decreased 0.0% out decreased 0.0%
Snapshots
| Effective | Input | Output | Cached in | Note | Source |
|---|---|---|---|---|---|
| 3 Sep 2026 | $0.05 | $0.2 | $0.03 | Imported from OpenRouter | openrouter.ai |
| 27 Aug 2026 | $0.05 | $0.2 | $0.025 | Imported from OpenRouter | openrouter.ai |
| 17 Aug 2026 | $0.05 | $0.2 | $0.03 | Imported from OpenRouter | openrouter.ai |
| 10 Aug 2026 | $0.05 | $0.2 | $0.025 | Imported from OpenRouter | openrouter.ai |
| 31 Jul 2026 | $0.05 | $0.2 | $0.03 | Imported from OpenRouter | openrouter.ai |
| 29 Jul 2026 | $0.05 | $0.2 | $0.025 | Imported from OpenRouter | openrouter.ai |
| 11 Jun 2026 | $0.05 | $0.2 | — | Imported from OpenRouter | openrouter.ai |