A 30-billion-parameter mixture-of-experts model from NVIDIA with 3B active parameters per token and a 262K token context window.
- Strengths
- Fast inference throughput due to its small active parameter count relative to total model size, making it well-suited for latency-sensitive applications.
- Best for
- High-throughput agentic workloads and specialized tasks where inference speed is prioritized over reasoning depth.
- Limitations
- The 3B active parameter constraint means it trades off reasoning capability and task complexity handling compared to larger NVIDIA models like Nemotron 3 Ultra or Nemotron 3 Super.
Input / 1M
$0.1
Output / 1M
$0.25
Cached input / 1M
$0.05
Context window
262K
Price history
Snapshots
| Effective | Input | Output | Cached in | Note | Source |
|---|---|---|---|---|---|
| 11 Aug 2026 | $0.1 | $0.25 | $0.05 | Imported from OpenRouter | openrouter.ai |