prices synced 2026-09-21
N

Nemotron 3 Ultra (batch)

NVIDIA · Released Jun 2026

Compare

A 55-billion-parameter mixture-of-experts model from NVIDIA with a hybrid Transformer-Mamba architecture and 512K token context window.

Strengths
Handles very long documents and extended reasoning tasks with its large context window and hybrid architecture that combines Transformer and Mamba components for efficiency.
Best for
Batch processing of long-form content, document analysis, and multi-turn reasoning where the expanded context window is critical.
Limitations
Superseded by the faster Nemotron 3.5 Lightning released two months later for latency-sensitive workloads, and not optimized for real-time inference like newer models in the lineup.

Input / 1M

$0.6

Output / 1M

$3.60

Cached input / 1M

$0.2

Context window

512K

Price history

Input (solid)Output (dashed)
Price per 1M tokens over timeOutput went from $1.80 to $3.60; Input went from $0.3 to $0.6 between Aug 2026 and Aug 2026.$0$1.00$2.00$3.00$4.00Output on 6 Aug 2026: $1.80Output on 19 Aug 2026: $3.60Out $3.60Input on 6 Aug 2026: $0.3Input on 19 Aug 2026: $0.6In $0.6Aug 2026Aug 2026

Price change

30d
in out
90d
in increased 100.0% out increased 100.0%
1y
in increased 100.0% out increased 100.0%
Since launch
in increased 100.0% out increased 100.0%

Snapshots

Effective Input Output Cached in Note Source
19 Aug 2026 $0.6 $3.60 $0.2 Imported from OpenRouter openrouter.ai
6 Aug 2026 $0.3 $1.80 $0.1 Imported from OpenRouter openrouter.ai

More from NVIDIA

Report a problem