prices synced 2026-08-07
N

Nemotron 3 Ultra (batch)

NVIDIA · Released Jun 2026

Compare

A 55B-parameter mixture-of-experts model from NVIDIA with a hybrid Transformer-Mamba architecture and 512K token context window for reasoning and task orchestration.

Strengths
The combination of Transformer and Mamba layers in a sparse MoE design enables efficient reasoning while maintaining a very large context window for handling lengthy inputs.
Best for
Long-context reasoning tasks, multi-step problem-solving, and applications that require both extended context awareness and efficient inference across reasoning workloads.
Limitations
As a June 2026 release, this is a newer model in NVIDIA's lineup and may have less production validation than earlier Nemotron releases; the MoE architecture requires compatible hardware and inference infrastructure to realize efficiency gains.

Input / 1M

$0.3

Output / 1M

$1.80

Cached input / 1M

$0.1

Context window

512K

Price history

Price per 1M tokens over timeOutput $1.80; Input $0.3 as of Aug 2026.$0$0.5$1.00$1.50$2.00Output on 6 Aug 2026: $1.80Out $1.80Input on 6 Aug 2026: $0.3In $0.3Aug 2026

Snapshots

Effective Input Output Cached in Note Source
6 Aug 2026 $0.3 $1.80 $0.1 Imported from OpenRouter openrouter.ai

More from NVIDIA

Report a problem