prices synced 2026-09-11
N

Llama 3.3 Nemotron Super 49B V1.5

NVIDIA · Released Oct 2025

Compare

A 49-billion-parameter model from NVIDIA with a 131K token context window, positioned between the earlier Nemotron 3 Super and the newer mixture-of-experts models in NVIDIA's lineup.

Strengths
Offers a balance of model scale and reasonable context length for handling longer documents and multi-turn conversations without the overhead of the larger mixture-of-experts variants.
Best for
General-purpose tasks where a moderately large model with extended context is useful, but where the efficiency gains of mixture-of-experts or the ultra-long context of larger models are not required.
Limitations
Smaller context window than Nemotron 3 Ultra (512K) and the newer Lightning model (262K), and lacks the mixture-of-experts active parameter efficiency of the newer NVIDIA releases in this size class.

Input / 1M

$0.4

Output / 1M

$0.4

Cached input / 1M

Context window

131K

Price history

Price per 1M tokens over timeOutput $0.4; Input $0.4 as of Jun 2026.$0$0.1$0.2$0.3$0.4Output on 11 Jun 2026: $0.4Out $0.4Input on 11 Jun 2026: $0.4In $0.4Jun 2026

Snapshots

Effective Input Output Cached in Note Source
11 Jun 2026 $0.4 $0.4 Imported from OpenRouter openrouter.ai

More from NVIDIA

Report a problem