prices synced 2026-07-28
N

Llama 3.3 Nemotron Super 49B V1.5

NVIDIA · Released Oct 2025

Compare

A 49B-parameter model optimized for reasoning and agentic workflows, derived from Llama 3.3 and supporting 131K token context.

Strengths
Handles structured tasks like tool calling, code generation, and RAG workflows effectively through post-training on reasoning and function-calling patterns.
Best for
Building agents and applications that require tool use, code execution, retrieval-augmented generation, or multi-step reasoning across a large document context.
Limitations
Designed primarily for English and reasoning tasks; may not perform as well on creative writing, translation, or non-English languages compared to its larger parent model.

Input / 1M

$0.4

Output / 1M

$0.4

Cached input / 1M

Context window

131K

Price history

Price per 1M tokens over timeOutput $0.4; Input $0.4 as of Jun 2026.$0$0.1$0.2$0.3$0.4Output on 11 Jun 2026: $0.4Out $0.4Input on 11 Jun 2026: $0.4In $0.4Jun 2026

Snapshots

Effective Input Output Cached in Note Source
11 Jun 2026 $0.4 $0.4 Imported from OpenRouter openrouter.ai

More from NVIDIA

Report a problem