prices synced 2026-08-12
N

Nemotron 3.5 Lightning

NVIDIA · Released Aug 2026

Compare

A 30-billion-parameter mixture-of-experts model from NVIDIA with 3B active parameters per token and a 262K token context window.

Strengths
Fast inference throughput due to its small active parameter count relative to total model size, making it well-suited for latency-sensitive applications.
Best for
High-throughput agentic workloads and specialized tasks where inference speed is prioritized over reasoning depth.
Limitations
The 3B active parameter constraint means it trades off reasoning capability and task complexity handling compared to larger NVIDIA models like Nemotron 3 Ultra or Nemotron 3 Super.

Input / 1M

$0.1

Output / 1M

$0.25

Cached input / 1M

$0.05

Context window

262K

Price history

Price per 1M tokens over timeOutput $0.25; Input $0.1 as of Aug 2026.$0$0.1$0.2Output on 11 Aug 2026: $0.25Out $0.25Input on 11 Aug 2026: $0.1In $0.1Aug 2026

Snapshots

Effective Input Output Cached in Note Source
11 Aug 2026 $0.1 $0.25 $0.05 Imported from OpenRouter openrouter.ai

More from NVIDIA

Report a problem