prices synced 2026-09-23

DeepSeek V4.1 Flash (batch)

DeepSeek · Multimodal · Released Sep 2026

Compare

A sparse mixture-of-experts model from DeepSeek built on its Causal Encoder-Decoder architecture, activating 8 billion parameters on input and 16 billion on output from a 284-billion-parameter pool, with a 1-million-token context window.

Strengths
Delivers frontier-class capabilities at low inference cost through aggressive parameter sparsity and efficient routing, with the extended context window enabling processing of long documents and multi-turn interactions.
Best for
Batch processing workloads where latency is not a constraint and cost-per-token efficiency matters, or applications requiring reasoning over extended documents within the million-token window.
Limitations
As the newer iteration of the V4 Flash line, this model trades some established track record for incremental improvements; sparse activation means certain complex reasoning tasks may benefit from denser model alternatives like V4 Pro.

Input / 1M

$0.112

Output / 1M

$0.336

Cached input / 1M

$0.0034

Context window

1.05M

Price history

Price per 1M tokens over timeOutput $0.336; Input $0.112 as of Sep 2026.$0$0.1$0.2$0.3Output on 22 Sep 2026: $0.336Out $0.336Input on 22 Sep 2026: $0.112In $0.112Sep 2026

Snapshots

Effective Input Output Cached in Note Source
22 Sep 2026 $0.112 $0.336 $0.0034 Imported from OpenRouter openrouter.ai

More from DeepSeek

Report a problem