DeepSeek V4.1 Flash (batch)
DeepSeek · Multimodal · Released Sep 2026
A sparse mixture-of-experts model from DeepSeek built on its Causal Encoder-Decoder architecture, activating 8 billion parameters on input and 16 billion on output from a 284-billion-parameter pool, with a 1-million-token context window.
- Strengths
- Delivers frontier-class capabilities at low inference cost through aggressive parameter sparsity and efficient routing, with the extended context window enabling processing of long documents and multi-turn interactions.
- Best for
- Batch processing workloads where latency is not a constraint and cost-per-token efficiency matters, or applications requiring reasoning over extended documents within the million-token window.
- Limitations
- As the newer iteration of the V4 Flash line, this model trades some established track record for incremental improvements; sparse activation means certain complex reasoning tasks may benefit from denser model alternatives like V4 Pro.
Input / 1M
$0.112
Output / 1M
$0.336
Cached input / 1M
$0.0034
Context window
1.05M
Price history
Snapshots
| Effective | Input | Output | Cached in | Note | Source |
|---|---|---|---|---|---|
| 22 Sep 2026 | $0.112 | $0.336 | $0.0034 | Imported from OpenRouter | openrouter.ai |