Step 3.7 Flash is a sparse Mixture of Experts model from StepFun that activates a subset of its parameters per token, designed for fast inference across a 256K token context.
- Strengths
- The model balances latency and capability through sparse parameter activation, making it efficient for high-throughput workloads while maintaining a large context window.
- Best for
- Applications requiring fast response times with long context handling, such as summarization, retrieval-augmented generation, and high-volume inference tasks.
- Limitations
- As a newer Flash variant released after Step 3.5 Flash, it trades some raw capability density for speed; dense models or those with higher parameter activation may handle complex reasoning more effectively.
Input / 1M
$0.2
Output / 1M
$1.15
Cached input / 1M
$0.04
Context window
256K
Price history
Snapshots
| Effective | Input | Output | Cached in | Note | Source |
|---|---|---|---|---|---|
| 11 Jun 2026 | $0.2 | $1.15 | $0.04 | Imported from OpenRouter | openrouter.ai |