prices synced 2026-09-11
S

Step 3.7 Flash

StepFun · Multimodal · Released May 2026

Compare

Step 3.7 Flash is a sparse Mixture of Experts model from StepFun that activates a subset of its parameters per token, designed for fast inference across a 256K token context.

Strengths
The model balances latency and capability through sparse parameter activation, making it efficient for high-throughput workloads while maintaining a large context window.
Best for
Applications requiring fast response times with long context handling, such as summarization, retrieval-augmented generation, and high-volume inference tasks.
Limitations
As a newer Flash variant released after Step 3.5 Flash, it trades some raw capability density for speed; dense models or those with higher parameter activation may handle complex reasoning more effectively.

Input / 1M

$0.2

Output / 1M

$1.15

Cached input / 1M

$0.04

Context window

256K

Price history

Price per 1M tokens over timeOutput $1.15; Input $0.2 as of Jun 2026.$0$0.5$1.00Output on 11 Jun 2026: $1.15Out $1.15Input on 11 Jun 2026: $0.2In $0.2Jun 2026

Snapshots

Effective Input Output Cached in Note Source
11 Jun 2026 $0.2 $1.15 $0.04 Imported from OpenRouter openrouter.ai

More from StepFun

Report a problem