prices synced 2026-07-28
S

Step 3.7 Flash

StepFun · Multimodal · Released May 2026

Compare

StepFun's multimodal Mixture-of-Experts model that processes text, images, and video with a 196B parameter backbone while activating approximately 11B parameters per inference.

Strengths
Efficiently handles multimodal inputs including video understanding while maintaining low computational overhead through sparse activation.
Best for
Applications requiring fast image and video analysis alongside text processing, particularly where latency and token efficiency matter.
Limitations
As a sparse model, it may not match dense models on complex reasoning tasks requiring sustained deep computation across very long contexts.

Input / 1M

$0.2

Output / 1M

$1.15

Cached input / 1M

$0.04

Context window

256K

Price history

Price per 1M tokens over timeOutput $1.15; Input $0.2 as of Jun 2026.$0$0.5$1.00Output on 11 Jun 2026: $1.15Out $1.15Input on 11 Jun 2026: $0.2In $0.2Jun 2026

Snapshots

Effective Input Output Cached in Note Source
11 Jun 2026 $0.2 $1.15 $0.04 Imported from OpenRouter openrouter.ai

More from StepFun

Report a problem