StepFun's multimodal Mixture-of-Experts model that processes text, images, and video with a 196B parameter backbone while activating approximately 11B parameters per inference.
- Strengths
- Efficiently handles multimodal inputs including video understanding while maintaining low computational overhead through sparse activation.
- Best for
- Applications requiring fast image and video analysis alongside text processing, particularly where latency and token efficiency matter.
- Limitations
- As a sparse model, it may not match dense models on complex reasoning tasks requiring sustained deep computation across very long contexts.
Input / 1M
$0.2
Output / 1M
$1.15
Cached input / 1M
$0.04
Context window
256K
Price history
Snapshots
| Effective | Input | Output | Cached in | Note | Source |
|---|---|---|---|---|---|
| 11 Jun 2026 | $0.2 | $1.15 | $0.04 | Imported from OpenRouter | openrouter.ai |