A multimodal mixture-of-experts model from Baidu with 424B total parameters and 47B active per token, trained on both text and image data.
- Strengths
- Handles both text and image inputs with a large parameter base and efficient routing through mixture-of-experts architecture.
- Best for
- Tasks requiring vision and language understanding across a 123k token context, such as document analysis with images or visual reasoning at scale.
- Limitations
- As a large multimodal MoE model, it may have higher latency and memory requirements compared to smaller alternatives, and performance depends on the quality of image encoding.
Input / 1M
$0.42
Output / 1M
$1.25
Cached input / 1M
—
Context window
123K
Price history
Snapshots
| Effective | Input | Output | Cached in | Note | Source |
|---|---|---|---|---|---|
| 11 Jun 2026 | $0.42 | $1.25 | — | Imported from OpenRouter | openrouter.ai |