Xiaomi's native omnimodal model that processes text, images, and video in a single forward pass.
- Strengths
- Handles multimodal perception across image and video understanding with efficient inference, requiring roughly half the compute of comparable models.
- Best for
- Agentic workflows and applications that need to reason over mixed text, image, and video inputs without separate modality pipelines.
- Limitations
- 32000-token context window may be restrictive for tasks requiring long document processing or extended video sequences.
Input / 1M
$0.14
Output / 1M
$0.28
Cached input / 1M
$0.0028
Context window
1.05M
Price history
Input (solid)Output (dashed)
Price change
- 30d
- in increased 33.3% out decreased 0.0%
- 90d
- in decreased 0.0% out decreased 0.0%
- 1y
- in decreased 0.0% out decreased 0.0%
- Since launch
- in decreased 0.0% out decreased 0.0%
Snapshots
| Effective | Input | Output | Cached in | Note | Source |
|---|---|---|---|---|---|
| 15 Jul 2026 | $0.14 | $0.28 | $0.0028 | Imported from OpenRouter | openrouter.ai |
| 7 Jul 2026 | $0.105 | $0.28 | $0.028 | Imported from OpenRouter | openrouter.ai |
| 24 Jun 2026 | $0.105 | $0.28 | — | Imported from OpenRouter | openrouter.ai |
| 11 Jun 2026 | $0.14 | $0.28 | $0.0028 | Imported from OpenRouter | openrouter.ai |