MiniMax M3 is a multimodal foundation model that processes text, image, and video inputs to produce text output, with a 524K token context window.
- Strengths
- Handles video alongside text and image inputs, enabling analysis and reasoning across multiple modalities in a single request.
- Best for
- Workflows that require multimodal understanding, such as video analysis, image-to-text reasoning, or document interpretation with visual content.
- Limitations
- Batch-only access limits it to offline, non-interactive use cases where results are needed after processing delay.
Input / 1M
$0.3
Output / 1M
$1.20
Cached input / 1M
$0.06
Context window
524K
Price history
Snapshots
| Effective | Input | Output | Cached in | Note | Source |
|---|---|---|---|---|---|
| 11 Jun 2026 | $0.3 | $1.20 | $0.06 | Imported from OpenRouter | openrouter.ai |