MiniMax M3 is a multimodal foundation model that processes text, image, and video inputs to produce text output, with a 524K token context window.
- Strengths
- Handles multimodal inputs across text, image, and video in a single model, enabling workflows that span multiple content types.
- Best for
- Long-context tasks involving mixed media—analysis of video with text, reasoning over documents with embedded images, or agentic work requiring visual grounding.
- Limitations
- The batch variant is designed for asynchronous processing; latency-sensitive or interactive applications should use standard deployment options.
Input / 1M
$0.15
Output / 1M
$0.6
Cached input / 1M
$0.03
Context window
524K
Price history
Snapshots
| Effective | Input | Output | Cached in | Note | Source |
|---|---|---|---|---|---|
| 29 Jul 2026 | $0.15 | $0.6 | $0.03 | Imported from OpenRouter | openrouter.ai |