MiniMax M3 is a multimodal foundation model that processes text, image, and video inputs to produce text output, with a 524K token context window.
- Strengths
- Handles multimodal inputs across text, image, and video in a single model, enabling workflows that span multiple content types.
- Best for
- Long-context tasks involving mixed media—analysis of video with text, reasoning over documents with embedded images, or agentic work requiring visual grounding.
- Limitations
- The batch variant is designed for asynchronous processing; latency-sensitive or interactive applications should use standard deployment options.
Input / 1M
$0.3
Output / 1M
$1.20
Cached input / 1M
$0.06
Context window
524K
Price history
Input (solid)Output (dashed)
Price change
- 30d
- in increased 100.0% out increased 100.0%
- 90d
- in increased 100.0% out increased 100.0%
- 1y
- in increased 100.0% out increased 100.0%
- Since launch
- in increased 100.0% out increased 100.0%
Snapshots
| Effective | Input | Output | Cached in | Note | Source |
|---|---|---|---|---|---|
| 19 Aug 2026 | $0.3 | $1.20 | $0.06 | Imported from OpenRouter | openrouter.ai |
| 29 Jul 2026 | $0.15 | $0.6 | $0.03 | Imported from OpenRouter | openrouter.ai |