MiniMax M3 is a multimodal foundation model that processes text, image, and video inputs to produce text output, with a 524K token context window.
- Strengths
- Handles long contexts and supports multiple input modalities including video, making it capable of processing diverse content types in a single request.
- Best for
- Extended-context tasks like video analysis, long-document coding work, and agentic systems that need to reason over mixed media.
- Limitations
- Limited to text output only, so it cannot generate images or video; performance on specialized domains may vary compared to dedicated single-modality models.
Input / 1M
$0.3
Output / 1M
$1.20
Cached input / 1M
$0.06
Context window
524K
Price history
Snapshots
| Effective | Input | Output | Cached in | Note | Source |
|---|---|---|---|---|---|
| 11 Jun 2026 | $0.3 | $1.20 | $0.06 | Imported from OpenRouter | openrouter.ai |