prices synced 2026-07-28
M

MiniMax M3

MiniMax · Multimodal · Released May 2026

Compare

MiniMax M3 is a multimodal foundation model that processes text, image, and video inputs to produce text output, with a 524K token context window.

Strengths
Handles long contexts and supports multiple input modalities including video, making it capable of processing diverse content types in a single request.
Best for
Extended-context tasks like video analysis, long-document coding work, and agentic systems that need to reason over mixed media.
Limitations
Limited to text output only, so it cannot generate images or video; performance on specialized domains may vary compared to dedicated single-modality models.

Input / 1M

$0.3

Output / 1M

$1.20

Cached input / 1M

$0.06

Context window

524K

Price history

Price per 1M tokens over timeOutput $1.20; Input $0.3 as of Jun 2026.$0$0.5$1.00Output on 11 Jun 2026: $1.20Out $1.20Input on 11 Jun 2026: $0.3In $0.3Jun 2026

Snapshots

Effective Input Output Cached in Note Source
11 Jun 2026 $0.3 $1.20 $0.06 Imported from OpenRouter openrouter.ai

More from MiniMax

Report a problem