prices synced 2026-09-11
M

MiniMax M3

MiniMax · Multimodal · Released May 2026

Compare

MiniMax M3 is a multimodal foundation model that processes text, image, and video inputs to produce text output, with a 524K token context window.

Strengths
Handles video alongside text and image inputs, enabling analysis and reasoning across multiple modalities in a single request.
Best for
Workflows that require multimodal understanding, such as video analysis, image-to-text reasoning, or document interpretation with visual content.
Limitations
Batch-only access limits it to offline, non-interactive use cases where results are needed after processing delay.

Input / 1M

$0.3

Output / 1M

$1.20

Cached input / 1M

$0.06

Context window

524K

Price history

Price per 1M tokens over timeOutput $1.20; Input $0.3 as of Jun 2026.$0$0.5$1.00Output on 11 Jun 2026: $1.20Out $1.20Input on 11 Jun 2026: $0.3In $0.3Jun 2026

Snapshots

Effective Input Output Cached in Note Source
11 Jun 2026 $0.3 $1.20 $0.06 Imported from OpenRouter openrouter.ai

More from MiniMax

Report a problem