prices synced 2026-09-11

GPT Audio Mini

OpenAI · Omnimodal · Released Jan 2026

Compare

A multimodal model from OpenAI that processes text, images, and audio inputs with a 128k token context window.

Strengths
Handles audio, image, and text inputs in a single model, enabling workflows that combine speech recognition, visual analysis, and language understanding.
Best for
Applications requiring native audio processing alongside text and image analysis, such as transcription with context, audio-visual search, or multimodal document understanding.
Limitations
The 128k token context is smaller than coding-specialized models like GPT-5.3-Codex (400k), and it lacks the advanced reasoning capabilities of the GPT-5.4 frontier series.

Input / 1M

$0.6

Output / 1M

$2.40

Cached input / 1M

Context window

128K

Also billed

Audio input
$0.6 / 1M

Price history

Input (solid)Output (dashed)
Price per 1M tokens over timeOutput went from $2.40 to $2.40; Input went from $0.6 to $0.6 between Jun 2026 and Jun 2026.$0$0.5$1.00$1.50$2.00$2.50Output on 11 Jun 2026: $2.40Output on 29 Jun 2026: $2.40Out $2.40Input on 11 Jun 2026: $0.6Input on 29 Jun 2026: $0.6In $0.6Jun 2026Jun 2026

Price change

30d
in out
90d
in decreased 0.0% out decreased 0.0%
1y
in decreased 0.0% out decreased 0.0%
Since launch
in decreased 0.0% out decreased 0.0%

Snapshots

Effective Input Output Cached in Note Source
29 Jun 2026 $0.6 $2.40 Imported from OpenRouter openrouter.ai
11 Jun 2026 $0.6 $2.40 Imported from OpenRouter openrouter.ai

More from OpenAI

Report a problem