prices synced 2026-09-11

GPT Audio

OpenAI · Omnimodal · Released Jan 2026

Compare

A multimodal model from OpenAI that processes text, images, and audio inputs with a 128k token context window.

Strengths
Handles audio alongside text and image inputs, enabling workflows that require speech transcription, audio analysis, or multimedia reasoning in a single model.
Best for
Applications that combine audio processing with text or image understanding, such as transcription with context, audio-driven Q&A, or multimodal analysis tasks.
Limitations
Smaller context window than text-focused peers and less capable than GPT-5.3 Chat or newer generations on pure language reasoning tasks; GPT Audio Mini offers a lighter variant if audio capability alone is needed.

Input / 1M

$2.50

Output / 1M

$10.00

Cached input / 1M

Context window

128K

Also billed

Audio input
$32.00 / 1M

Price history

Input (solid)Output (dashed)
Price per 1M tokens over timeOutput went from $10.00 to $10.00; Input went from $2.50 to $2.50 between Jun 2026 and Jun 2026.$0$2.00$4.00$6.00$8.00$10.00Output on 11 Jun 2026: $10.00Output on 29 Jun 2026: $10.00Out $10.00Input on 11 Jun 2026: $2.50Input on 29 Jun 2026: $2.50In $2.50Jun 2026Jun 2026

Price change

30d
in out
90d
in decreased 0.0% out decreased 0.0%
1y
in decreased 0.0% out decreased 0.0%
Since launch
in decreased 0.0% out decreased 0.0%

Snapshots

Effective Input Output Cached in Note Source
29 Jun 2026 $2.50 $10.00 Imported from OpenRouter openrouter.ai
11 Jun 2026 $2.50 $10.00 Imported from OpenRouter openrouter.ai

More from OpenAI

Report a problem