GPT Audio
OpenAI · Omnimodal · Released Jan 2026
A multimodal model from OpenAI that processes text, images, and audio inputs with a 128k token context window.
- Strengths
- Handles audio alongside text and image inputs, enabling workflows that require speech transcription, audio analysis, or multimedia reasoning in a single model.
- Best for
- Applications that combine audio processing with text or image understanding, such as transcription with context, audio-driven Q&A, or multimodal analysis tasks.
- Limitations
- Smaller context window than text-focused peers and less capable than GPT-5.3 Chat or newer generations on pure language reasoning tasks; GPT Audio Mini offers a lighter variant if audio capability alone is needed.
Input / 1M
$2.50
Output / 1M
$10.00
Cached input / 1M
—
Context window
128K
Also billed
- Audio input
- $32.00 / 1M
Price history
Input (solid)Output (dashed)
Price change
- 30d
- in — out —
- 90d
- in decreased 0.0% out decreased 0.0%
- 1y
- in decreased 0.0% out decreased 0.0%
- Since launch
- in decreased 0.0% out decreased 0.0%
Snapshots
| Effective | Input | Output | Cached in | Note | Source |
|---|---|---|---|---|---|
| 29 Jun 2026 | $2.50 | $10.00 | — | Imported from OpenRouter | openrouter.ai |
| 11 Jun 2026 | $2.50 | $10.00 | — | Imported from OpenRouter | openrouter.ai |