GPT Audio
OpenAI · Omnimodal · Released Jan 2026
GPT Audio is OpenAI's multimodal model that processes text, images, and audio inputs within a 128k token context window.
- Strengths
- It handles audio, image, and text inputs in a single model, enabling workflows that combine speech, visual, and textual data.
- Best for
- Applications requiring native audio processing alongside text and images, such as transcription with context, audio analysis, or multimodal reasoning tasks.
- Limitations
- Its 128k context window is smaller than most text-only variants in OpenAI's lineup, and multimodal capability may come with different performance characteristics than dedicated text or audio specialist models.
Input / 1M
$2.50
Output / 1M
$10.00
Cached input / 1M
—
Context window
128K
Also billed
- Audio input
- $32.00 / 1M
Price history
Input (solid)Output (dashed)
Price change
- 30d
- in decreased 0.0% out decreased 0.0%
- 90d
- in decreased 0.0% out decreased 0.0%
- 1y
- in decreased 0.0% out decreased 0.0%
- Since launch
- in decreased 0.0% out decreased 0.0%
Snapshots
| Effective | Input | Output | Cached in | Note | Source |
|---|---|---|---|---|---|
| 29 Jun 2026 | $2.50 | $10.00 | — | Imported from OpenRouter | openrouter.ai |
| 11 Jun 2026 | $2.50 | $10.00 | — | Imported from OpenRouter | openrouter.ai |