GPT Audio Mini
OpenAI · Omnimodal · Released Jan 2026
A multimodal model from OpenAI that processes text, images, and audio inputs with a 128k token context window.
- Strengths
- Handles audio, image, and text inputs in a single model, enabling workflows that combine speech recognition, visual analysis, and language understanding.
- Best for
- Applications requiring native audio processing alongside text and image analysis, such as transcription with context, audio-visual search, or multimodal document understanding.
- Limitations
- The 128k token context is smaller than coding-specialized models like GPT-5.3-Codex (400k), and it lacks the advanced reasoning capabilities of the GPT-5.4 frontier series.
Input / 1M
$0.6
Output / 1M
$2.40
Cached input / 1M
—
Context window
128K
Also billed
- Audio input
- $0.6 / 1M
Price history
Input (solid)Output (dashed)
Price change
- 30d
- in — out —
- 90d
- in decreased 0.0% out decreased 0.0%
- 1y
- in decreased 0.0% out decreased 0.0%
- Since launch
- in decreased 0.0% out decreased 0.0%
Snapshots
| Effective | Input | Output | Cached in | Note | Source |
|---|---|---|---|---|---|
| 29 Jun 2026 | $0.6 | $2.40 | — | Imported from OpenRouter | openrouter.ai |
| 11 Jun 2026 | $0.6 | $2.40 | — | Imported from OpenRouter | openrouter.ai |