GPT Audio Mini
OpenAI · Omnimodal · Released Jan 2026
A lightweight multimodal model from OpenAI that processes text, images, and audio inputs, positioned as a smaller variant of GPT Audio released in January 2026.
- Strengths
- Handles audio, image, and text inputs within a single inference, making it suitable for tasks that combine multiple input modalities.
- Best for
- Applications requiring audio transcription, description, or analysis alongside text and image inputs where model size and latency matter more than maximum capability.
- Limitations
- The 128k token context window limits performance on long document processing; for reasoning-heavy tasks or extended code sessions, larger models in the lineup offer more capability.
Input / 1M
$0.6
Output / 1M
$2.40
Cached input / 1M
—
Context window
128K
Also billed
- Audio input
- $0.6 / 1M
Price history
Input (solid)Output (dashed)
Price change
- 30d
- in decreased 0.0% out decreased 0.0%
- 90d
- in decreased 0.0% out decreased 0.0%
- 1y
- in decreased 0.0% out decreased 0.0%
- Since launch
- in decreased 0.0% out decreased 0.0%
Snapshots
| Effective | Input | Output | Cached in | Note | Source |
|---|---|---|---|---|---|
| 29 Jun 2026 | $0.6 | $2.40 | — | Imported from OpenRouter | openrouter.ai |
| 11 Jun 2026 | $0.6 | $2.40 | — | Imported from OpenRouter | openrouter.ai |