Voxtral Small 24B 2507
Mistral · Multimodal · Released Oct 2025
A 24B multimodal model from Mistral that processes both text and audio inputs, with a 32K token context window.
- Strengths
- Handles speech transcription, translation, and audio understanding alongside traditional text tasks without significant performance degradation on text.
- Best for
- Applications that need to process audio content (transcription, translation, or analysis) while maintaining text-based reasoning capabilities.
- Limitations
- The 32K context window is relatively constrained for handling long documents or extensive conversation history, and audio processing performance may vary depending on audio quality and complexity.
Input / 1M
$0.1
Output / 1M
$0.3
Cached input / 1M
$0.01
Context window
32K
Also billed
- Audio input
- $100.00 / 1M
Price history
Input (solid)Output (dashed)
Price change
- 30d
- in decreased 0.0% out decreased 0.0%
- 90d
- in decreased 0.0% out decreased 0.0%
- 1y
- in decreased 0.0% out decreased 0.0%
- Since launch
- in decreased 0.0% out decreased 0.0%
Snapshots
| Effective | Input | Output | Cached in | Note | Source |
|---|---|---|---|---|---|
| 29 Jun 2026 | $0.1 | $0.3 | $0.01 | Imported from OpenRouter | openrouter.ai |
| 11 Jun 2026 | $0.1 | $0.3 | $0.01 | Imported from OpenRouter | openrouter.ai |