Speech-to-text API pricing, per minute of audio.
16 transcription models — native per-minute rates, dated and sourced from provider price pages.
Price table
| Model | Provider | Price | Released | ||
|---|---|---|---|---|---|
|
G
The faster Whisper Large v3 Turbo distilled model on GroqCloud.
—/—I/O
|
Groq | $0.000667 /min as of Jul 2026 | — | ||
|
R
Rev AI's faster, cheaper asynchronous transcription model.
—/—I/O
|
Rev AI | $0.00167 /min as of Jul 2026 | — | ||
|
G
OpenAI's open-weight Whisper Large v3 served on GroqCloud.
—/—I/O
|
Groq | $0.00185 /min as of Jul 2026 | — | ||
|
S
Speechmatics' multilingual speech-to-text model.
—/—I/O
|
Speechmatics | $0.00215 /min as of Jul 2026 | — | ||
|
A
AssemblyAI's mainstream asynchronous transcription model; also its streaming tier.
—/—I/O
|
AssemblyAI | $0.0025 /min as of Jul 2026 | — | ||
| OpenAI | $0.003 /min as of Jul 2026 | — | |||
|
R
Rev AI's asynchronous speech-to-text model.
—/—I/O
|
Rev AI | $0.0033 /min as of Jul 2026 | — | ||
|
A
AssemblyAI's highest-accuracy asynchronous transcription model.
—/—I/O
|
AssemblyAI | $0.0035 /min as of Jul 2026 | — | ||
|
E
ElevenLabs' speech-to-text model, batch and real-time.
—/—I/O
|
ElevenLabs | $0.00367 /min as of Jul 2026 | — | ||
| OpenAI | $0.006 /min as of Jul 2026 | — | |||
| OpenAI | $0.006 /min as of Jul 2026 | — | |||
|
D
Deepgram's streaming-first English ASR model.
—/—I/O
|
Deepgram | $0.0065 /min as of Jul 2026 | — | ||
|
D
Deepgram's flagship English speech-to-text model, batch and streaming.
—/—I/O
|
Deepgram | $0.0077 /min as of Jul 2026 | — | ||
|
G
Gladia's multilingual speech-to-text model, with diarization and features bundled into the rate.
—/—I/O
|
Gladia | $0.0102 /min as of Jul 2026 | — | ||
| $0.016 /min as of Jul 2026 | — | ||||
|
M
Microsoft Azure's speech-to-text service, real-time and batch.
—/—I/O
|
Microsoft Azure | $0.0167 /min as of Jul 2026 | — |