Best speech-to-text models
Transcription models by word error rate, lowest first.
GPT-4o Mini Transcribe, OpenAI leads from OpenAI, ahead of 56 others.
- 1OpenAI
GPT-4o Mini Transcribe, OpenAI
0.00 word error rate
- 2Speechmatics
Speechmatics Enhanced
0.00 word error rate
- 3Mistral
Voxtral Small, Mistral
0.00 word error rate
ranked by word error rate · 57 models · updated Aug 30 · 1–50
| # | Model | Creator | Word error rate |
|---|---|---|---|
| 1 | GPT-4o Mini Transcribe, OpenAI | OpenAI | 0.00 |
| 2 | Speechmatics Enhanced | Speechmatics | 0.00 |
| 3 | Voxtral Small, Mistral | Mistral | 0.00 |
| 4 | Solaria-1, Gladia | Gladia | 0.00 |
| 5 | Smallest AI Pulse Pro | Smallest.ai | 0.00 |
| 6 | Gemini 3 Flash (High), Google | 0.00 | |
| 7 | Nova 2 Pro, Amazon | Amazon | 0.00 |
| 8 | StepAudio 2.5 ASR, StepFun | StepFun | 0.00 |
| 9 | GPT-4o Transcribe, OpenAI | OpenAI | 0.00 |
| 10 | GPT Transcribe, OpenAI | OpenAI | 0.00 |
| 11 | Gemini 2.0 Flash Lite, Google | 0.00 | |
| 12 | Gemini 3.1 Pro Preview (High) | 0.00 | |
| 13 | Fun-Realtime-ASR-preview | Alibaba | 0.00 |
| 14 | Inworld STT 1 | Inworld | 0.00 |
| 15 | MAI-Transcribe-1 | Microsoft AI | 0.00 |
| 16 | Chirp 3, Google | 0.00 | |
| 17 | Whisper Large v2, OpenAI | OpenAI | 0.00 |
| 18 | Universal, AssemblyAI | AssemblyAI | 0.00 |
| 19 | Gemini 3.1 Flash-Lite Preview (Minimal) | 0.00 | |
| 20 | Qwen3.5 Omni Plus | Alibaba | 0.00 |
| 21 | Soniox V4 | Soniox | 0.00 |
| 22 | Gemini 2.5 Pro, Google | 0.00 | |
| 23 | Gemini 3.1 Pro Preview (Low) | 0.00 | |
| 24 | Voxtral Mini Transcribe 2, Mistral | Mistral | 0.00 |
| 25 | Scribe v2, ElevenLabs | ElevenLabs | 0.00 |
| 26 | Universal-3 Pro, AssemblyAI | AssemblyAI | 0.00 |
| 27 | Whisper Large v3 Turbo, OpenAI | OpenAI | 0.00 |
| 28 | Cohere Transcribe | Cohere | 0.00 |
| 29 | Smallest AI Pulse | Smallest.ai | 0.00 |
| 30 | MAI-Transcribe-1.5 | Microsoft AI | 0.00 |
| 31 | LLM Speech, Azure | Microsoft | 0.00 |
| 32 | Modulate STT Batch English VFast | Modulate | 0.00 |
| 33 | Universal-3.5 Pro, AssemblyAI | AssemblyAI | 0.00 |
| 34 | Grok Speech to Text, SpaceXAI | SpaceXAI | 0.00 |
| 35 | Melia | Speechmatics | 0.00 |
| 36 | Resonant-1 | Reson8 | 0.00 |
| 37 | Canary Qwen 2.5B, NVIDIA | NVIDIA | 0.00 |
| 38 | Solaria-3, Gladia | Gladia | 0.00 |
| 39 | Amazon Transcribe | Amazon | 0.00 |
| 40 | Gemini 3.5 Transcribe | 0.00 | |
| 41 | Inkling (256K), Thinking Machines | Thinking Machines | 0.00 |
| 42 | Soniox v5 Async | Soniox | 0.00 |
| 43 | Whisper Large v3, OpenAI | OpenAI | 0.00 |
| 44 | Gemini 2.5 Flash, Google | 0.10 | |
| 45 | Nova 2 Omni, Amazon | Amazon | 0.10 |
| 46 | Gemini 2.5 Lite, Google | 0.10 | |
| 47 | Speechmatics Standard | Speechmatics | 0.10 |
| 48 | Parakeet RNNT 1.1B, NVIDIA | NVIDIA | 0.10 |
| 49 | Parakeet TDT 0.6B V2, NVIDIA | NVIDIA | 0.10 |
| 50 | Gemma 4 12B, Google | 0.10 |
what word error rate means
The share of words a model gets wrong when transcribing. Lower is better, so this board ranks upward from zero.
