Best speech-to-text models
Transcription models by word error rate, lowest first.
ranked by word error rate · 57 models · updated Aug 30 · 51–57
| # | Model | Creator | Word error rate |
|---|---|---|---|
| 51 | Gradium Speech-to-Text | Gradium | 0.10 |
| 52 | Nova-3, Deepgram | Deepgram | 0.10 |
| 53 | Qwen3.5 Omni Flash | Alibaba | 0.10 |
| 54 | Chirp 2, Google | 0.10 | |
| 55 | Rev AI | Rev AI | 0.10 |
| 56 | Qwen3 ASR Flash, Alibaba | Alibaba | 0.10 |
| 57 | Cloud Speech-To-Text (Chirp), Google | 0.30 |
what word error rate means
The share of words a model gets wrong when transcribing. Lower is better, so this board ranks upward from zero.
