Best text-to-speech models

Voice models ranked by head-to-head human preference.

ranked by elo · 100 models · updated Aug 30 · 51–100

#ModelElo
51Maya 2 Flash
1,069
52Magpie-Multilingual 357M (Feb 2026)
1,067
53Polly Generative
1,065
54SIMBA 1.6
1,065
55MiMo-V2.5-TTS
1,061
56Maya 2 Global
1,061
57Kokoro 82M v1.0
1,060
58Coda
1,060
59Gemini 2.5 Flash TTS (Dec 2025)
1,058
60Chirp 3: HD
1,057
61Octave 2
1,056
62OpenAudio S1 Mini
1,051
63Async Flash v1.0
1,050
64Maya1
1,046
65Sonic English (Oct 2024)
1,043
66Higgs Audio V3 TTS
1,043
67Polly Long-Form
1,043
68SIMBA 1.0
1,038
69Gemini 2.5 Pro (Dec 2025)
1,037
70T2A-01-Turbo
1,035
71MAI-Voice-1
1,032
72Azure Neural
1,032
73Octave TTS
1,031
74Lightning v3.1
1,023
75MiMo-V2-TTS
1,023
76Chatterbox
1,022
77Raon SpeechLM
1,008
78Arcana v3
1,006
79Magpie-Multilingual 357M
1,004
80Zonos-v0.1
1,000
81Murf Speech Gen 2
981
82LMNT
979
83VibeVoice 1.5B
969
84VibeVoice 7B
969
85OpenVoice v2
955
86Magpie Multilingual
942
87Qwen3 TTS Flash
940
88Neuphonic TTS
935
89Qwen3 TTS
927
90Kugel 3
924
91XTTS v2
922
92WaveNet
914
93Mist V2
898
94StyleTTS 2
893
95Neural2
892
96Polly Neural
890
97Standard
884
98Noiz TTS
869
99MetaVoice v1
845
100Polly Standard
820

what elo means

An Elo rating from head-to-head votes: two outputs are shown side by side and people pick the better one. Higher is better, and a gap of about 100 points means the higher model wins roughly two thirds of the time.