Open TTS Leaderboard: Scalable Evaluation for Multilingual Text-to-Speech and Voice Cloning
Open-source text-to-speech (TTS) model releases have surged, with more than 8,000 TTS models on the Hugging Face Hub as of September 30, 2026. Evaluation, however, has lagged behind, remaining fragmented and unstandardized. Human preference scores like MOS or MUSHRA are the gold standard, and arena-based leaderboards such as TTS Arena v2 and Artificial Analysis Voice Arena have become key references by collecting pairwise user votes and computing Elo scores via the Bradley–Terry model. Yet arenas cannot scale with the pace of TTS releases, which helps explain why open-source models are underrepresented: only 16 of the 92 models on Artificial Analysis are open-weights, with a similar skew on Voice Arena. This is partly because adding an API model requires little more than an API key, while an open model must be hosted and served by the arena operator, and commercial providers have stronger incentives to seek placement than open-source authors. Another limitation is voter consistency—no arena can ensure the same voters with the same criteria evaluate models over time, and even individual preferences shift.
To address these gaps, the Open TTS Leaderboard evaluates models using objective metrics across complementary performance aspects. Intelligibility is measured by word/character error rate (WER and CER) between the prompt and the generated audio's transcript, using Qwen3 ASR, the top-ranking open-source model on the Open ASR Leaderboard. Speed is captured by inverse real-time factor (RTFx) for batched offline inference on an H200 GPU, and by time-to-first-audio (TTFA) for streaming batch size 1 latency on an H200 GPU and CPU. Speaker similarity is computed as cosine similarity (SIM) between WavLM speaker embeddings of the generated audio and the reference clip.
Relying on these objective metrics reduces the evaluation time for a model from a couple of weeks of collecting votes to a couple of hours. Importantly, the Open TTS Leaderboard does not replace human preference ranking: ASR-based WER is only a proxy for intelligibility, and speaker similarity estimates voice identity preservation—neither directly measures naturalness, expressiveness, or listener preference. Nonetheless, the leaderboard can help inform voting-based leaderboards about which models to include in their evaluations. The creators intend for the leaderboard to be shaped by the community.