In this article
Open TTS Leaderboard: Scalable Evaluation for Multilingual Text-to-Speech and Voice Cloning
Current evaluation methods for text-to-speech are fragmented and lack standardisation. While human preference scores like MOS or MUSHRA remain the gold standard, they cannot keep pace with the speed of new model releases. To address this, several arena-based leaderboards have emerged as useful reference points. These arenas present users with outputs from two different models and ask them to choose the better one. Once a sufficient number of votes are collected, an Elo score is calculated using the Bradley–Terry model to rank the models.
Human preference is the ultimate decider, but arenas struggle to scale. This partly explains why open-source models are underrepresented. As of Sep 30, 2026, only 16 of the 92 models on Artificial Analysis were open-weights, with a similar skew on Voice Arena. Adding an API model requires little more than an API key, whereas an open model must be hosted and served by the arena operator. Commercial providers have more incentive to seek placement than open-source authors. Another limitation is voter consistency. No arena can ensure the same voters apply the same criteria over time. Even the preferences of a single person change over time.
To address this, the Open TTS Leaderboard uses objective metrics to evaluate models on complementary aspects of performance:
- Intelligibility: word/character error rate (WER and CER) between the prompt and the generated audio’s transcript, using Qwen3 ASR (top ranking open-source model on the Open ASR Leaderboard).
- Speed: inverse real-time factor (RTFx) for batched offline inference on an H200 GPU, and time-to-first-audio (TTFA) for quantifying streaming batch size 1 latency on an H200 GPU and CPU.
- Speaker similarity by computing the cosine similarity (SIM) between WavLM speaker embeddings of the generated audio and the reference clip.
Relying on objective metrics reduces evaluation time from a couple of weeks to a couple of hours.
The Open TTS Leaderboard does not replace human preference ranking. ASR-based WER provides a proxy for intelligibility, while speaker similarity estimates voice identity preservation. Neither directly measures naturalness, expressiveness, or listener preference. Nevertheless, they can inform voting-based leaderboards which models to include in their evaluations.
The intention is for the leaderboard to be shaped by the community. Feedback is welcome to keep evaluations relevant. The following sections provide an overview of main features.
Multilingual + voice cloning evaluation
From the default view, models are ranked by macro-average WER on the English splits of Seed TTS Eval and CV3 Eval. hexgrad/Kokoro-82M, Supertone/supertonic-3, and fishaudio/s2-pro lead the pack on English WER when averaged on these two splits. Pareto plots visualise which models strike a good balance between WER, batched inference (RTFx), and size.
English performance does not necessarily translate to other languages. Multiple languages can be toggled to rank models on multilingual performance. Seed TTS Eval only has audio for English and Chinese, so the other languages are simply the score on CV3 Eval. Note that Chinese, Japanese, and Korean are character-based languages and so character error rate (CER) is reported, and the “Average WER” across languages is a macro-average across languages.
k2-fsa/OmniVoice, fishaudio/s2-pro, and FunAudioLLM/Fun-CosyVoice3-0.5B-2512 are strong multilingual models.
By toggling “Voice cloning”, the models that support this functionality can be compared.
A SIM column for speaker similarity now appears in the table, as well as two more Pareto plots for visualising the tradeoff between SIM, batched inference, and size.
The average WER of some models, such as bosonai/higgs-tts-3-4b and openbmb/VoxCPM2, improves under voice cloning, namely when a reference audio is provided.
Compare and vote on TTS outputs
Numbers only tell part of the story. Human preference is the ultimate decider. From the “Listen” tab, you can compare the generated outputs behind the metrics to find which model you prefer.
Pick the language/dataset you are interested in, whether you want to compare voice cloning, and optionally pick the models or listen to outputs from a random selection.
The “Listen” tab fills an important gap in existing TTS leaderboards: a space to explore model outputs of various models.
You can even give feedback on the generated outputs. As more votes are collected, this data may be included on the leaderboard. Vote! Please login with your HF account to help weed out spam/bots.
Streaming performance
The “Streaming” tab compares streaming capabilities. Models are ranked by TTFA, which quantifies how long a user waits after probing a model to obtain audio that can be played. This is important for voice agents and other interactive apps.
For streaming models, it is the time until the first audio chunk arrives. For non-streaming models, it is the time until the whole utterance is generated, because playback cannot start any earlier. Every model runs one audio at a time on the same 50 English prompts from CV3-Eval, on the same hardware and in its default voice. The first 3 runs are dropped as warm-up and the median TTFA is reported.
The default view compares performance on an H200 GPU. Results are also available for CPU for a small set of models.
kyutai/pocket-tts is a great model for streaming on both GPU and CPU.
Conclusion
The goal is to keep up with the pace of TTS model releases and to be shaped by the community. Feedback is welcome. Let us know which datasets, models, and metrics you want to see.
For now, the focus is on:
- Open-source models, to put forward many great models that have been neglected by arena-style evaluations.
- Multilingual, since English performance is not a suitable proxy for other languages.
Soon the evaluation scripts will be open-sourced, much like the Open ASR Leaderboard repo, so that you can directly provide feedback via GitHub Issues and PRs. Let’s shape TTS evaluations together.




