TLDR 👉 新的语音合成排行榜专注于开源和多语言
开源文本转语音(TTS)模型的发布速度令人惊叹。截至2026年9月30日,在Hugging Face Hub上已有超过8000个可用的TTS模型🚀
然而,评估工作并未跟上这一节奏:它仍然分散且缺乏标准化。黄金标准是人类偏好评分,如MOS或MUSHRA(详见指标部分)。为此,几个基于竞技场模式的排行榜已确立为社区有用的参考点:
TTS Arena v2
Artificial Analysis
Voice Arena
这些竞技场通过向用户展示来自两个模型的TTS输出,并让他们选择其中一个来比较模型。在收集到足够的投票后,计算Elo分数以对模型进行排名,通常使用Bradley–Terry模型(参见Voice Arena方法论)。
虽然人类偏好是最终的裁决者,但竞技场无法扩展以跟上TTS发布的节奏。这可能在一定程度上解释了为什么开源模型在竞技场式排行榜中代表性不足:截至2026年9月30日,Artificial Analysis上的92个模型中仅有16个是开源权重的,Voice Arena上也存在类似的偏差。这可能反映了实际因素:添加API模型只需一个API密钥,而开源模型必须由竞技场运营商托管和提供服务,且商业提供商比开源作者更有理由寻求上榜。竞技场式评估的另一个局限性在于投票者的一致性:没有任何竞技场能确保同一批具有相同“更好”标准的投票者能够随着时间的推移一致地评估模型。即使是同一个人的偏好也会随时间变化(正如赫拉克利特那句名言所说:“人不能两次踏进同一条河流”)。
为此,我们构建了Open TTS Leaderboard,它使用客观指标从性能的不同互补方面评估模型:
可懂度:提示与生成音频转录之间的词/字符错误率(WER和CER),使用Qwen3 ASR(Open ASR Leaderboard上排名最高的开源模型)。
速度:在H200 GPU上进行批量离线推理的逆实时因子(RTFx),以及在H200 GPU和CPU上量化流式批处理大小为1的延迟的时间到首个音频(TTFA)。
说话人相似度:通过计算生成音频与参考片段的WavLM说话人嵌入之间的余弦相似度(SIM)得出。
依靠客观指标评估,模型评估时间从几周(用于收集投票)缩短到几小时⚡
重要的是,Open TTS Leaderboard并不取代人类偏好排名。基于ASR的WER提供可懂度的代理指标,而说话人相似度估计语音身份保留程度。它们都不能直接衡量自然度、表现力或听众偏好。尽管如此,它们甚至可以为基于投票的排行榜提供哪些模型应纳入评估的建议。
我们建立这个排行榜的初衷是让它由社区塑造;我们希望听到您的反馈,以便评估保持相关性和洞察力。接下来的几节将概述Open TTS Leaderboard的主要功能。
多语言+语音克隆评估
从排行榜的默认视图来看,模型根据Seed TTS Eval(论文)和CV3 Eval(零样本)(论文)的英语分片的宏观平均WER进行排名。
在综合这两个分片上的英语WER时,hexgrad/Kokoro-82M、Supertone/supertonic-3和fishaudio/s2-pro领先,而帕累托图则可视化了哪些模型在WER、批量推理(RTFx)和模型大小之间取得了良好的平衡。
英语表现并不一定适用于其他语言。可以通过切换语言来对多语言性能进行模型排名。Seed TTS Eval 仅包含英语和中文的音频,因此其他语言的得分直接来自 CV3 Eval(零样本)。请注意,中文、日文和韩文属于基于字符的语言,因此报告的是字符错误率(CER),而跨语言的“平均 WER”是各语言之间的宏平均值。
k2-fsa/OmniVoice 、 fishaudio/s2-pro 以及 FunAudioLLM/Fun-CosyVoice3-0.5B-2512 是强大的多语言模型。
通过切换“语音克隆”,可以比较支持该功能(在所选语言上)的模型。
此外,表格中现在出现了用于说话人相似度的 SIM 列,以及两个额外的帕累托图,用于可视化 SIM、批量推理和模型大小之间的权衡。
某些模型(如 bosonai/higgs-tts-3-4b 和 openbmb/VoxCPM2 )的平均 WER 在启用语音克隆时有所改善,即在提供参考音频的情况下。
比较并投票评选 TTS 输出
数字仅能说明故事的一部分,正如前面提到的,人类偏好才是最终的裁决者。通过“Listen”标签页,您可以比较指标背后的生成输出,找出您更喜欢的模型!
选择您感兴趣的语言/数据集,无论您是否想比较语音克隆,并可选择特定模型或收听随机选择的输出。
“Listen”标签页填补了现有 TTS 排行榜的一个重要空白:一个探索各种模型输出的空间。
您甚至可以对生成的输出提供反馈。随着我们收集到更多社区投票,我们可能会将这些数据纳入排行榜。所以请投票!但请使用您的 HF 账户登录,以帮助我们剔除垃圾信息/机器人。
流式传输性能
“Streaming”标签页比较了流式传输能力。模型根据 TTFA(首音频时间)进行排名,该指标量化了用户探测模型后到获得可播放音频所需等待的时间。这对于语音代理和其他交互式应用非常重要。
对于流式模型(“Streaming API”下显示✅),这是直到第一个音频块到达的时间。对于非流式模型,这是直到整个语句生成完毕的时间,因为播放无法更早开始。每个模型在相同的 50 个 CV3-Eval 英语提示上以单次音频(批量大小为 1)运行,使用相同的硬件和默认语音。我们舍弃前 3 次运行作为预热,并报告其余部分的 TTFA 中位数。
默认视图比较的是在 H200 GPU 上的性能。对于一小部分(但正在增长)的模型,也提供了 CPU 的结果!
kyutai/pocket-tts 是一款在 GPU 和 CPU 上流式传输表现极佳的模型!
结论
Open TTS Leaderboard 的目标不仅是跟上 TTS 模型发布的惊人速度,还要由社区塑造;我们希望听到您的反馈,以使评估保持相关性和洞察力。请告诉我们您希望看到哪些数据集、模型和指标!
目前,我们专注于:
开源模型 ,以突出许多被竞技场式评估所忽视的优秀模型。
多语言 ,因为英语表现不能作为其他语言的合适代理。
我们将很快开源评估脚本,类似于 Open ASR Leaderboard 仓库 ,以便您可以直接通过 GitHub Issues 和 PRs 提供反馈和建议!让我们一起塑造 TTS 评估 🤗
TLDR 👉 new TTS leaderboard focused on open-source and multilingual
The pace of open-source text-to-speech (TTS) model releases has been incredible. On the Hugging Face Hub (as of Sep 30, 2026) there are more than
8K TTS models available 🚀
Evaluation, however, hasn't kept pace: it remains fragmented and unstandardized. The gold standard is human preference scores such as MOS or MUSHRA (more on metrics ). To this end, several arena-based leaderboards have established themselves as useful reference points for the community:
TTS Arena v2
Artificial Analysis
Voice Arena
These arenas compare models by presenting users with TTS outputs from two models, and asking them to choose one over the other. After collecting a sufficient number of votes, an Elo score is computed to rank models, typically with the Bradley–Terry model (see Voice Arena methodology ).
While human preference is the ultimate decider, arenas cannot scale to keep up with the pace of TTS releases . This may partly explain why open-source models are underrepresented on arena-style leaderboards: as of Sep 30, 2026, only 16 of the 92 models on Artificial Analysis are open-weights, with a similar skew on Voice Arena . This likely reflects practical factors: adding an API model requires little more than an API key, whereas an open model must be hosted and served by the arena operator, and commercial providers have more reason to seek placement than open-source authors. Another limitation with arena-style evaluation is voter consistency: no arena can ensure that the same voters with the same criteria of “better” can consistently evaluate models over time. Even the preferences of a single person change over time (“A man cannot step into the same river twice” as famously said by Heraclitus).
To this end, we've built the Open TTS Leaderboard , which uses objective metrics to evaluate models on complementary aspects of performance:
Intelligibility : word/character error rate (WER and CER) between the prompt and the generated audio's transcript, using Qwen3 ASR (top ranking open-source model on the Open ASR Leaderboard ).
Speed : inverse real-time factor (RTFx) for batched offline inference on an H200 GPU, and time-to-first-audio (TTFA) for quantifying streaming batch size 1 latency on an H200 GPU and CPU.
Speaker similarity by computing the cosine similarity (SIM) between WavLM speaker embeddings of the generated audio and the reference clip.
By relying on objective metrics evaluating a model drops from a couple weeks (for collecting votes) to a couple hours ⚡
Importantly, the Open TTS Leaderboard does not replace human preference ranking. ASR-based WER provides a proxy for intelligibility, while speaker similarity estimates voice identity preservation. Neither directly measures naturalness, expressiveness, or listener preference. Nevertheless, they can even inform voting-based leaderboards which models to include in their evaluations.
Our intention with this leaderboard is for it to be shaped by the community ; we want to hear your feedback so the evaluations stay relevant and insightful. The next few sections give an overview of main features of the Open TTS Leaderboard.
Multilingual + voice cloning evaluation
From the default view of the leaderboard, models are ranked by macro-average WER on the English splits of Seed TTS Eval ( paper ) and CV3 Eval (zero shot) ( paper ).
hexgrad/Kokoro-82M , Supertone/supertonic-3 , and fishaudio/s2-pro lead the pack on English WER when averaged on these two splits, while the Pareto plots visualize which models strike a good balance between WER, batched inference (RTFx), and size.
English performance doesn't necessarily translate to other languages. Multiple languages can be toggled to rank models on multilingual performance. Seed TTS Eval only has audio for English and Chinese, so the other languages are simply the score on CV3 Eval (zero shot). Note that Chinese, Japanese, and Korean are character-based languages and so character error rate (CER) is reported, and the “Average WER” across languages is a macro-average across languages.
k2-fsa/OmniVoice , fishaudio/s2-pro , and FunAudioLLM/Fun-CosyVoice3-0.5B-2512 are strong multilingual models.
By toggling “Voice cloning”, the models that support this functionality (on the selected languages) can be compared.
Moreover, a SIM column for speaker similarity now appears in the table, as well as two more Pareto plots for visualizing the tradeoff between SIM, batched inference, and size.
The average WER of some models, such as bosonai/higgs-tts-3-4b and openbmb/VoxCPM2 , improve under voice cloning, namely when a reference audio is provided.
Compare and vote on TTS outputs
Numbers only tell part of the story, and as mentioned earlier human preference is the ultimate decider . From the “Listen” tab, you can compare the generated outputs that are behind the metrics, to find which model(s) you prefer!
Pick the language/dataset you're interested in, whether you want to compare voice cloning , and optionally pick the models or listen to outputs from a random selection.
The “Listen” tab fills an important gap in existing TTS leaderboards: a space to explore model outputs of various models.
You can even give feedback on the generated outputs. As we collect more votes from the community, we may include this data on the leaderboard. So vote! But please login with your HF account to help us weed out spam/bots.
Streaming performance
The “Streaming” tab compares the streaming capabilities. Models are ranked by TTFA (time-to-first-audio), which quantifies how long a user waits after probing a model in order to obtain audio that can be played. This is important for voice agents and other interactive apps.
For streaming models (✅ under “Streaming API”) it's the time until the first audio chunk arrives. For non-streaming models, it's the time until the whole utterance is generated, because playback can't start any earlier. Every model runs one audio at a time (batch size 1), on the same 50 English prompts from CV3-Eval, on the same hardware and in its default voice. We drop the first 3 runs as warm-up and report the median TTFA across the rest.
The default view compares performance on an H200 GPU. Results are also available for CPU for a small (but growing) set of models!
kyutai/pocket-tts is a great model for streaming on both GPU and CPU!
Conclusion
The goal of the Open TTS Leaderboard is not only to keep up with the incredible pace of TTS model releases, but to be shaped by the community; we want to hear your feedback so the evaluations stay relevant and insightful. Let us know which datasets, models, and metrics you want to see!
For now, we've focused on:
Open-source models , to put forward many great models that have been neglected by arena-style evaluations.
Multilingual , since English performance is not a suitable proxy for other languages.
We will soon open-source the evaluation scripts, much like the Open ASR Leaderboard repo , so that you can directly provide your feedback and suggestions via GitHub Issues and PRs! Let's shape TTS evaluations together 🤗
| 刊期 | 得分 | 排名 | 结果 |
|---|---|---|---|
| 2026-10-05 | 8.4 | 18 | 入选 |
| 2026-10-04 | 8.49 | 31 | 未入选 |
| 2026-10-03 | 8.77 | 38 | 未入选 |
| 2026-10-02 | 9.22 | 36 | 未入选 |