Alibaba’s Qwen Audio 3.0 TTS Plus tops the competition in the text-to-speech rankings

Alibaba Cloud Model Studio has released Qwen-Audio-3.0-TTS-Plus, a new text-to-speech model that currently ranks first on the Artificial Analysis Speech Arena leaderboard…

By Vane July 21, 2026 1 min read
Alibaba’s Qwen Audio 3.0 TTS Plus tops the competition in the text-to-speech rankings

Alibaba Cloud Model Studio has released Qwen-Audio-3.0-TTS-Plus, a new text-to-speech model that currently ranks first on the Artificial Analysis Speech Arena leaderboard with an Elo score of 1,236. This places it two points ahead of the nearest competitor, Simba 3.2, while also surpassing Gemini 3.1 Flash TTS and Sonic 3.5 in the current standings. The system operates in two distinct versions, with the Plus variant prioritising high-fidelity output over speed. It supports sixteen languages, including several Chinese dialects and less common options such as Tagalog, Malay, Thai, and Vietnamese. Users can control speaking styles using natural language instructions or specific tags like [angry] or [giggles]. The model also claims improved performance when cloning voices from noisy or echo-heavy reference recordings.

The primary distinction lies in the trade-off between quality and latency. While the Flash version targets real-time interaction with approximately 300 milliseconds of delay, the Plus version sacrifices speed for clarity. At sixteen characters per second, the Plus model lags significantly behind Sonic 3.5 at 120 characters and Simba 3.2 at 30.2 characters. Pricing for the service is set at $27.60 per million characters. This positions the tool for specific use cases requiring precise vocal nuance rather than instant response times.

  • Supports sixteen languages including Chinese dialects and Tagalog
  • Latency of 16 characters per second for the Plus version
  • Costs $27.60 per million characters via Alibaba Cloud
Scroll to Top