Gradium AI has switched its new text-to-speech model to default across its API and Studio on August 31, 2026. The company reports an 81.0% human-rated pass rate on a 500-sentence hard-case set spanning five languages. This beats Cartesia Sonic 3.6 at 75.1% and ElevenLabs v3 Conversational at 65.4%. Time to first audio is 216 ms at P50 on Coval, 170 ms faster than the model it replaces.
In this article
Is it deployable?
Yes, today, with no migration. Gradium switched the model on as the default across its API and Studio on August 31, 2026. Existing voices, including custom clones, keep working unchanged.
The accuracy number
Gradium built a 500-sentence evaluation set and open-sourced it on Hugging Face under CC BY 4.0: 100 items across 10 criteria in five languages (EN, DE, FR, ES, PT). Seven atomic criteria cover spelling, acronyms, alphanumeric tokens, dates, regular numbers, large and floating numbers, and email. Three composite criteria (Orders, IT Ticket, Claims) stack several of those into one realistic agent turn.
Scoring is human and strict. A sentence passes only if an independent native-speaker rater hears every element pronounced correctly and completely; one dropped digit fails the sentence. Audio was loudness-normalized, order randomized, and raters capped at 40 comparisons with an enforced break.
Pooled across the ten criteria and averaged over the five languages with equal weight: Gradium TTS 81.0%, Cartesia Sonic 3.6 75.1%, ElevenLabs v3 Conversational 65.4%, Fish Audio S2.1 Pro 49.5%, Inworld TTS 1.5 Max 46.5%. All generated in August 2026 with default settings.
The latency number
On Coval’s TTS benchmark, Gradium reports a 216 ms P50 time to first audio, 170 ms faster than the model it replaces. The more useful figure is the spread: a 30 ms p75-p25 interquartile range across 480 runs, the tightest of the five models tested. Cartesia Sonic 3.6 sits at 454 ms median with a 165 ms spread, 36% of its own median, and callers experience tail turns rather than medians.
Gradium is not the fastest model on that chart. Inworld TTS 2 posts a 166 ms median; Fish Audio S2.1 Pro (291 ms) and ElevenLabs v3 Conversational (329 ms) trail Gradium. The claim being made is about joint position: the lowest hard-case failure rate at sub-250 ms first audio, with very little variance.
What it means
For teams building voice agents, this update changes the baseline for reliability in high-stakes interactions. Previously, systems struggled to read phone numbers, email addresses, or reference codes without manual text normalization. Now, the model handles these tokens directly, meaning callers receive correct information without agents needing to rephrase or correct the output. The tight latency spread ensures consistent performance even during peak usage, removing the frustration of variable wait times that plagues many current solutions.
Getting started
Existing users need do nothing. New teams install the Python SDK, point at the WebSocket TTS endpoint, and reuse existing voice IDs. Gradium is offering 1M credits for complete hard-case failure reports on its Discord.




