Cartesia has released Sonic-3.6, a new version of its real-time text-to-speech model that now sits at the top of both Artificial Analysis speech leaderboards. The update arrived three months after Sonic-3.5, focusing on naturalness and independent verification.
In this article
Performance metrics
The model currently holds #1 on the Provider Voice board with a score of 1,283 Elo and 1,123 on the Controlled Voice board. The second figure carries more weight because that board clones every model onto the same eight reference voices. This setup isolates the synthesis engine from the voice catalog.
Sonic-3.6 leads the Controlled Voice board, with Sonic-3.5 in second place and ElevenLabs Eleven v3 third.
Deployment and availability
Cartesia states the model is available in beta as a hosted API. It is not offered as self-hosted weights. Sonic is a closed, commercial product with no open weights and no Hugging Face repository. Users rent access.
Targeted at solo developers, startups, scaleups running contact centers, and regulated enterprises requiring DPAs, BAAs, and SSO, the service covers financial services, healthcare, retail, logistics, recruiting, SaaS support, consumer companion apps, and media localization.
Typical applications include inbound support agents, outbound qualification calls, IVR replacement, appointment reminders, sales-training simulators, audio localization, and in-product voice UI.
Under the hood
Sonic runs on state space models rather than transformers. Cartesia frames the usual tradeoffs between speed and naturalness, or accuracy and cost, as architectural choices rather than inevitable limitations.
The practical output is time-to-first-audio. Cartesia claims sub-90ms TTS latency and 100ms transcript latency for its Ink-2 speech-to-text model. Both figures are vendor-stated model latency, not measured end-to-end round trips.
Production features
Sonic exposes controls built for agent transcripts rather than narration:
- Inline expression tags. Non-verbal expressions like [laughter] go directly in the transcript.
- Instant voice cloning from about 10 seconds of audio.
- Custom pronunciation dictionaries, including IPA overrides such as <<s|ə|ˈ|p|i|n|ə>> for subpoena.
- Speed, volume, and emotion parameters exposed through the API and integrations like the LiveKit Agents plugin.
- Native alphanumerics. Order numbers, phone numbers, and confirmation codes read correctly without preprocessing.
Cartesia’s launch demos show English with natural pauses and filler words, plus Hinglish code-switching between Hindi and English.
Cost
Artificial Analysis normalises Sonic 3.6 at $49.00 per 1M characters. That is half the cost of ElevenLabs Eleven v3 at $100.00, and well above Speechify Simba 3.2 at $10.00 for a 1,240 Elo score.
Cartesia sells credits, not characters. The Scale plan at $299 per month includes roughly 10,667 TTS minutes and 15 concurrent requests. Line voice agents bill separately at $0.06 per minute.
What it means
Developers building voice agents get a tool that scores highest on independent benchmarks for engine quality, not just voice variety. The sub-90ms latency claims require careful benchmarking of the full round trip, but the architecture suggests a viable path for real-time interaction. Pricing sits in the mid-range, offering a cheaper alternative to ElevenLabs for high-volume use cases.




