In this article
ElevenLabs launches v4 speech model with better direction-following and voice consistency
ElevenLabs has released Eleven v4, a speech model designed to follow direction cues more accurately and maintain vocal consistency across extended productions. A new architecture also powers the Turbo variant, which targets real-time voice agents.
The update generates laughter, whispers, and sound effects like slamming doors more reliably than its predecessor. The Turbo variant begins producing speech in about 150 milliseconds. Eleven v3, released just over a year ago, already supported these audio tags but followed them less accurately.
Improved consistency and character limits
ElevenLabs says v4 uses a new architecture that analyses a script’s tone, pacing, and context. Users can give directions through tags or plain sentences and use phonetic spelling to set the pronunciation of names and technical terms. The company says pronunciation controls now work more reliably. Narrators and characters should sound consistent throughout a production, even when users regenerate individual lines several times.
Eleven v4 handles up to 10,000 characters per request, roughly ten minutes of audio. Longer works like audiobooks use multiple segments, with pacing and delivery expected to stay consistent across transitions. In dialogue, AI speakers respond to the context of the entire scene rather than delivering each line in isolation.
The new version supports more than 90 languages, up from about 70 with v3. Cloned voices should speak other languages with native accents without drifting back to their original accents over time. Professional Voice Clones work again after being unsupported in v3, while an “Instant Voice Clone” needs only ten seconds of audio.
Turbo aims for faster expressive voice agents
ElevenLabs is also releasing a faster v4 variant for real-time uses like customer service calls or game characters. The company says voice agent developers previously had to choose between speed and expression, and v4 Turbo is meant to offer both.
In ElevenLabs’ tests, Turbo starts producing audible speech in 150 milliseconds, compared with 262 milliseconds for Cartesia Sonic 3.6. OpenAI‘s GPT-4o mini TTS takes 814 milliseconds. ElevenLabs optimized Turbo together with its ElevenAgents platform.
Eleven v4 ranks ahead of Cartesia Sonic 3.6 and Google’s Gemini 3.8 Flash TTS on Artificial Analysis’ Provider Voice Arena leaderboard. It scores 91.7 percent on the pronunciation benchmark, up from v3’s 85.6 percent. In ElevenLabs’ blind tests, about three-quarters of listeners preferred v4 over models from Cartesia, Inworld, and Google.
Higher-quality voice clones make v4 useful for dubbing, with ElevenLabs promoting the ability to use an actor’s voice across all supported languages. The company licenses voices from the people behind them. Voice actors can offer an extensively trained clone of their voice in ElevenLabs’ library and earn money when paying users use it. ElevenLabs already offers access to celebrity voices like Michael Caine’s through a dedicated marketplace.
Launch prices and data handling
ElevenLabs’ standard API pricing is $80 per million characters for v4 and $40 for Turbo. Through October 12, those rates fall to $22 and $11. According to ElevenLabs, users on the $22 monthly Creator plan or higher can use v4 in ElevenCreative at no extra cost for two weeks. Usage is capped at twice their monthly credits. Artificial Analysis lists Sonic 3.6 at $49 per million characters and Gemini 3.8 Flash TTS at $16.49.
Both models are available now in ElevenAgents, ElevenCreative, and through the API. ElevenLabs stores customer data in the US by default, according to its documentation. Enterprise customers can store data in isolated environments in the EU, India, or Singapore, though some processing may take place outside the chosen region. In the EU, customers can keep API processing within the region by using a mode that does not retain data.
ElevenLabs also released its Music 2.5 model in mid-September, designed to produce denser, more natural-sounding songs.
What it means
For people making audio, the main change is that long-form projects no longer require constant manual adjustments to keep a voice sounding the same. Users can now generate longer segments in a single go and rely on the system to maintain tone and pronunciation. Voice actors can monetise their cloned voices more effectively, and developers building voice agents have access to a model that responds quickly while still sounding natural.




