Google has launched Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, two new text-to-speech models within its Gemini Audio family. Both tools allow developers to direct delivery line by line using natural language instructions. Flash TTS focuses on creative direction and character voices, while Flash-Lite TTS targets high-volume, cost-efficient production.
In this article
The models are available now via the Gemini API and Google AI Studio. Access remains API-only; there are no open weights for self-hosting. Enterprise API access through Gemini Enterprise is listed as coming soon.
Two tiers, shared controls
The release splits the technology into two tiers. Both share the same direction controls.
- Gemini 3.8 Flash TTS targets deep creative direction and character design. Use cases include gaming, immersive audiobooks, podcasts and interactive media. It offers granular control over acting cues, pacing, dialect shifts and backchanneling.
- Gemini 3.8 Flash-Lite TTS targets high-volume, cost-efficient scale. Google positions it for dubbing, audio content creation and expressive voice agents. It offers fine-grained control over tone, pacing and expressive nuance.
In AI Studio, the playground links use the model identifiers gemini-3.8-flash-tts and gemini-3.8-flash-lite-tts.
Voice design from a text prompt
Previous Gemini TTS offered 30 original voices. The 3.8 release moves to a much larger voice system.
- Generative voice design: Flash TTS creates new voices from prompts describing role, accent and voice characteristics. This works across more than 100 languages and dialects. Google’s demos include a Melbourne DJ, a monotone robot and a Japanese dragon.
- Voice library: Developers get 2,000+ production-ready voices. Coverage includes regional varieties like Mexican Spanish, Quebec French and Scots English.
- Save and scale: Custom voices can be saved and reused, with minimal drift across projects.
- Voice remixing (coming soon): Users will adjust a library voice’s timbre, pitch, pace and accent through prompts.
Directing the performance
Both models accept stage directions written in the script. Gemini can also steer delivery from natural script cues.
- Long-form generation: Voice quality, pacing and timbre hold across hours of continuous audio.
- Native 2-speaker staging: A single script drives a multi-turn conversation with distinct, separated voices.
- Vocal bursts: Non-verbal cues like <laughs>, <sigh> and <gasp> add conversational texture.
- Backchanneling: Active-listening interjections like |mhm| and |yeah| control reaction beats and comedic timing.
Voice replication and safety controls
Voice replication builds a consistent vocal profile from a 30-second audio sample. The sample must be your voice or one you have rights to use. Replication requires a verbal consent recording from the voice owner, matched against the reference speaker.
Every clip from Gemini Audio models carries a SynthID watermark. This imperceptible mark is embedded directly in the audio output. Replicated voices also carry C2PA content credentials. Google points to the Gemini 3.8 Audio model card for its broader safety approach.
Benchmark results
Google reports these results for the new models:
- Hume AI Voice Design Benchmark: Flash TTS ranks #1 overall with a score of 71.4, per Hume AI.
- Accent modeling: Flash TTS leads with a score of 60.8.
- Hume AI Overall Quality Index: Flash TTS ranks #1 and Flash-Lite TTS ranks #2.
- Voice Arena blind preference: Both models take top positions in Japanese, Brazilian Portuguese, Vietnamese, Modern Standard Arabic, Mexican Spanish and Hindi.
What it means
Developers can now write stage directions in plain English to control pacing, emotion and character identity. This removes the need to manually select from a fixed list of pre-recorded voices. The ability to generate specific accents and character archetypes on demand makes production faster for gaming and dubbing workflows. However, voice cloning remains restricted to those with explicit consent recordings, and the technology is not available in the UK, EEA or India for that specific feature.




