Gemini 3.8 TTS Playground

Google released two new text-to-speech models today, gemini-3.8-flash-tts and gemini-3.8-flash-lite-tts. These versions offer a library of over 2,000 voices and allow users…

By Vane September 23, 2026 1 min read

Google released two new text-to-speech models today, gemini-3.8-flash-tts and gemini-3.8-flash-lite-tts. These versions offer a library of over 2,000 voices and allow users to create custom voices using a thirty-second audio sample. Simon Willison built a playground interface to test these models, demonstrating how easily they can generate multi-speaker conversations with distinct delivery styles.

The practical value lies in the speed and cost of production for audio content. A one minute and eighteen second clip took approximately twenty seconds to generate at a cost of 2.74 cents. The API handles complex dialogue structures without requiring separate prompts for each speaker.

* Custom voices require only a short audio sample
* Delivery styles are specified directly in the prompt
* Costs remain low for extended generation runs

Scroll to Top