Google’s Gemini 3.5 Transcribe turns speech to text in 85 languages while auto-correcting your verbal stumbles

Google has launched Gemini 3.5 Transcribe, a new speech-to-text model designed for real-time transcription across over 85 languages. The system automatically strips…

By Vane August 27, 2026 1 min read
Google’s Gemini 3.5 Transcribe turns speech to text in 85 languages while auto-correcting your verbal stumbles

Google has launched Gemini 3.5 Transcribe, a new speech-to-text model designed for real-time transcription across over 85 languages. The system automatically strips filler words and corrects verbal slips while formatting the resulting text. Google reports a word error rate of 4.0 percent for streaming audio and 2.6 percent for recorded input. This represents a 70 percent reduction in latency compared to its predecessor, Chirp 3. Through function calling, the model can delegate tasks such as image generation or web searches to other Gemini models.

The release includes two distinct interfaces for different use cases. The Live API handles real-time streaming with very low latency, while the Interactions API processes recorded audio with speaker attribution and timestamps. The model is currently available in Google AI Studio and the Gemini Enterprise Agent Platform. It has already been integrated into Gboard for Android via the Rambler feature and the Gemini app on macOS, with Chrome support arriving soon. This update addresses practical needs for accuracy and speed in communication tools.

* Live API supports real-time streaming with minimal delay
* Interactions API adds speaker attribution and timestamps
* Chrome browser integration is scheduled for release soon

Scroll to Top