Microsoft AI releases new transcription and text-to-speech models for voice agents

Microsoft AI has launched MAI-Transcribe-2-Streaming, a new model designed for real-time transcription that currently ranks first for accuracy on Artificial Analysis. The…

By Vane October 2, 2026 1 min read
Microsoft AI releases new transcription and text-to-speech models for voice agents

Microsoft AI has launched MAI-Transcribe-2-Streaming, a new model designed for real-time transcription that currently ranks first for accuracy on Artificial Analysis. The system processes audio in sixty languages and returns initial partial results within one hundred milliseconds, allowing voice agents to begin responding before a speaker finishes a sentence. Through the end of the year, an hour of audio processing costs $0.54 at the introductory price. The company also released two new text-to-speech models, with MAI-Voice-2.1 capable of speaking twenty-three languages using a single voice identity that adapts to native accents in each region. A specific variant known as MAI-Voice-2.1-Flash achieves a latency of one hundred fifty milliseconds and reduces costs to $15 per million characters compared to the previous $22 rate. Both models can generate a voice clone from just a few seconds of reference audio and include built-in safeguards intended to prevent misuse. They are available through Microsoft Foundry, the MAI Playground, and OpenRouter. In one test involving four thousand participants, approximately half believed the generated voices belonged to real people.

The practical value lies in reducing the delay between input and output for automated voice interactions. Lower latency allows customer service bots to interrupt or respond faster, which improves the natural flow of conversation. The ability to clone voices from short samples enables more flexible applications without requiring extensive recording sessions.

  • Initial transcription results arrive in under 100 milliseconds
  • MAI-Voice-2.1 supports 23 languages with a single voice
  • Voice cloning requires only a few seconds of reference audio
Scroll to Top