OpenAI has launched GPT Transcribe and GPT Live Transcribe, two new speech recognition models accessible via its API. GPT Transcribe processes pre-recorded audio files approximately 34 times faster than real time, while GPT Live Transcribe handles real-time streaming with low latency. Independent testing by Artificial Analysis shows GPT Transcribe achieved a word error rate of 3.31 percent. This represents a 0.7 percentage point improvement over the year-old GPT-4o Transcribe model. Pricing for the service dropped 25 percent to land at $0.0045 per minute of audio. Both models accept text as transcription context, keywords, and multiple input languages.
Despite this progress, OpenAI still trails several competitors in the AA-WER ranking. ElevenLabs Scribe v2 leads with a 2.3 percent error rate, followed by Google’s Gemini 3 Pro at 2.9 percent and Mistral’s Voxtral Small at 3 percent. Mistral recently undercut the market with Voxtral Transcribe V2, starting at just $0.003 per minute. The new transcription models complement OpenAI’s recently announced Realtime model generation, which also includes the real-time transcription model GPT-Realtime-Whisper.
* GPT Transcribe costs $0.0045 per minute of audio
* ElevenLabs Scribe v2 holds the lowest error rate at 2.3 percent
* Mistral’s Voxtral Transcribe V2 starts at $0.003 per minute




