Alibaba’s Qwen team has released Qwen-Audio-3.1, a five-model audio stack covering speech recognition, text-to-speech, and real-time interaction. The primary component is Qwen-Audio-3.1-Realtime, a full-duplex voice model designed for agents that trigger tools. The launch includes price cuts of approximately 85% for Realtime, 70% for TTS, and up to 95% for ASR.
In this article
Availability and Access
The system is available as a managed API on QwenCloud via WebSocket. The specific endpoint is qwen-audio-3.1-realtime-plus. No open weights have been announced.
What Ships on QwenCloud
The model page lists both text and audio as inputs and outputs. The context window holds 262,000 tokens, with a maximum input of 245,000 and a maximum output of 16,000. Default limits cap usage at 60 requests and 100,000 tokens per minute. Pricing is set at $6.4 per million audio input tokens and $0.8 per million text input tokens. Output costs $24 per million tokens, though output text is not charged separately. Key features include function calling, web search, structured outputs, context caching, and fine-tuning.
A companion model, Qwen-Audio-3.1-ASR-Flash-Filetrans, targets offline transcription of long audio files. It supports hot words, speaker separation, punctuation, and recognition of multiple languages plus Chinese dialects. Costs are $0.15 for input and $0.47 for output per million tokens.
Architecture: Two Models Behind One Voice
The system runs two models sharing the same Audio Encoder and Large Language Model design. A full-duplex decision model predicts whether to keep listening, speak, stop, or resume. A speech-to-text model generates the response content as text. A context-aware voice renderer then converts that text into streaming speech, conditioning on conversation history, voice cues, and acoustic context.
Training is organised into three layers: Think, Act, and Speak and Coordinate.
Think: M²-OPD
Core-Cocktail SFT re-anchors the audio model to its source text LLM using million-hour-scale paired data. Multimodality OPD follows. A Text Teacher and a frozen Audio Reference score each token of the student’s own trajectory. This is on-policy distillation, not imitation of pre-written answers. Domain experts for empathy, pragmatic intent, and acoustic scenes are then trained with GRPO. Multi-Teacher OPD merges them into one deployable model.
Act: Executable Environments
Each training domain bundles a tool pool, a stateful JSON database, and a natural-language business policy. Domains are seeded from open-source tool and MCP server definitions. Every task defines one of three outcomes: a write, a justified refusal, or an unsupported request. Scoring checks terminal state, then permitted writes, then behavioural assertions. A fluent reply cannot rescue a failed state check.
GRPO receives rewards at dialogue, milestone, and turn levels. Search training penalises redundant queries using the formula rquery = q min(1, nref / npred). Mean queries per search call fell from 4.37 to 1.05. Trigger F1 slipped from 60.87% to 58.61%.
Speak and Coordinate
This layer decides whether, when, and how to speak. On Full-Duplex-Bench v1.5, replies to people talking to someone else fell from 0.13 to 0.03. On v3.0, the filler rate dropped from 0.7590 to 0.2960. There are trade-offs. After interruptions, the unwanted resume rate rose from 0.035 to 0.130. Interruption stop latency is 1.116 seconds, versus 0.383 for GPT-Realtime-2.
Interactive Explainer
Explore the Think, Act, Speak loop, duplex decisions, a scored training episode, and the search reward.
Benchmarks at a Glance
Against version 3.0, Audio MultiChallenge rises from 47.12 to 52.21. The 14-language BBA average climbs from 81.7% to 88.1%. FLEURS WER falls from 9.01 to 3.98. The τ-Voice figures use a half-duplex speech-to-text adaptation. They are not comparable to official full-duplex results. GPT-Realtime-2 still leads the 50-session human red-team study, with 96.00% versus 92.00%.
How It Compares
| Feature | Qwen-Audio-3.1-Realtime-Plus | OpenAI GPT-Realtime-2 | Google Gemini 3.8 Live |
|---|---|---|---|
| Input | Text, audio | Text, audio, image | Text, images, audio, video |
| Output | Text, audio | Text, audio | Text and audio |
| Context / input limit | 262K | 128K | 131,072 |
| Max output | 16K | 32K | 65,536 |
| Function calling | Yes | Yes | Yes (async by default) |
| Built-in web search | Yes | Not listed | Google Search grounding |
| Reasoning | Thinking mode, 2K max reasoning | Configurable effort | Interleaved reasoning |
| Audio input / 1M tokens | $6.4 | $32 | $3.00 |
| Audio output / 1M tokens | $24 (text and audio) | $64 | $12.00 |
| Open weights | No | No | No |
| Source | QwenCloud | OpenAI docs | Model page, pricing |
Prices are list rates checked September 28, 2026. Token rates are not directly comparable, since each provider tokenises audio differently.
Key Takeaways
- Task success on a τ-Voice adaptation rises from 78.4% to 82.0% over version 3.0.
- Replies to background speech drop from 73.0% to 13.0% on Full-Duplex-Bench v1.5.
- 262K context, function calling and web search at $6.4 per 1M audio input tokens.
- Tool use is trained with GRPO inside self-evolving executable environments.
- Multi-turn attack success falls to 26.0% (Chinese) and 23.5% (English).
FAQ
- What is Qwen-Audio-3.1-Realtime? A full-duplex speech model from Alibaba’s Qwen team for voice agents that reason, call tools and manage turn-taking.
- Can I self-host it? No open weights were announced. Access is through the QwenCloud API as
qwen-audio-3.1-realtime-plus. - How much does it cost? $6.4 per 1M audio input tokens and $24 per 1M output tokens for text and audio.
Read the Paper, check Qwen-Audio-3.1-ASR, and view Qwen-Audio-3.1-Realtime. All credit goes to the researcher of this project.




