Alibaba Qwen Releases Qwen-Audio-3.1-Realtime: A Full-Duplex Voice Model Trained to Think, Act, and Decide When to Speak

Disclosure: Some links in this article are affiliate links. AI Maestro may earn a commission if you make a purchase. We do…

By Vane September 29, 2026 4 min read
Alibaba Qwen Releases Qwen-Audio-3.1-Realtime: A Full-Duplex Voice Model Trained to Think, Act, and Decide When to Speak

Alibaba’s Qwen team has released Qwen-Audio-3.1, a five-model audio stack covering speech recognition, text-to-speech, and real-time interaction. The primary component is Qwen-Audio-3.1-Realtime, a full-duplex voice model designed for agents that trigger tools. The launch includes price cuts of approximately 85% for Realtime, 70% for TTS, and up to 95% for ASR.

Availability and Access

The system is available as a managed API on QwenCloud via WebSocket. The specific endpoint is qwen-audio-3.1-realtime-plus. No open weights have been announced.

What Ships on QwenCloud

The model page lists both text and audio as inputs and outputs. The context window holds 262,000 tokens, with a maximum input of 245,000 and a maximum output of 16,000. Default limits cap usage at 60 requests and 100,000 tokens per minute. Pricing is set at $6.4 per million audio input tokens and $0.8 per million text input tokens. Output costs $24 per million tokens, though output text is not charged separately. Key features include function calling, web search, structured outputs, context caching, and fine-tuning.

A companion model, Qwen-Audio-3.1-ASR-Flash-Filetrans, targets offline transcription of long audio files. It supports hot words, speaker separation, punctuation, and recognition of multiple languages plus Chinese dialects. Costs are $0.15 for input and $0.47 for output per million tokens.

Architecture: Two Models Behind One Voice

The system runs two models sharing the same Audio Encoder and Large Language Model design. A full-duplex decision model predicts whether to keep listening, speak, stop, or resume. A speech-to-text model generates the response content as text. A context-aware voice renderer then converts that text into streaming speech, conditioning on conversation history, voice cues, and acoustic context.

Training is organised into three layers: Think, Act, and Speak and Coordinate.

Think: M²-OPD

Core-Cocktail SFT re-anchors the audio model to its source text LLM using million-hour-scale paired data. Multimodality OPD follows. A Text Teacher and a frozen Audio Reference score each token of the student’s own trajectory. This is on-policy distillation, not imitation of pre-written answers. Domain experts for empathy, pragmatic intent, and acoustic scenes are then trained with GRPO. Multi-Teacher OPD merges them into one deployable model.

Act: Executable Environments

Each training domain bundles a tool pool, a stateful JSON database, and a natural-language business policy. Domains are seeded from open-source tool and MCP server definitions. Every task defines one of three outcomes: a write, a justified refusal, or an unsupported request. Scoring checks terminal state, then permitted writes, then behavioural assertions. A fluent reply cannot rescue a failed state check.

GRPO receives rewards at dialogue, milestone, and turn levels. Search training penalises redundant queries using the formula rquery = q min(1, nref / npred). Mean queries per search call fell from 4.37 to 1.05. Trigger F1 slipped from 60.87% to 58.61%.

Speak and Coordinate

This layer decides whether, when, and how to speak. On Full-Duplex-Bench v1.5, replies to people talking to someone else fell from 0.13 to 0.03. On v3.0, the filler rate dropped from 0.7590 to 0.2960. There are trade-offs. After interruptions, the unwanted resume rate rose from 0.035 to 0.130. Interruption stop latency is 1.116 seconds, versus 0.383 for GPT-Realtime-2.

Interactive Explainer

Explore the Think, Act, Speak loop, duplex decisions, a scored training episode, and the search reward.

Benchmarks at a Glance

Against version 3.0, Audio MultiChallenge rises from 47.12 to 52.21. The 14-language BBA average climbs from 81.7% to 88.1%. FLEURS WER falls from 9.01 to 3.98. The τ-Voice figures use a half-duplex speech-to-text adaptation. They are not comparable to official full-duplex results. GPT-Realtime-2 still leads the 50-session human red-team study, with 96.00% versus 92.00%.

How It Compares

FeatureQwen-Audio-3.1-Realtime-PlusOpenAI GPT-Realtime-2Google Gemini 3.8 Live
InputText, audioText, audio, imageText, images, audio, video
OutputText, audioText, audioText and audio
Context / input limit262K128K131,072
Max output16K32K65,536
Function callingYesYesYes (async by default)
Built-in web searchYesNot listedGoogle Search grounding
ReasoningThinking mode, 2K max reasoningConfigurable effortInterleaved reasoning
Audio input / 1M tokens$6.4$32$3.00
Audio output / 1M tokens$24 (text and audio)$64$12.00
Open weightsNoNoNo
SourceQwenCloudOpenAI docsModel page, pricing

Prices are list rates checked September 28, 2026. Token rates are not directly comparable, since each provider tokenises audio differently.

Key Takeaways

  • Task success on a τ-Voice adaptation rises from 78.4% to 82.0% over version 3.0.
  • Replies to background speech drop from 73.0% to 13.0% on Full-Duplex-Bench v1.5.
  • 262K context, function calling and web search at $6.4 per 1M audio input tokens.
  • Tool use is trained with GRPO inside self-evolving executable environments.
  • Multi-turn attack success falls to 26.0% (Chinese) and 23.5% (English).

FAQ

  • What is Qwen-Audio-3.1-Realtime? A full-duplex speech model from Alibaba’s Qwen team for voice agents that reason, call tools and manage turn-taking.
  • Can I self-host it? No open weights were announced. Access is through the QwenCloud API as qwen-audio-3.1-realtime-plus.
  • How much does it cost? $6.4 per 1M audio input tokens and $24 per 1M output tokens for text and audio.

Read the Paper, check Qwen-Audio-3.1-ASR, and view Qwen-Audio-3.1-Realtime. All credit goes to the researcher of this project.

Scroll to Top