Qwen3.8-Omni-Flash enters the market as the first multimodal model designed specifically for AI agents, capable of processing audio and video simultaneously to edit vlogs or summarise films. The system handles a context window of one million tokens and claims performance on audio-video tasks that approaches that of Gemini 3.8 Flash.
The primary distinction lies in cost structure. API pricing stands at $0.15 per million input tokens and $0.47 per million output tokens. This contrasts sharply with Gemini 3.8 Flash, which charges $0.75 for input and $3.75 for output per million tokens at its introductory rate. Qwen estimates audio input costs under $0.01 per hour, whereas 720p video with audio at one frame per second runs about $0.20. Availability extends through Qwen Studio and Qwen Cloud, with open-source plugins enabling video editing and speaker recognition for agents like Claude Code and Gemini CLI.
* Pricing excludes response costs for video processing
* Prices scheduled to double on January 1, 2027
* Qwen-Live Harness enables real-time camera and microphone interaction




