Alibaba releases Qwen3.8-Omni-Flash, an omnimodal agent model with 1M context and 93-98% cheaper audio/video input
Alibaba's Qwen team launched Qwen3.8-Omni-Flash, a native omnimodal model (text, image, audio and video in; 1M-token context) built for audio/video agent work such as video editing, film commentary and meeting summaries. Qwen reports a >25% average gain over Qwen3.5-Omni-Plus on 29 evaluations, audio performance above Gemini 3.8 Flash, and API prices per hour of audio (audio-visual) input more than 98% (93%) lower. It also open-sourced the Qwen-Live Harness and expanded Qwen-MM-Plugins.
Key facts
- Qwen blog index dates the post 2026-09-18 (the page header shows Sept 14 with a [draft] marker); OpenRouter listing 2026-09-21
- Inputs: text, image, audio, video; 1M-token context; text quality comparable to a text-only model of the same size (Qwen)
- Average score up >25% vs Qwen3.5-Omni-Plus across 29 audio, audio-visual and agent benchmarks
- WildClawBench-MM +36.5 pts, AgenticVBench +22.3 pts, UniClawBench 69.6; LongAudioSpan +8.3, OmniVideoBench +9.6
- AliMeeting DER / cpWER fell from 88.11 / 89.61 to 3.35 / 17.18
- Qwen claims audio-visual performance close to Gemini 3.8 Flash and overall audio performance above it
- Price per hour of audio input down >98%, per hour of audio-visual input down >93% (Qwen's methodology: 720p at 1 fps)
- API ids qwen3.8-omni-flash and qwen3.8-omni-flash-realtime; $0.15 input / $0.47 output per 1M tokens (Model Studio Intl, see model file)
- Open-sourced: Qwen-Live Harness (runtime for real-time omnimodal interaction); Video2Note added to Qwen-MM-Plugins
What happened
Qwen3.8-Omni-Flash is the Qwen3.8-generation successor to the Qwen Omni line. Qwen positions it as a step from understanding audio and video to acting on them: planning tasks, calling tools and delivering finished work (edited videos, music videos, commentary, PDF notes from tutorials) on its own. It also improves long-audio understanding and multi-speaker recognition. Alongside the model, Qwen open-sourced Qwen-Live Harness, a runtime for continuous real-time omnimodal interaction, and added plugins such as Video2Note. It is closed-weights and served on Alibaba Cloud Model Studio and the Qianwen app.
Why it matters
The model sells audio and video understanding at Flash-tier prices and competes directly with Gemini 3.8 Flash on the multimodal agents that Chinese and US labs are both racing to ship.
All benchmarks and price comparisons are Qwen's own. The exact release day is uncertain (Sept 14 to 18).
Changelog
- 2026-09-30: created (the model file existed but there was no timeline entry)
Models
- Qwen3.8-Omni-Flash Alibaba (Qwen) · current
Related events
- Alibaba launches Qwen-Audio-3.1 five-model voice stack and cuts audio API prices up to 95% ★★★
- Apsara 2026: Alibaba says Qwen 4 is in training, targets 5-10T-parameter Qwen 4.5/5, reports self-improvement runs and unveils Zhenwu V900 chip ★★★
- Qwen3.8-Flash-Next: 125B MoE with only 6B active previews Qwen 4 architecture ★★★
Sources (4)
- officialQwen - Qwen3.8-Omni-Flash: Omni Senses. Agentic Delivery.
- codeGitHub - QwenLM/Qwen-Live-Harness
- codeGitHub - QwenLM/Qwen-MM-Plugins
- docsAlibaba Cloud Model Studio - Qwen3.8-Omni-Flash
id: 2026-09-18-qwen3-8-omni-flash · updated 2026-09-30 · open in the interactive timeline