Qwen-Audio-3.1-Realtime (Plus)
Languages: de, en, es, fr, id, it, ja, ko, pt, ru, zh (Mandarin, Cantonese and 18+ Chinese varieties). Predecessors qwen-audio-3.0-realtime-plus / -flash (July 2026) still listed. Press (MarkTechPost) reports interruption-stop latency 1.116 s vs 0.383 s for GPT-Realtime-2 and higher red-team refusal for GPT-Realtime-2; not verified on an official page. Release date is the announcement date (Qwen X post / Apsara); Model Studio pricing for this id not verified.
- Context window
- 262,144 tokens
- Max output
- 16,384 tokens
- Input
- audio, text
- Output
- audio, text
- License
- proprietary
- Pricing
- audio input: $6.4 · text input: $0.8 · text output: $6.4 · audio output: $24 (USD per 1M tokens (QwenCloud list price; audio/text output $24 when audio is generated, $6.40 text-only)) source
- Verified
- 2026-09-29
How to call it
| Provider | Model id | Endpoint / URL | Docs |
|---|---|---|---|
| QwenCloud (Realtime WebSocket) | qwen-audio-3.1-realtime-plus | wss://maas.qwencloudapi.com/api-ws/v1/realtime?model=qwen-audio-3.1-realtime-plus | docs |
| Alibaba Cloud Model Studio (Singapore / Beijing) | qwen-audio-3.1-realtime-plus | wss://{WorkspaceId}.sg-singapore.maas.aliyuncs.com/api-ws/v1/realtime | docs |
Notable capabilities (3)
- Full-duplex agentic voice ("Think, Act, Speak and Coordinate"): Listens while speaking, decides whether to keep listening, speak, stop or resume; function calling and built-in web search. Task success 82.0% vs 78.4% for the previous version; replies to background speech fell from 73.0% to 13.0% (Full-Duplex-Bench v1.5). source
- Three turn-taking modes and voice cloning: server_vad, semantic smart_turn and push-to-talk modes; system voices plus cloned custom voices; 16 kHz PCM in, 24 kHz PCM out. source
- ~85% price cut at launch: Alibaba cut Realtime prices about 85% with the 3.1 release (TTS ~70%, ASR up to 95%). source
Alibaba's hosted real-time voice agent model (WebSocket Realtime API, OpenAI-Realtime-style events).
Sources: https://www.qwencloud.com/models/qwen-audio-3.1-realtime-plus · https://help.aliyun.com/en/model-studio/qwen-audio-realtime-user-guides · https://arxiv.org/abs/2609.25176
Timeline entry
- Alibaba launches Qwen-Audio-3.1 five-model voice stack and cuts audio API prices up to 95% ★★★
Around its 2026 Apsara Conference Alibaba's Qwen team released Qwen-Audio-3.1: upgraded ASR, TTS and full-duplex Realtime models plus two new ones (ASR-Next for audio understanding, TTS-Next for one-pass speech+SFX+ambience generation), with price cuts of ~70% (TTS), ~85% (Realtime) and up to 95%…
Other Alibaba (Qwen) models
Qwen-Audio-3.1-ASR (Flash) · Qwen-Audio-3.1-TTS-Next · Qwen3.8-LiveTranslate (Flash Realtime) · Qwen3.8-Omni-Flash · Qwen3.8-27B · Qwen3.8-Flash · Qwen3.8-Max · Qwen-Image-3.0 (Pro) · Qwen3.7-Plus · Qwen3-ASR (0.6B / 1.7B) + Qwen3-ForcedAligner · Qwen3-TTS (open weights 0.6B / 1.7B; API qwen3-tts-flash)