Post-Cutoff.com
  1. Home
  2. Models
  3. Qwen-Audio-3.1-ASR (Flash)

Qwen-Audio-3.1-ASR (Flash)

Alibaba (Qwen)currentaudio/speechQwen-Audio 3.1

Secondary sources report 30 languages + Chinese dialects and ~160 ms latency (unverified). Sibling Qwen-Audio-3.1-ASR-Next adds speaker diarization with timestamps, emotion and sound-event detection (API id not verified). Previous: qwen-audio-3.0-asr-flash; open-weights alternative Qwen3-ASR (see qwen3-asr). Pricing not verified on an official page.

Input
audio
Output
text
License
proprietary
Verified
2026-09-29

How to call it

ProviderModel idEndpoint / URLDocs
Alibaba Cloud Model Studio (streaming)qwen-audio-3.1-asr-flash-streaming—docs
Alibaba Cloud Model Studio / QwenCloud (file transcription)qwen-audio-3.1-asr-flash-filetrans—docs

Notable capabilities (1)

Timeline entry

  1. Alibaba launches Qwen-Audio-3.1 five-model voice stack and cuts audio API prices up to 95% ★★★

    Around its 2026 Apsara Conference Alibaba's Qwen team released Qwen-Audio-3.1: upgraded ASR, TTS and full-duplex Realtime models plus two new ones (ASR-Next for audio understanding, TTS-Next for one-pass speech+SFX+ambience generation), with price cuts of ~70% (TTS), ~85% (Realtime) and up to 95%…

Other Alibaba (Qwen) models

Qwen-Audio-3.1-Realtime (Plus) · Qwen-Audio-3.1-TTS-Next · Qwen3.8-LiveTranslate (Flash Realtime) · Qwen3.8-Omni-Flash · Qwen3.8-27B · Qwen3.8-Flash · Qwen3.8-Max · Qwen-Image-3.0 (Pro) · Qwen3.7-Plus · Qwen3-ASR (0.6B / 1.7B) + Qwen3-ForcedAligner · Qwen3-TTS (open weights 0.6B / 1.7B; API qwen3-tts-flash)