Alibaba open-sources Qwen3-TTS (voice design, 3-second cloning, 97 ms streaming) and, a week later, Qwen3-ASR
On 2026-01-22 Alibaba's Qwen team released Qwen3-TTS under Apache-2.0 (0.6B and 1.7B checkpoints plus a 12 Hz tokenizer). It offers voice design from text descriptions, voice cloning from about 3 s of audio in 10 languages, and ~97 ms streaming latency. On 2026-01-29 Qwen3-ASR followed (0.6B/1.7B plus a forced aligner, 30 languages and 22 Chinese dialects). Both became among the most-downloaded open speech models of 2026.
Key facts
- Qwen3-TTS repos: Qwen3-TTS-12Hz-{1.7B,0.6B}-{Base,CustomVoice}, 1.7B-VoiceDesign, Qwen3-TTS-Tokenizer-12Hz; tech report arXiv 2601.15621
- Languages (TTS): zh, en, ja, ko, de, fr, ru, pt, es, it; end-to-end latency as low as 97 ms; one model for streaming and non-streaming
- Qwen3-ASR (2026-01-29): 1.7B and 0.6B plus Qwen3-ForcedAligner-0.6B; 30 languages + 22 Chinese dialects, singing/music robust; self-reported AISHELL-2 WER 2.71 vs 5.06 for Whisper-large-v3
- Hugging Face downloads in the month to 2026-09-29: Qwen3-TTS-12Hz-1.7B-CustomVoice ~2.4M, Qwen3-ASR-1.7B ~1.76M
- Hosted equivalents: qwen3-tts-flash / qwen3-tts-instruct-flash on Model Studio; superseded in Alibaba's API lineup by Qwen-Audio-3.0 (Jul 2026) and 3.1 (Sep 2026)
What happened
Qwen released a complete open TTS family with voice design, cloning and low-latency streaming, then an open ASR family with a forced aligner for timestamps a week later, both under Apache-2.0.
Why it matters
Voice cloning and voice design had mostly been proprietary (ElevenLabs and others). Qwen3-TTS made them freely self-hostable, and it became one of the most-downloaded speech models of 2026.
Changelog
- 2026-09-29: created
Models
- Qwen3-ASR (0.6B / 1.7B) + Qwen3-ForcedAligner Alibaba (Qwen) · current
- Qwen3-TTS (open weights 0.6B / 1.7B; API qwen3-tts-flash) Alibaba (Qwen) · current
Related events
- Alibaba's Qwen-Audio-3.0-TTS takes #1 on the Artificial Analysis text-to-speech leaderboard ★★★
- Alibaba launches Qwen-Audio-3.1 five-model voice stack and cuts audio API prices up to 95% ★★★
- Fish Audio open-sources S2: expressive 80+ language TTS with inline emotion tags ★★★
Sources (5)
- codeGitHub - QwenLM/Qwen3-TTS
- paperarXiv 2601.15621 - Qwen3-TTS technical report
- codeHugging Face - Qwen3-TTS-12Hz-1.7B-CustomVoice
- codeGitHub - QwenLM/Qwen3-ASR
- codeHugging Face - Qwen3-ASR-1.7B
id: 2026-01-22-qwen3-tts-asr-open-weights · updated 2026-09-29 · open in the interactive timeline