Qwen3-TTS (open weights 0.6B / 1.7B; API qwen3-tts-flash)
HF repos: Qwen3-TTS-12Hz-{1.7B,0.6B}-{Base,CustomVoice}, 1.7B-VoiceDesign, Qwen3-TTS-Tokenizer-12Hz. API snapshots: qwen3-tts-flash (=2025-11-27), qwen3-tts-flash-2025-09-18, qwen3-tts-instruct-flash-2026-01-26, qwen3-tts-vd-2026-01-26 (voice design), qwen3-tts-vc-2026-01-22 (voice clone). Superseded in Alibaba's hosted lineup by Qwen-Audio-3.0-TTS (Jul 2026) and Qwen-Audio-3.1-TTS (Sep 2026). API pricing not verified.
- Input
- text, audio
- Output
- audio
- License
- apache-2.0
- Verified
- 2026-09-29
How to call it
| Provider | Model id | Endpoint / URL | Docs |
|---|---|---|---|
| Hugging Face | Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice | huggingface.co/Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice | — |
| Hugging Face | Qwen/Qwen3-TTS-12Hz-1.7B-VoiceDesign | huggingface.co/Qwen/Qwen3-TTS-12Hz-1.7B-VoiceDesign | — |
| GitHub | — | github.com/QwenLM/Qwen3-TTS | — |
| Alibaba Cloud Model Studio | qwen3-tts-flash | — | docs |
| Alibaba Cloud Model Studio (instruct / voice design / voice clone) | qwen3-tts-instruct-flash | — | docs |
Notable capabilities (2)
- Open-weights voice design and 3-second cloning: Voice design from natural-language descriptions and voice cloning from ~3 s of audio, in 10 languages (zh, en, ja, ko, de, fr, ru, pt, es, it). source
- 97 ms streaming latency: 12 Hz multi-codebook tokenizer; first audio packet after a single input character, end-to-end latency as low as 97 ms; one model for streaming and non-streaming. source
Apache-2.0 multilingual TTS you can run locally; hosted versions on Model Studio.
Sources: https://github.com/QwenLM/Qwen3-TTS · https://arxiv.org/abs/2601.15621 · https://www.alibabacloud.com/help/en/model-studio/qwen-tts
Timeline entry
- Alibaba open-sources Qwen3-TTS (voice design, 3-second cloning, 97 ms streaming) and, a week later, Qwen3-ASR ★★★
On 2026-01-22 Alibaba's Qwen team released Qwen3-TTS under Apache-2.0 (0.6B and 1.7B checkpoints plus a 12 Hz tokenizer). It offers voice design from text descriptions, voice cloning from about 3 s of audio in 10 languages, and ~97 ms streaming latency. On 2026-01-29 Qwen3-ASR followed (0.6B/1.7B…
Other Alibaba (Qwen) models
Qwen-Audio-3.1-ASR (Flash) · Qwen-Audio-3.1-Realtime (Plus) · Qwen-Audio-3.1-TTS-Next · Qwen3.8-LiveTranslate (Flash Realtime) · Qwen3.8-Omni-Flash · Qwen3.8-27B · Qwen3.8-Flash · Qwen3.8-Max · Qwen-Image-3.0 (Pro) · Qwen3.7-Plus · Qwen3-ASR (0.6B / 1.7B) + Qwen3-ForcedAligner