Qwen-Audio-3.0-TTS (Flash / Plus)
Flash tier targets real-time use (~300 ms first packet, press); Plus targets quality (throughput ~16 chars/s, press). Languages: ar, zh, en, fr, de, id, it, ja, ko, ms, pt, ru, es, tl, th, vi. Companion qwen-audio-3.0-realtime-plus/-flash and qwen-audio-3.0-asr-flash also exist. Superseded by Qwen-Audio-3.1 (2026-09-23), but as of 2026-09-29 the Model Studio catalog still lists qwen-audio-3.0-tts-plus as its TTS model, and no 3.1 TTS id is published in the international docs.
- Input
- text, audio
- Output
- audio
- License
- proprietary
- Pricing
- per 1m characters: $27.59 (USD per 1M characters for the Plus tier at launch (MarkTechPost). Alibaba cut TTS prices about 70% with the 3.1 generation (2026-09-23), so check current pricing) source
- Verified
- 2026-09-29
How to call it
| Provider | Model id | Endpoint / URL | Docs |
|---|---|---|---|
| Alibaba Cloud Model Studio (Singapore / Beijing) | qwen-audio-3.0-tts-flash | https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api/v1/services/audio/tts/SpeechSynthesizer | docs |
| Alibaba Cloud Model Studio | qwen-audio-3.0-tts-plus | — | docs |
Notable capabilities (2)
- #1 on Artificial Analysis TTS arena at launch: Qwen-Audio-3.0-TTS-Plus ranked first on the Artificial Analysis Text-to-Speech leaderboard in July 2026 (Elo ~1,236-1,237, just ahead of Speechify Simba 3.2 at ~1,234). It was later overtaken (Eleven v4 was #1 by late Sept 2026). source
- Controllable, robust multilingual synthesis: 12.5 Hz speech tokenizer plus a five-stage LM + flow-matching training recipe; natural-language instructions and inline tags; 16 languages and 20 Chinese dialect regions; one-pass long-form output up to 3 minutes; voice cloning works from noisy or reverberant references. source
Alibaba's hosted flagship TTS from July 2026.
Sources: https://arxiv.org/abs/2607.23938 · https://www.alibabacloud.com/help/en/model-studio/qwen-tts · https://www.alibabacloud.com/help/en/model-studio/models · https://www.marktechpost.com/2026/07/20/alibabas-tongyi-lab-releases-qwen-audio-3-0-tts-a-hosted-text-to-speech-model-in-flash-and-plus-tiers-across-16-languages/
Timeline entry
- Alibaba's Qwen-Audio-3.0-TTS takes #1 on the Artificial Analysis text-to-speech leaderboard ★★★
On 2026-07-20 Alibaba's Tongyi Lab released Qwen-Audio-3.0-TTS in Flash (real-time) and Plus (quality) tiers. It supports 16 languages and 20 Chinese dialect regions. The Plus tier ranked first on the independent Artificial Analysis TTS leaderboard while costing about $27.6 per 1M characters…