Post-Cutoff.com
  1. Home
  2. Timeline
  3. 2026
  4. Alibaba's Qwen-Audio-3.0-TTS takes #1 on the Artificial…

Alibaba's Qwen-Audio-3.0-TTS takes #1 on the Artificial Analysis text-to-speech leaderboard

★★★after cutoffmodel-releaseAlibabaQwenconfidence: high

On 2026-07-20 Alibaba's Tongyi Lab released Qwen-Audio-3.0-TTS in Flash (real-time) and Plus (quality) tiers. It supports 16 languages and 20 Chinese dialect regions. The Plus tier ranked first on the independent Artificial Analysis TTS leaderboard while costing about $27.6 per 1M characters, roughly a quarter of Eleven v3's price. It was the first Chinese hosted TTS to top that arena.

Key facts

What happened

Alibaba released a new generation of hosted TTS models built on a low-frame-rate tokenizer and a multi-stage training recipe, with strong control features (instructions, inline tags, dialects, long-form output). Its Plus tier topped the Artificial Analysis blind-listening arena at launch.

Why it matters

A Chinese lab led the main independent TTS leaderboard at a fraction of ElevenLabs' price, which started the summer-2026 TTS price and quality race. Alibaba followed two months later with Qwen-Audio-3.1 and price cuts of about 70%.

Changelog

  • 2026-09-29: created

Models

Related events

  1. Alibaba launches Qwen-Audio-3.1 five-model voice stack and cuts audio API prices up to 95% ★★★
  2. ElevenLabs launches Eleven v4 and Eleven v4 Turbo, #1 on Artificial Analysis TTS arena ★★★★
  3. Inworld Realtime TTS-2 reaches GA with audio-aware, prompt-directed speech ★★★
  4. Cartesia Sonic-3.6 goes GA and tops the Artificial Analysis Speech Arena ★★★
  5. Alibaba open-sources Qwen3-TTS (voice design, 3-second cloning, 97 ms streaming) and, a week later, Qwen3-ASR ★★★

Sources (4)

id: 2026-07-20-qwen-audio-3-0-tts · updated 2026-09-29 · open in the interactive timeline