Post-Cutoff.com
  1. Home
  2. Models
  3. Qwen-Audio-3.0-TTS (Flash / Plus)

Qwen-Audio-3.0-TTS (Flash / Plus)

Alibaba (Qwen / Tongyi Lab)currentaudio/speechQwen-Audio 3.0

Flash tier targets real-time use (~300 ms first packet, press); Plus targets quality (throughput ~16 chars/s, press). Languages: ar, zh, en, fr, de, id, it, ja, ko, ms, pt, ru, es, tl, th, vi. Companion qwen-audio-3.0-realtime-plus/-flash and qwen-audio-3.0-asr-flash also exist. Superseded by Qwen-Audio-3.1 (2026-09-23), but as of 2026-09-29 the Model Studio catalog still lists qwen-audio-3.0-tts-plus as its TTS model, and no 3.1 TTS id is published in the international docs.

Input
text, audio
Output
audio
License
proprietary
Pricing
per 1m characters: $27.59 (USD per 1M characters for the Plus tier at launch (MarkTechPost). Alibaba cut TTS prices about 70% with the 3.1 generation (2026-09-23), so check current pricing) source
Verified
2026-09-29

How to call it

ProviderModel idEndpoint / URLDocs
Alibaba Cloud Model Studio (Singapore / Beijing)qwen-audio-3.0-tts-flashhttps://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api/v1/services/audio/tts/SpeechSynthesizerdocs
Alibaba Cloud Model Studioqwen-audio-3.0-tts-plus—docs

Notable capabilities (2)

Timeline entry

  1. Alibaba's Qwen-Audio-3.0-TTS takes #1 on the Artificial Analysis text-to-speech leaderboard ★★★

    On 2026-07-20 Alibaba's Tongyi Lab released Qwen-Audio-3.0-TTS in Flash (real-time) and Plus (quality) tiers. It supports 16 languages and 20 Chinese dialect regions. The Plus tier ranked first on the independent Artificial Analysis TTS leaderboard while costing about $27.6 per 1M characters…