Post-Cutoff.com
  1. Home
  2. Models
  3. Qwen3-ASR (0.6B / 1.7B) + Qwen3-ForcedAligner

Qwen3-ASR (0.6B / 1.7B) + Qwen3-ForcedAligner

Alibaba (Qwen)currentaudio/speechQwen3-ASRopen weights

Native Transformers (-hf repos) support added 2026-06-26. Hosted ASR is now Qwen-Audio-3.x-ASR (see qwen-audio-3-1-asr).

Input
audio
Output
text
License
apache-2.0
Verified
2026-09-29

How to call it

ProviderModel idEndpoint / URLDocs
Hugging FaceQwen/Qwen3-ASR-1.7Bhuggingface.co/Qwen/Qwen3-ASR-1.7B—
Hugging FaceQwen/Qwen3-ASR-0.6Bhuggingface.co/Qwen/Qwen3-ASR-0.6B—
Hugging FaceQwen/Qwen3-ForcedAligner-0.6Bhuggingface.co/Qwen/Qwen3-ForcedAligner-0.6B—
GitHub—github.com/QwenLM/Qwen3-ASR—

Notable capabilities (2)

Apache-2.0 speech recognition models for self-hosting.

Sources: https://github.com/QwenLM/Qwen3-ASR

Timeline entry

  1. Alibaba open-sources Qwen3-TTS (voice design, 3-second cloning, 97 ms streaming) and, a week later, Qwen3-ASR ★★★

    On 2026-01-22 Alibaba's Qwen team released Qwen3-TTS under Apache-2.0 (0.6B and 1.7B checkpoints plus a 12 Hz tokenizer). It offers voice design from text descriptions, voice cloning from about 3 s of audio in 10 languages, and ~97 ms streaming latency. On 2026-01-29 Qwen3-ASR followed (0.6B/1.7B…

Other Alibaba (Qwen) models

Qwen-Audio-3.1-ASR (Flash) · Qwen-Audio-3.1-Realtime (Plus) · Qwen-Audio-3.1-TTS-Next · Qwen3.8-LiveTranslate (Flash Realtime) · Qwen3.8-Omni-Flash · Qwen3.8-27B · Qwen3.8-Flash · Qwen3.8-Max · Qwen-Image-3.0 (Pro) · Qwen3.7-Plus · Qwen3-TTS (open weights 0.6B / 1.7B; API qwen3-tts-flash)