Post-Cutoff.com
  1. Home
  2. Posts
  3. Qwen

Qwen

Qwen @Alibaba_Qwen · x · 2026-09-23 · ★★★ · archived

Open the original ↗

Cited as a source by: qwen-audio-3-1-realtime

Summary

Archived text

⚡ Meet Qwen-Audio-3.1! ASR, TTS & Realtime are fully upgraded, joined by two new models: TTS-Next for audio creation and ASR-Next for audio understanding.

Five models, one complete audio stack: understanding, generation, interaction & creation.

Plus big price cuts across the lineup: TTS ~70% off, Realtime ~85% off, and ASR up to 95% off.

Highlights: 🥳

  • ASR: stronger multilingual & dialect recognition, plus native polishing that auto-removes fillers & repetitions for cleaner, more logical transcripts.
  • ASR-Next: supports multi-speaker ASR with speaker labels, timestamps & aligned transcripts, and understands emotions, ambient & machine sounds for sound captioning, event localization, audio QA & reasoning.
  • TTS: multilingual & dialect synthesis with natural cross-lingual voice transfer; control emotion, speed & style via simple instructions.
  • TTS-Next: unified LM + diffusion framework generating voice, sound effects & background audio in one pass for audiobooks, podcasts, games & ads.
  • Realtime: speak & listen at once with anytime interruption, just like a real call; it even slows down and responds empathetically when it senses a low mood.

Unlock the full potential of Qwen-Audio-3.1! 👇

Media: https://pbs.twimg.com/media/HS4-pIvbYAAQL0M.jpg?name=orig

views 250132 · likes 4008 · reposts 357 · replies 136 (at fetch time)

Archived 2026-09-29 via fxtwitter (unofficial).

Archived text

⚡ Meet Qwen-Audio-3.1! ASR, TTS & Realtime are fully upgraded, joined by two new models: TTS-Next for audio creation and ASR-Next for audio understanding.

Five models, one complete audio stack: understanding, generation, interaction & creation.

Plus big price cuts across the lineup: TTS ~70% off, Realtime ~85% off, and ASR up to 95% off.

Highlights: 🥳

  • ASR: stronger multilingual & dialect recognition, plus native polishing that auto-removes fillers & repetitions for cleaner, more logical transcripts.
  • ASR-Next: supports multi-speaker ASR with speaker labels, timestamps & aligned transcripts, and understands emotions, ambient & machine sounds for sound captioning, event localization, audio QA & reasoning.
  • TTS: multilingual & dialect synthesis with natural cross-lingual voice transfer; control emotion, speed & style via simple instructions.
  • TTS-Next: unified LM + diffusion framework generating voice, sound effects & background audio in one pass for audiobooks, podcasts, games & ads.
  • Realtime: speak & listen at once with anytime interruption, just like a real call; it even slows down and responds empathetically when it senses a low mood.

Unlock the full potential of Qwen-Audio-3.1! 👇

Media: https://pbs.twimg.com/media/HS4-pIvbYAAQL0M.jpg?name=orig

views 250132 · likes 4008 · reposts 357 · replies 136 (at fetch time)

Archived 2026-09-29 via fxtwitter (unofficial).

All posts · id: x-alibaba_qwen-2102687258990026993