Kyutai TTS 1.6B / Kyutai STT + Unmute
STT open-sourced 2025-06-19, TTS + Unmute open-sourced 2025-07-03 (Kyutai blog). Weights CC-BY-4.0. For CPU TTS see kyutai-pocket-tts.
- Input
- text, audio
- Output
- audio, text
- License
- cc-by-4.0
- Verified
- 2026-09-29
How to call it
| Provider | Model id | Endpoint / URL | Docs |
|---|---|---|---|
| Hugging Face (TTS) | kyutai/tts-1.6b-en_fr | huggingface.co/kyutai/tts-1.6b-en_fr | — |
| Hugging Face (STT) | kyutai/stt-2.6b-en | huggingface.co/kyutai/stt-2.6b-en | — |
| Hugging Face (STT) | kyutai/stt-1b-en_fr | huggingface.co/kyutai/stt-1b-en_fr | — |
| GitHub (Unmute) | — | github.com/kyutai-labs/unmute | — |
Notable capabilities (2)
- Text-streaming TTS: Delayed-streams architecture (~1.8B params incl. 600M depth transformer) starts speaking before the full text is available, English + French; voices only via pre-computed embeddings (no raw cloning, by design). source
- Streaming STT with semantic VAD: stt-2.6b-en (English, 2.5 s delay) and stt-1b-en_fr (0.5 s delay) transcribe as audio arrives; used in Unmute, which wraps any text LLM with real-time STT+TTS. source
Sources: https://kyutai.org/blog/ , https://huggingface.co/kyutai/tts-1.6b-en_fr , https://huggingface.co/kyutai/stt-2.6b-en
Other Kyutai models
Kyutai Pocket TTS · Kyutai Moshi / Hibiki-Zero (full-duplex speech models)