Fish Audio S2 Pro / S2.1 Pro
S2 Pro held #1 open-weights on Artificial Analysis until Breeze TTS 2 (Aug 2026); now #2 open (~1119 Elo). S2.1 Pro weights are NOT open. OpenRouter lists S2.1 Pro release as 2026-07-29 (API availability there). Predecessor OpenAudio S1 (`s1`) still supported.
- Input
- text, audio
- Output
- audio
- License
- Fish Audio Research License (S2 Pro weights; non-commercial, commercial license on request)
- Pricing
- per 1m utf8 bytes: $15 (USD per 1M UTF-8 bytes (s1, s2-pro, s2.1-pro); s2.1-pro-free $0 under fair use through 2026-11-30; ASR transcribe-1 $0.36/hr) source
- Verified
- 2026-09-29
How to call it
| Provider | Model id | Endpoint / URL | Docs |
|---|---|---|---|
| Fish Audio API | s2.1-pro | https://api.fish.audio/v1/tts (model passed in `model` header) | docs |
| Fish Audio API (free tier) | s2.1-pro-free | https://api.fish.audio/v1/tts | — |
| Fish Audio API | s2-pro | — | — |
| Hugging Face | fishaudio/s2-pro | huggingface.co/fishaudio/s2-pro | — |
| OpenRouter | — | openrouter.ai/fish-audio/s2.1-pro | — |
| GitHub | — | github.com/fishaudio/fish-speech | — |
Notable capabilities (3)
- Inline natural-language emotion/paralinguistic tags: Free-form bracket cues like [whisper], [laugh], [emphasis]; multi-speaker dialogue in one pass; 80+ languages from 10M+ hours of training audio. source
- Open model with production inference stack: Dual-AR (4B slow + 400M fast) on a Qwen3-4B backbone released with fine-tuning code and SGLang serving; RTF 0.195, ~100 ms TTFA; Seed-TTS Eval WER 0.54% zh / 0.99% en. source
- Free production API (S2.1 Pro) (found after launch): S2.1 Pro (closed, 2026-06-23) offered free under fair use with ~90 ms TTFA, 83 languages; 61% win rate vs S2 Pro. source
Sources: https://docs.fish.audio/developer-guide/models-pricing/pricing-and-rate-limits , https://huggingface.co/fishaudio/s2-pro , https://fish.audio/blog/s2-1-pro-free-api/
Timeline entry
- Fish Audio open-sources S2: expressive 80+ language TTS with inline emotion tags ★★★
Fish Audio released S2 (S2 Pro) on 2026-03-09 with weights, fine-tuning code and an SGLang-based production inference stack: a Dual-AR TTS on a Qwen3-4B backbone trained on 10M+ hours in ~80 languages, with free-form [bracket] emotion and paralinguistic cues and multi-speaker dialogue. It led…