Artificial Analysis
Artificial Analysis @ArtificialAnlys · x · 2026-07-08 · ★★★ · archived
Cited as a source by: 2026-08-24-artificial-analysis-speech-agent-arena
Summary
Archived text
Announcing the Controlled Voice Arena Leaderboard comparing Text to Speech models on the same set of 8 cloned voices
The Controlled Voice Arena standardizes, through voice cloning, the set of voices that each model’s performance is evaluated on - separating specific voice preference from broader aspects of model quality. It complements our Provider Voice Arena, where each model uses a select set of its own available voices.
We have generated speech samples on models that offer voice cloning abilities using the same voice categories as our existing Provider Voice Arena, namely: 2 US Male voices, 2 US Female voices, 2 UK Male voices, 2 UK Female voices. Each model has been cloned on the same 1-2 minute audio recordings for each voice.
Key results ➤ Overall: @cartesia Sonic 3.5 leads (1,122 Elo), followed by @ElevenLabs Eleven v3 (1,088) and @inworld Realtime TTS-2 - Research Preview (1,070) ➤ US accent: Cartesia Sonic 3.5 leads (1,139 Elo), followed by ElevenLabs Eleven v3 (1,104) and Inworld Realtime TTS-2 - Research Preview (1,059) ➤ UK accent: Cartesia Sonic 3.5 also leads (1,103 Elo), with Inworld Realtime TTS-2 - Research Preview (1,075) moving ahead of ElevenLabs Eleven v3 (1,067) into 2nd ➤ Open weights: @FishAudio S2 Pro leads (1,034 Elo), followed by @MistralAI Voxtral TTS (1,024) and @resembleai Chatterbox (930)
See more details below ⬇️
Media: https://pbs.twimg.com/media/HMt6U1PagAAj7QS.png?name=orig
views 26034 · likes 155 · reposts 18 · replies 5 (at fetch time)
Archived 2026-09-29 via fxtwitter (unofficial).
Archived text
Announcing the Controlled Voice Arena Leaderboard comparing Text to Speech models on the same set of 8 cloned voices
The Controlled Voice Arena standardizes, through voice cloning, the set of voices that each model’s performance is evaluated on - separating specific voice preference from broader aspects of model quality. It complements our Provider Voice Arena, where each model uses a select set of its own available voices.
We have generated speech samples on models that offer voice cloning abilities using the same voice categories as our existing Provider Voice Arena, namely: 2 US Male voices, 2 US Female voices, 2 UK Male voices, 2 UK Female voices. Each model has been cloned on the same 1-2 minute audio recordings for each voice.
Key results ➤ Overall: @cartesia Sonic 3.5 leads (1,122 Elo), followed by @ElevenLabs Eleven v3 (1,088) and @inworld Realtime TTS-2 - Research Preview (1,070) ➤ US accent: Cartesia Sonic 3.5 leads (1,139 Elo), followed by ElevenLabs Eleven v3 (1,104) and Inworld Realtime TTS-2 - Research Preview (1,059) ➤ UK accent: Cartesia Sonic 3.5 also leads (1,103 Elo), with Inworld Realtime TTS-2 - Research Preview (1,075) moving ahead of ElevenLabs Eleven v3 (1,067) into 2nd ➤ Open weights: @FishAudio S2 Pro leads (1,034 Elo), followed by @MistralAI Voxtral TTS (1,024) and @resembleai Chatterbox (930)
See more details below ⬇️
Media: https://pbs.twimg.com/media/HMt6U1PagAAj7QS.png?name=orig
views 26034 · likes 155 · reposts 18 · replies 5 (at fetch time)
Archived 2026-09-29 via fxtwitter (unofficial).
Related events
All posts · id: x-artificialanlys-2074886571166462405