Introducing V4 and V4 Turbo for developers
ElevenLabs Developers · 2026-09-28 · official · 12,707 views
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary
ElevenLabs developer advocate Tadas introduces Eleven v4 and Eleven v4 Turbo, the company's next-generation text-to-speech models built on a completely new architecture. He demonstrates their voice cloning fidelity, prompt directing with inline bracket tags, multilingual capabilities, phonetic pronunciation control, developer API integrations (REST, WebSockets, SDKs, CLI, and MCP), and conversational agent performance.
What is shown
- [00:08] A voice clone of the presenter speaking while the presenter drinks from a mug, trained on 10 minutes of audio.
- [00:14] Overview of Eleven v4 targeting long-form production, character work, voiceovers, and dubbing, followed by Eleven v4 Turbo at [00:24] for low-latency conversational agents.
- [00:34] Diagram explaining the new architecture interpreting tone, pacing, emotion, character, and general context.
- [00:48] Artificial Analysis Text to Speech Leaderboard ranking Eleven v4 at #1 with an Elo of 1319.
- [01:05] Demonstration of inline performance tags inside square brackets (
[whispers],[laughs],[said angrily in British accent],[door slams],[light rain], and[phone buzzing]). - [01:37] Multilingual synthesis demonstrated in Polish for a hotel assistant script, followed by phonetic spelling using the International Phonetic Alphabet (IPA) to correctly pronounce the presenter's Lithuanian name "Tadas" at [01:53].
- [02:09] Request stitching visualization handling requests over 10,000 characters seamlessly.
- [02:20] API code snippet and live testing showing REST endpoint usage (
POST /v1/text-to-speech/{voice_id}witheleven_v4), Python/TypeScript SDK snippets, CLI options, and streaming dialogue over WebSockets with v4 Turbo at [02:44]. - [03:05] ElevenLabs Model Context Protocol (MCP) server demonstrated inside Claude (using Claude Fable 5.1).
- [03:19] Walkthrough of the ElevenCreative web platform and the Eleven Agents dashboard showing Eleven v4 Turbo latency metrics (~86 ms to 100 ms median).
Claims & numbers
- Eleven v4 is ranked #1 on the Artificial Analysis Text to Speech Leaderboard (Provider Voices) with an Elo score of 1319 (ahead of Cartesia Sonic 3.6 at 1276 and Google Gemini 3.8 Flash TTS at 1267).
- The presenter states that a voice clone can be trained on just 10 minutes of audio.
- The model supports over 90 languages.
- A single TTS request can handle up to 10,000 characters, with automated request stitching linking sequential chunks into a single seamless audio file.
- Eleven v4 Turbo delivers live conversational voice synthesis with a median latency of approximately 100 ms (and as low as ~86 ms in the shown interface).
Notable quotes
- [00:08] "In fact, for this sentence, I decided to let the model show you. This is a voice trained on 10 minutes of my audio."
- [00:36] "They're built from the ground up with a brand new architecture."
- [03:33] "It keeps the full expressive range and responds with a median of 100 milliseconds, which makes live conversation feel more fluid."
Assessment
This is an official product launch and developer walkthrough from ElevenLabs. The presentation features concrete, working audio generations and UI demonstrations across the web platform, REST API, WebSockets, and Claude MCP tool-use, backed by verified benchmarks from Artificial Analysis.
Described by gemini-3.8-flash on 2026-09-29 from the video's audio and frames.