Introducing Eleven v4 and Eleven v4 Turbo
ElevenLabs · 2026-09-28 · official · 69,683 views
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary This is an official launch video by ElevenLabs introducing its speech foundation models, Eleven v4 and Eleven v4 Turbo. Narrated by a synthetic voiceover against minimalist typographic and particle-based visuals, the video highlights conversational realism, expressive non-verbal vocalizations, voice cloning fidelity, and low-latency multilingual switching.
What is shown
- [00:00 - 00:08] Opening disclaimer stating that all audio was generated directly from the shown text prompts without edits or modifications using Eleven v4.
- [00:08 - 00:51] A multi-speaker dramatic dialogue demo set on a film set, demonstrating complex non-verbal audio prompt tags (e.g.,
[chatter],[nervous],[whispering nervously],[commanding],[clapperboard snap],[voice breaking],[crying],[sniffs],[light chuckle],[British accent]). - [00:52 - 01:06] Narration explaining tone, texture, and speaker similarity in professional voice cloning, accompanied by abstract spherical animations.
- [01:07 - 01:31] A fast-paced Australian radio presenter demonstration navigating prompt annotations including natural pauses, laughter, and tone shifts (
[building tension],[chuckle],[laughs],[sarcastic chuckle]). - [01:32 - 01:44] Feature overview announcing infinite text duration consistency, support across 100 languages, and the ultra-low-latency model "Eleven v4 Turbo".
- [01:45 - 02:27] An interactive customer service phone call demo using v4 Turbo where a representative confirms a medication prior authorization and fluently switches from English to Mandarin Chinese (
[professionally] 当然可以...). - [02:28 - 02:37] ElevenLabs outro branding and title card for Eleven v4.
Claims & numbers
- The narrator claims everything heard was generated directly from prompts without edits or modifications using Eleven v4.
- The narrator states the model delivers "significantly better speaker similarity" with professional voice clones.
- The narrator claims voice consistency "over an infinite text duration."
- The narrator states the model is native across 100 languages.
- An ultra-low latency version, Eleven v4 Turbo, is introduced for real-time interactions.
Notable quotes
- [00:07] "A speech model that doesn't just speak, it performs."
- [00:59] "With professional voice clones, you don't just imitate a voice, you embody it..."
- [02:29] "Eleven v4: the next frontier of human-level communication."
Assessment This is an official promotional product announcement showcasing pre-rendered text-to-speech audio outputs generated from detailed prompt annotations. While the audio samples demonstrate impressive emotional inflection and multilingual capabilities, they represent curated showcase demonstrations rather than interactive live interface tests.
Described by gemini-3.8-flash on 2026-09-29 from the video's audio and frames.