Post-Cutoff.com
  1. Home
  2. Timeline
  3. 2026
  4. ElevenLabs launches Eleven v4 and Eleven v4 Turbo, #1 on…

ElevenLabs launches Eleven v4 and Eleven v4 Turbo, #1 on Artificial Analysis TTS arena

★★★★after cutoffmodel-releaseElevenLabsconfidence: high

On 2026-09-28 ElevenLabs released Eleven v4 (eleven_v4), a text-to-speech model on an entirely new architecture that performs scripts with context-aware emotion, and Eleven v4 Turbo (eleven_v4_turbo, ~100 ms median inference latency) for voice agents. v4 took #1 on the Artificial Analysis TTS arena (Elo ~1315-1319), supports 90+ languages, clones voices from ~10 s of audio and launched with a 72% API discount.

Key facts

What happened

ElevenLabs released Eleven v4 and Eleven v4 Turbo on 2026-09-28. The blog, YouTube launch video (07:01 PT) and X announcement came out the same day. v4 replaces Eleven v3 (June 2025 alpha, GA February 2026) as the flagship. ElevenLabs says it is built on "an entirely new architecture that reads a script the way a voice actor would". Turbo is aimed at ElevenAgents and other live uses. Model files: data/models/elevenlabs-v4.md.

Caveats: the blind-test preference and latency comparisons come from ElevenLabs. The docs still recommend 1-2 minutes of audio for Instant Voice Clones, while the marketing says 10 seconds.

Why it matters

ElevenLabs had fallen behind Cartesia, Google and others on the Artificial Analysis arena with v3 (Elo ~1169). v4 puts it back at #1, and Turbo brings expressive, tag-directed speech to sub-200 ms voice agents at a launch price well below v3.

Changelog

  • 2026-09-29: created

Models

Videos (2)

Introducing Eleven v4 and Eleven v4 Turbo

ElevenLabs · 2026-09-28 · official

Description by Gemini, which watched the video:

Summary This is an official launch video by ElevenLabs introducing its speech foundation models, Eleven v4 and Eleven v4 Turbo. Narrated by a synthetic voiceover against minimalist typographic and particle-based visuals, the video highlights conversational realism, expressive non-verbal vocalizations, voice cloning fidelity, and low-latency multilingual switching.

What is shown

  • [00:00 - 00:08] Opening disclaimer stating that all audio was generated directly from the shown text prompts without edits or modifications using Eleven v4.
  • [00:08 - 00:51] A multi-speaker dramatic dialogue demo set on a film set, demonstrating complex non-verbal audio prompt tags (e.g., [chatter], [nervous], [whispering nervously], [commanding], [clapperboard snap], [voice breaking], [crying], [sniffs], [light chuckle], [British accent]).
  • [00:52 - 01:06] Narration explaining tone, texture, and speaker similarity in professional voice cloning, accompanied by abstract spherical animations.
  • [01:07 - 01:31] A fast-paced Australian radio presenter demonstration navigating prompt annotations including natural pauses, laughter, and tone shifts ([building tension], [chuckle], [laughs], [sarcastic chuckle]).
  • [01:32 - 01:44] Feature overview announcing infinite text duration consistency, support across 100 languages, and the ultra-low-latency model "Eleven v4 Turbo".
  • [01:45 - 02:27] An interactive customer service phone call demo using v4 Turbo where a representative confirms a medication prior authorization and fluently switches from English to Mandarin Chinese ([professionally] 当然可以...).
  • [02:28 - 02:37] ElevenLabs outro branding and title card for Eleven v4.

Claims & numbers

  • The narrator claims everything heard was generated directly from prompts without edits or modifications using Eleven v4.
  • The narrator states the model delivers "significantly better speaker similarity" with professional voice clones.
  • The narrator claims voice consistency "over an infinite text duration."
  • The narrator states the model is native across 100 languages.
  • An ultra-low latency version, Eleven v4 Turbo, is introduced for real-time interactions.

Notable quotes

  • [00:07] "A speech model that doesn't just speak, it performs."
  • [00:59] "With professional voice clones, you don't just imitate a voice, you embody it..."
  • [02:29] "Eleven v4: the next frontier of human-level communication."

Assessment This is an official promotional product announcement showcasing pre-rendered text-to-speech audio outputs generated from detailed prompt annotations. While the audio samples demonstrate impressive emotional inflection and multilingual capabilities, they represent curated showcase demonstrations rather than interactive live interface tests.

Described by gemini-3.8-flash on 2026-09-29 from the video's audio and frames.

Introducing V4 and V4 Turbo for developers

ElevenLabs Developers · 2026-09-28 · official

Description by Gemini, which watched the video:

Summary
ElevenLabs developer advocate Tadas introduces Eleven v4 and Eleven v4 Turbo, the company's next-generation text-to-speech models built on a completely new architecture. He demonstrates their voice cloning fidelity, prompt directing with inline bracket tags, multilingual capabilities, phonetic pronunciation control, developer API integrations (REST, WebSockets, SDKs, CLI, and MCP), and conversational agent performance.

What is shown

  • [00:08] A voice clone of the presenter speaking while the presenter drinks from a mug, trained on 10 minutes of audio.
  • [00:14] Overview of Eleven v4 targeting long-form production, character work, voiceovers, and dubbing, followed by Eleven v4 Turbo at [00:24] for low-latency conversational agents.
  • [00:34] Diagram explaining the new architecture interpreting tone, pacing, emotion, character, and general context.
  • [00:48] Artificial Analysis Text to Speech Leaderboard ranking Eleven v4 at #1 with an Elo of 1319.
  • [01:05] Demonstration of inline performance tags inside square brackets ([whispers], [laughs], [said angrily in British accent], [door slams], [light rain], and [phone buzzing]).
  • [01:37] Multilingual synthesis demonstrated in Polish for a hotel assistant script, followed by phonetic spelling using the International Phonetic Alphabet (IPA) to correctly pronounce the presenter's Lithuanian name "Tadas" at [01:53].
  • [02:09] Request stitching visualization handling requests over 10,000 characters seamlessly.
  • [02:20] API code snippet and live testing showing REST endpoint usage (POST /v1/text-to-speech/{voice_id} with eleven_v4), Python/TypeScript SDK snippets, CLI options, and streaming dialogue over WebSockets with v4 Turbo at [02:44].
  • [03:05] ElevenLabs Model Context Protocol (MCP) server demonstrated inside Claude (using Claude Fable 5.1).
  • [03:19] Walkthrough of the ElevenCreative web platform and the Eleven Agents dashboard showing Eleven v4 Turbo latency metrics (~86 ms to 100 ms median).

Claims & numbers

  • Eleven v4 is ranked #1 on the Artificial Analysis Text to Speech Leaderboard (Provider Voices) with an Elo score of 1319 (ahead of Cartesia Sonic 3.6 at 1276 and Google Gemini 3.8 Flash TTS at 1267).
  • The presenter states that a voice clone can be trained on just 10 minutes of audio.
  • The model supports over 90 languages.
  • A single TTS request can handle up to 10,000 characters, with automated request stitching linking sequential chunks into a single seamless audio file.
  • Eleven v4 Turbo delivers live conversational voice synthesis with a median latency of approximately 100 ms (and as low as ~86 ms in the shown interface).

Notable quotes

  • [00:08] "In fact, for this sentence, I decided to let the model show you. This is a voice trained on 10 minutes of my audio."
  • [00:36] "They're built from the ground up with a brand new architecture."
  • [03:33] "It keeps the full expressive range and responds with a median of 100 milliseconds, which makes live conversation feel more fluid."

Assessment
This is an official product launch and developer walkthrough from ElevenLabs. The presentation features concrete, working audio generations and UI demonstrations across the web platform, REST API, WebSockets, and Claude MCP tool-use, backed by verified benchmarks from Artificial Analysis.

Described by gemini-3.8-flash on 2026-09-29 from the video's audio and frames.

Related events

  1. ElevenLabs releases Music v2.5 ★★★
  2. ElevenLabs raises $500M Series D at $11B valuation (Sequoia) ★★★
  3. ElevenLabs Dubbing v2: direct speech-to-speech dubbing in 90+ languages ★★★
  4. Alibaba's Qwen-Audio-3.0-TTS takes #1 on the Artificial Analysis text-to-speech leaderboard ★★★
  5. Artificial Analysis launches the Speech Agent Arena for speech-to-speech voice agents ★★
  6. BreezeBlue releases Breeze TTS 2, the new top open-weights text-to-speech model ★★★
  7. Cartesia Sonic-3.6 goes GA and tops the Artificial Analysis Speech Arena ★★★
  8. Inworld Realtime TTS-2 reaches GA with audio-aware, prompt-directed speech ★★★
  9. ElevenLabs Studio 4.0 turns ElevenCreative into an agentic AI video editor ★★

Sources (11)

id: 2026-09-28-elevenlabs-eleven-v4 · updated 2026-09-29 · open in the interactive timeline