Brand Logo

Eleven v4 Explained: Inside ElevenLabs' New Voice Architecture and Benchmarks

Aerin Kim

Written by

Aerin Kim

ElevenLabs shipped Eleven v4 and v4 Turbo on September 28, 2026. Here is the real architecture story, verified benchmark numbers, pricing, and a working API example.

If you built a voiceover workflow, a dubbing pipeline, or a voice agent on top of ElevenLabs any time in the last year, the model underneath it just changed. On September 28, 2026, ElevenLabs shipped Eleven v4 and Eleven v4 Turbo, which the company describes as a genuinely new text to speech architecture rather than an incremental update to Eleven v3 [1]. That is nine days before this post, which matters because Eleven v3 and its Multilingual v2 predecessor have been the default engine behind a huge share of AI voiceover, audiobook narration, and real-time voice agent products since mid-2025. A model swap at that layer is not a cosmetic update. It changes what a cloned voice sounds like, what languages are safe to ship in, how much an audiobook chapter costs to generate, and how fast a voice agent can respond on a live call.

This post goes deep on what actually changed in Eleven v4 and v4 Turbo: the architecture shift ElevenLabs is describing, how the new inline audio tag system actually works syntactically, what the real benchmark numbers say once you pull them from Artificial Analysis directly instead of trusting a single headline claim, what it costs per character compared to the model it replaces, and how to call it from the API with a working code example. Along the way we will also place Eleven v4 in the wider 2026 text to speech landscape, next to Cartesia's Sonic 3.6 and the autoregressive and diffusion approaches that came before it, because "entirely new architecture" means very little without that context.

If you make AI voiceovers, dub video content, or build with Text2Shorts style narration tools, the practical questions in this post (what changed, what it costs, what breaks if you switch models mid-project) are the ones worth answering before you touch a production pipeline.

What Actually Shipped on September 28

ElevenLabs released two models on the same day, aimed at two different jobs [1]:

  • Eleven v4 is the quality-first model, built for audio you produce once and publish: audiobooks, character voiceovers, dubbing, narration, and any content where you can afford a little more generation time in exchange for better delivery and consistency.
  • Eleven v4 Turbo is the low-latency sibling, built for voice agents and other real-time, back-and-forth conversational use, where every extra hundred milliseconds before the first sound reaches the listener is a UX problem, not a quality one.

Both models are available immediately inside ElevenAgents (ElevenLabs' voice agent platform), ElevenCreative (its creative content workspace for voice, music, and video) [11], and directly through the ElevenLabs API, and both are accessible on the free tier, not gated behind an enterprise plan [1] [7].

Headline numbers worth anchoring on before we go deeper, each sourced from ElevenLabs' own documentation and an independent benchmark, not a press release summary:

  • Language support grew from roughly 70 languages in Eleven v3 to more than 90 in Eleven v4, with ElevenLabs' technical docs listing support extending toward 99 languages at varying levels of fluency [2] [3].
  • Voice cloning from as little as 10 seconds of reference audio, for Instant Voice Clones, with Professional Voice Clone support restored after being unavailable in v3 [8].
  • Eleven v4 Turbo posts a median time to first speech of roughly 150 milliseconds in ElevenLabs' own September 2026 benchmark tests, run over WebSocket streaming against Cartesia Sonic 3.6, Google's Gemini Flash and Flash-Lite TTS, and OpenAI's GPT-4o mini TTS, with network latency measured and subtracted out for every system tested [1] [7].
  • On Artificial Analysis' independently run Provider Voice Arena leaderboard, Eleven v4 Turbo and Eleven v4 took the top two spots for September 2026, ahead of Cartesia, Google, and OpenAI's entries [4].

We will unpack every one of those numbers in detail below, including a wrinkle in the benchmark story that ElevenLabs' own marketing copy glosses over.

elevenlabs-eleven-v4-architecture-benchmarks-explained-2026-architecture-comparison.jpg

Step 1: Why ElevenLabs Needed a New Architecture

To understand why "new architecture" is a meaningful claim and not just copywriting, it helps to know what problem Eleven v3 actually had. Eleven v3, launched in alpha in mid-2025, was ElevenLabs' first real push into expressive, tag-driven speech: audio tags like [excited], [whispers], and [sighs] that let a script direct performance, not just pronunciation [3]. It was a genuine leap in expressiveness over the Multilingual v2 model that came before it, but it shipped with two well-documented tradeoffs that production teams ran into constantly: regenerating a single line in a multi-line script could shift the speaker's underlying voice characteristics enough to be noticeable, and ElevenLabs explicitly recommended staying on v2.5 Turbo or Flash for real-time, conversational use because v3 was not built for low-latency streaming [3].

Eleven v4's "entirely new architecture" claim is specifically targeted at those two gaps. ElevenLabs' documentation describes the core change as a new method for capturing and preserving a speaker's identity across a generation, which is what lets v4 stay consistent when you regenerate one sentence in a ten-paragraph audiobook chapter instead of drifting toward a slightly different-sounding voice [2]. On top of that, v4 is described as taking broader textual context into account while generating, rather than treating each sentence as an isolated unit, which is part of why tone and pacing carry forward more naturally across a paragraph instead of resetting at every full stop [1] [8].

ElevenLabs has not published a research paper or technical architecture diagram for Eleven v4, which is consistent with how it has historically kept its production model internals proprietary. That means anyone writing about "the Eleven v4 architecture" in concrete terms (diffusion steps, attention mechanism, token vocabulary) is speculating, and this post will not pretend otherwise. What is useful, and verifiable, is placing Eleven v4 against the three broad architecture families that currently compete for the same market, because the tradeoffs ElevenLabs is visibly optimizing for (quality-first v4 versus latency-first v4 Turbo, as two separate model releases rather than one model with a slider) map directly onto a known architectural tension in TTS research.

Autoregressive neural codec models. This is the lineage that includes Microsoft's VALL-E, which treats text to speech as a language modeling problem: instead of predicting a spectrogram frame by frame, the model predicts discrete audio codec tokens one step at a time, the same way a text LLM predicts the next word token [14]. VALL-E's genuine innovation was showing that a model trained this way could clone an unseen speaker from just a 3-second prompt by treating that prompt as in-context conditioning, the same trick that makes GPT-style models few-shot learners [14]. The tradeoff is that pure autoregressive generation is inherently sequential, generating one token after the previous one resolves, which caps how low your time-to-first-audio can go without architectural tricks layered on top.

Diffusion-based acoustic models. A separate line of work generates speech by iteratively denoising a signal (a spectrogram or a latent audio representation) rather than predicting it token by token, the way image diffusion models generate pixels. A 2023 survey on audio diffusion models for speech synthesis documents this approach across three places it can be applied: the acoustic model that turns text into an intermediate representation, the vocoder that turns that representation into a waveform, or an end-to-end system that does both in one pass [13]. Diffusion-style systems have tended to produce very high perceived naturalness, at the cost of needing multiple denoising passes per generation, which historically made them slower to reach first audio than a well-tuned autoregressive or state-space system, although distilled and single-step variants have narrowed that gap.

State-space models (SSMs). Cartesia, founded by researchers behind the S4 and Mamba state-space-model papers, built its Sonic model family on this third approach. Rather than attending back over every previous token the way a Transformer does, an SSM updates a compressed internal state vector as it processes a sequence, which lets compute scale roughly linearly with sequence length instead of quadratically [12]. That is the architectural reason Cartesia has consistently marketed Sonic on raw speed: Sonic 3.6, the version Eleven v4 was benchmarked against at launch, is built to minimize time-to-first-audio as its primary design goal.

ElevenLabs shipping two separate models instead of one unified model with a quality/speed dial is itself a signal: it strongly suggests Eleven v4 and v4 Turbo are not identical architectures at different settings, but two systems tuned toward opposite ends of the same tradeoff space that autoregressive, diffusion, and state-space approaches all have to navigate in their own way. Whatever ElevenLabs actually built, "new architecture" is doing real work in the sense that going from one general-purpose model (v3) to two purpose-built models (v4 and v4 Turbo) is itself an architectural decision, not just a marketing label.

Step 2: Audio Tags — Directing Performance Instead of Just Text

The single biggest usability change in Eleven v4 for anyone writing scripts by hand is how much more reliably it follows audio tags, the inline bracketed directions that tell the model how to perform a line rather than just what words to say [6]. ElevenLabs' own audio tags documentation lays out the syntax precisely enough to treat as a real spec, not a vague feature description:

  • Placement: a tag goes in square brackets immediately before the text it affects, for example [whispers] Don't let them hear us.
  • Carry-forward scope: once you set a tag, its effect carries forward across the rest of that line until you introduce a new tag, so you do not need to repeat [whispers] on every sentence of a hushed passage, only at the point where the delivery changes [6].
  • Stacking: you can combine multiple qualities in one tag by separating them with commas, for example [whispering, fearful], which is different from writing two separate bracketed tags back to back.
  • Open vocabulary: tags are not limited to a fixed list. ElevenLabs explicitly documents that you can invent your own by describing an emotion, a delivery style, a situation, or a character trait in natural language, and v4 is tuned to interpret that description rather than requiring an exact match against a predefined tag dictionary [6].

The documented tag categories give a sense of how far this goes beyond "happy" and "sad." ElevenLabs groups its example tags into emotion (from high-energy states like excited, amazed, and triumphant, through heated states like aggressive and bitter, to low-energy states like tired, bored, and despair), delivery and volume (whispers, shouts, hushed, booming), pacing (slowly, rushed, drawn out, snappy, long pause), reactions (laughs, sighs, gasps, clears throat, crying), accent and character (British accent, pirate voice), and sound effects that are not vocal at all (thunder rumbling, door creaking, footsteps, zipper opening) [6]. That last category matters for anyone building narrated video content: a single script can carry both the narration and simple ambient sound cues without a separate sound design pass, though for anything beyond a light touch you will still want a real sound effects or music layer underneath it.

Audio tags work across Eleven v4, Eleven v4 Turbo, and the earlier Eleven v3, through the standard text to speech API [6]. What changed in v4 is not the tag syntax itself, which carries over from v3, but how consistently the model actually follows it, particularly for stacked and custom tags, and how well it holds a performance across a long passage instead of the tag's effect fading out after a sentence or two.

elevenlabs-eleven-v4-architecture-benchmarks-explained-2026-audio-tags-pipeline.jpg

Here is a short example script showing stacked and shifting tags in practice, the kind of thing you would hand to Eleven v4 for a two-character dialogue line in a dubbed scene or an audiobook chapter:

[nervous] I don't think we should be here.
[whispering, fearful] Did you hear that?
[long pause]
[shouts] Run!
[footsteps]
[breathless] Okay. Okay, we're clear. [relieved sigh] That was close.

A practical note for anyone adapting an existing v3 script library to v4: audio tags are parsed as part of the text itself, which means they count toward character-based billing the same way spoken words do. A heavily tagged script (lots of short directional tags between sentences) will cost slightly more per minute of output than a lightly tagged one, something worth factoring in if you are converting a large back catalog of narration scripts rather than writing new ones.

Step 3: Voice Cloning From 10 Seconds of Audio

Voice cloning is the other area where v4 is a meaningfully different product from v3, not just a quality bump. ElevenLabs offers two cloning tiers, and both changed with this release [8]:

  • Instant Voice Clones (IVCs) need as little as 10 seconds of clean reference audio and produce a usable clone in roughly the time it takes to upload the file. ElevenLabs describes v4's IVC accuracy as "more accurate than ever," claiming the clone now captures a source voice's timbre, cadence, and delivery more faithfully than any prior model version [2].
  • Professional Voice Clones (PVCs), which train on a larger, higher-quality sample set for a more faithful, lower-artifact result, were not available on Eleven v3 at all. Eleven v4 restores full PVC support, with ElevenLabs stating that professional clones now perform across v4's full emotional and expressive range rather than being limited to flatter, more neutral delivery [8].

The 10-second figure is worth being precise about, because "voice cloning from 10 seconds" is the kind of claim that is easy to either overstate or dismiss. It is a real, documented minimum for a usable Instant Voice Clone, not a marketing exaggeration, but 10 seconds is a floor, not a recommendation. In practice, a clean 10-second sample with no background noise, consistent pacing, and a reasonable range of phonemes will produce a serviceable clone; a noisy or monotone 10-second clip will produce a worse one, and professional work (an audiobook narrator's voice, a brand's signature voiceover) still benefits from the larger Professional Voice Clone pipeline rather than relying on the instant minimum.

elevenlabs-eleven-v4-architecture-benchmarks-explained-2026-voice-cloning-pipeline.jpg

Another documented change is cross-lingual accent handling. ElevenLabs' v4 documentation specifically calls out that when you clone a voice from a reference in one language and then generate speech in a different target language, v4 is built to produce fluent, natural-sounding speech in that target language rather than carrying over the reference speaker's native accent by default [2]. That is directly relevant to dubbing: it means a cloned voice can narrate a video in a second language without sounding like a non-native speaker reading a script phonetically, which was a common artifact in earlier cross-lingual cloning systems. If you specifically want the accent preserved (for stylistic reasons, or because the character is meant to sound foreign), you now have to ask for that explicitly through an accent audio tag rather than getting it as an unavoidable side effect.

Step 4: Eleven v4 Turbo and the Latency Numbers

Latency is where Eleven v4 Turbo's whole reason for existing lives, so it is worth being specific about what was actually measured instead of repeating a single rounded number.

ElevenLabs ran its September 2026 benchmark over WebSocket streaming, using identical scripts and default settings across every system tested, and reporting time to first speech (the delay between sending a request and the first audible sound arriving) with each provider's own network latency measured and subtracted out, so the comparison is about model inference time, not whose servers happen to be closer to the test rig [1] [7]. The reported figures:

  • Eleven v4 Turbo: ~150 milliseconds median time to first speech.
  • Eleven v4 (standard): ~100 milliseconds median inference latency reported separately in ElevenLabs' technical coverage, measuring a related but not identical quantity to Turbo's time-to-first-speech figure, which is why the two numbers should not be read as "v4 is faster than v4 Turbo" [9].
  • Cartesia Sonic 3.6: ~262 milliseconds, in the same test conditions [9].
  • OpenAI GPT-4o mini TTS: ~814 milliseconds, the slowest system in the comparison [9].
  • Google's Gemini Flash-Lite TTS and xAI's TTS offering were both included in the comparison set ElevenLabs benchmarked against, landing somewhere in the 262 to 814 millisecond band along with Cartesia and OpenAI's entries, per ElevenLabs' own announcement [1].
elevenlabs-eleven-v4-architecture-benchmarks-explained-2026-latency-chart.jpg

A few things worth flagging before treating 150ms as a universal number you will see in production. First, this is a median, not a worst case; real-world time to first speech varies with script length, how many audio tags precede the first spoken word, and network conditions between your server and ElevenLabs' infrastructure, which is exactly why ElevenLabs' own test explicitly subtracted network latency rather than including it. Second, Cartesia's own public positioning for Sonic has historically claimed sub-90ms to sub-100ms time-to-first-audio depending on the specific Sonic variant and test setup [12], a figure that does not obviously square with the 262ms number from ElevenLabs' own test. Vendor-run benchmarks of a competitor's product are exactly the kind of number that deserves independent verification rather than being taken at face value from either side, which is the main reason the next section leans on Artificial Analysis' independently run leaderboard instead of either company's self-reported numbers alone.

For anyone building a voice agent where the conversation itself matters more than the model: ElevenLabs is positioning Eleven v4 Turbo specifically for use inside ElevenAgents, where it runs alongside separate transcription and turn-taking models in an integrated stack designed for natural interruption handling, not just a fast text to speech endpoint called in isolation [7]. That distinction matters for anyone comparing raw model latency across providers without accounting for how much of the end-to-end conversational delay actually comes from the surrounding pipeline (speech-to-text, turn detection, the LLM generating a response) rather than the TTS model itself.

Step 5: Where Eleven v4 Actually Ranks on Independent Benchmarks

This is the section where it pays to go past the press release. ElevenLabs' own announcement states that Eleven v4 "ranked first on the Artificial Analysis Provider Voice Arena Leaderboard" for September 2026 [1]. Pulling the actual leaderboard directly from Artificial Analysis clarifies, and slightly complicates, that claim.

Artificial Analysis runs two separate text to speech arenas, and they measure different things [4] [5]:

  • Provider Voice Arena: listeners compare each vendor's own voice catalog against the others, so a provider's result reflects both model quality and the breadth of its voice library.
  • Controlled Voice Arena: every provider is given the same eight cloned reference voices (four US, four UK), so only the underlying model varies, not the voice selection.

On the Provider Voice Arena, the leaderboard Artificial Analysis publishes for this period shows:

RankProvider / ModelElo ScorePrice per 1M Characters
1ElevenLabs Eleven v4 Turbo1334 ± 19$40.00
2ElevenLabs Eleven v41321 ± 18$80.00
4Cartesia Sonic 3.61278 ± 16$49.00
5Google Gemini 3.8 Flash TTS1275 ± 16$16.50
9Google Gemini 3.8 Flash-Lite TTS1242 ± 15$11.00

That table is the real basis for ElevenLabs' "#1" claim, and it holds up, with one nuance: it is Eleven v4 Turbo, not the standard Eleven v4 model, that actually takes the top spot, with standard v4 close behind in second place. ElevenLabs' marketing language treats "Eleven v4" as an umbrella covering both releases, which is accurate in spirit but worth knowing if you are deciding which of the two specific models to benchmark against a competitor yourself.

The picture shifts on the Controlled Voice Arena, where every provider clones the same eight reference voices rather than using its own catalog. There, Alibaba's Qwen-Audio-3.1-TTS-Plus takes the top spot, with Eleven v4 Turbo second and standard Eleven v4 third, still clearly ahead of Cartesia Sonic 3.6 but no longer in first place overall [5]. The practical takeaway: ElevenLabs' "#1" claim is true for the specific leaderboard it is citing, but it is not true across every independent benchmark Artificial Analysis runs, and a model that wins decisively when compared across each provider's best voices does not automatically win when every provider is forced to clone the exact same reference voice. If your use case depends heavily on voice cloning quality specifically, rather than overall catalog quality, the Controlled Voice numbers are the more relevant comparison to look at.

elevenlabs-eleven-v4-architecture-benchmarks-explained-2026-benchmark-chart.jpg

ElevenLabs' separate claim that roughly 75% of listeners preferred Eleven v4 in blind head-to-head tests against five rival services is worth flagging as a different kind of number from the leaderboard Elo scores above: that figure comes from ElevenLabs' own internal testing, not from Artificial Analysis, and preference-test percentages like this are commonly reported by the company whose model is being tested, which does not make the number false, but does mean it should be read as a self-reported result rather than independently audited the way the Artificial Analysis Elo rankings are [1] [7].

Step 6: Calling Eleven v4 From the API

Switching an existing integration from Eleven v3 or Multilingual v2 to Eleven v4 is mostly a one-line change: you swap the model_id parameter in your text to speech request. Here is a minimal, runnable example using ElevenLabs' official Python SDK, hitting the standard text to speech convert endpoint with the v4 model and a stacked audio tag in the script [10]:

python
/code # pip install elevenlabs from elevenlabs import ElevenLabs client = ElevenLabs(api_key="YOUR_XI_API_KEY") audio = client.text_to_speech.convert( voice_id="YOUR_VOICE_ID", # a cloned or library voice ID model_id="eleven_v4", # use "eleven_v4_turbo" for real-time agents output_format="mp3_44100_128", text=( "[nervous] I don't think we should be here. " "[whispering, fearful] Did you hear that? " "[long pause] [shouts] Run!" ), voice_settings={ "stability": 0.6, "similarity_boost": 0.85, "use_speaker_boost": True, }, ) with open("narration.mp3", "wb") as f: for chunk in audio: f.write(chunk)

If you are calling the REST API directly instead of through the SDK (useful for a language the SDK does not cover, or for debugging exactly what gets sent over the wire), the equivalent request body looks like this, posted to https://api.elevenlabs.io/v1/text-to-speech/{voice_id} with your key in the xi-api-key header [10]:

json
/code { "text": "[nervous] I don't think we should be here. [whispering, fearful] Did you hear that? [long pause] [shouts] Run!", "model_id": "eleven_v4", "voice_settings": { "stability": 0.6, "similarity_boost": 0.85, "use_speaker_boost": true }, "output_format": "mp3_44100_128" }

Two parameters are worth understanding before you tune them for production: stability, which controls how consistent delivery is across repeated generations of the same text (lower values allow more expressive variation, higher values lock in a more uniform, repeatable read), and similarity_boost, which controls how closely the output adheres to the reference voice's characteristics when you are using a cloned voice [10]. For narration and dubbing work, a moderately high stability setting with audio tags doing the expressive work tends to hold up better across a long script than relying on low stability to generate variation line by line, since the latter can make a single character sound like two different people across a chapter.

One documentation detail that will trip up anyone porting a v3 pipeline: SSML is not supported on Eleven v4, and the Style and Speed sliders that existed in the ElevenLabs web app for some earlier models are not available for v4 either [2]. If your existing scripts lean on SSML prosody tags for pacing and emphasis, you will need to rewrite that direction as natural-language audio tags instead, which is more work up front but tends to produce more reliably followed results given how much more consistently v4 tracks tags compared to how inconsistently most TTS engines have historically honored SSML.

Pricing: What Eleven v4 Actually Costs

ElevenLabs bills text to speech by credits, where one credit roughly equals one character of output on the Multilingual v2 model, and faster, lower-fidelity models like Flash and Turbo variants have historically run at a discount, around half a credit per character [9]. For Eleven v4 specifically, Artificial Analysis' leaderboard listing prices it at $80.00 per million characters, with Eleven v4 Turbo at $40.00 per million characters, which lines up with the standard-versus-turbo discount pattern ElevenLabs applies across its model lineup [4]. Coverage close to launch also reported a limited-time 72% launch discount running through October 12, 2026, which is worth checking against your own account before assuming the list price applies [9].

PlanMonthly PriceMonthly CreditsApprox. TTS Minutes
Free$010,000~10 minutes
Starter$630,000~30 minutes
Creator$22121,000~2 hours
Pro$99600,000~10 hours
Scale$2991,800,000~30 hours
Business$9906,000,000~100 hours

Both Eleven v4 and Eleven v4 Turbo are included across every subscription tier, including Free, rather than being locked behind Pro or Enterprise, though the free tier's 10,000 monthly credits translate to only a few minutes of generated audio on the standard model, enough to evaluate quality but not to run a real production workload [9]. For context on what that buys: at roughly 1,000 credits per minute of Multilingual v2-equivalent audio, the Creator tier's 121,000 monthly credits is enough for a little over two hours of narration a month before you hit the ceiling, and that math changes meaningfully if your scripts are tag-heavy, since every tag character is billed the same as spoken-word characters.

Where This Actually Gets Used: Dubbing, Agents, and Narration

The real test of any TTS model upgrade is not the benchmark chart, it is whether production workflows that already run at scale on the previous model actually get better. Three use cases line up directly with what changed in v4:

Dubbing and localization. The combination of better cross-lingual accent handling and expanded language coverage (90-plus languages, up from roughly 70) makes v4 a more credible choice for dubbing a video into a second language from a single cloned reference voice, without needing a native speaker to re-record the line [2]. This is the same broad use case Google Ads' own dubbing tooling targets for advertisers running multilingual video campaigns, and ElevenLabs' improvements here are a direct response to the same demand: more content being produced once and localized many times over, rather than reshot per market.

Voice agents. Eleven v4 Turbo's sub-200ms time to first speech, running inside ElevenAgents alongside purpose-built transcription and turn-taking models, is squarely aimed at the same market as Amazon's Nova 2 Sonic and OpenAI's real-time voice offerings: customer support, sales, and scheduling agents where a half-second of dead air after the caller finishes speaking reads as the system being broken, not just a little slow [7].

Narration and audiobooks. The speaker-identity consistency fix is the one that matters most here specifically because audiobook and long-form narration production involves constant re-generation: fixing a mispronunciation, re-recording a line after a script edit, regenerating a paragraph that came out too fast. If every regeneration risks a slightly different-sounding narrator, as teams reported happening with v3, that is a real production cost, not a cosmetic one, and it is the single change in v4 most likely to be invisible in a benchmark chart while mattering enormously to anyone shipping long-form audio [2].

If you are producing short-form video narration rather than long-form audio, the same underlying problem (getting a consistent, well-paced AI voiceover without manually directing every line) is exactly what Miraflow's Text2Shorts workflow is built to simplify: you pick a voice and speed, Miraflow generates the script and scene visuals, and the voiceover renders alongside the video in one pass rather than as a separate audio production step. It is a parallel tool to ElevenLabs rather than something built on top of it, aimed at creators who want a finished short-form video without managing a separate TTS pipeline at all. For a deeper look at voice cloning specifically, including consent and disclosure rules that apply regardless of which model you use, see Miraflow's guide to AI voice cloning for content creators. If your narration needs a music layer to sit under it, Miraflow's AI music generator prompts for ElevenLabs Music v2's genre-switching features covers the companion product in ElevenLabs' own lineup.

Common Mistakes Teams Make Switching to Eleven v4

Porting SSML scripts directly instead of rewriting them as audio tags. As covered above, Eleven v4 does not support SSML [2]. A script full of <break time="500ms"/> and <prosody rate="slow"> tags will either be read aloud literally or silently stripped, neither of which gives you the pacing you wanted. Budget real time to convert prosody direction into natural-language audio tags like [pause] and [slowly] rather than assuming the model will interpret SSML syntax it was never trained to parse.

Assuming "#1 on Artificial Analysis" means #1 on every relevant benchmark. As Step 5 covers, Eleven v4 Turbo tops the Provider Voice Arena but lands third on the Controlled Voice Arena behind Alibaba's Qwen-Audio-3.1-TTS-Plus and its own Turbo variant [5]. If voice cloning fidelity specifically, rather than overall catalog quality, is your priority, check the benchmark that actually isolates that variable before picking a model based on a single headline ranking.

Over-tagging scripts and inflating character billing. Because every audio tag counts as billed characters, a script with a tag before every sentence costs noticeably more per minute than one that uses tags sparingly at genuine delivery shifts, exactly as the tag carry-forward behavior is designed to support. Write tags at the moments where delivery actually changes, not reflexively at the start of every line.

Treating v4 and v4 Turbo as the same model at different quality settings. They are two separate releases with different latency profiles and, per Artificial Analysis, different competitive rankings even against each other [4]. Benchmark your specific use case against both before defaulting to whichever one a blog post (including this one) calls "the best."

Skipping a side-by-side regression test against your existing v3 voice library. A cloned voice that sounds right on v3 is not guaranteed to sound identical on v4, since the underlying identity-capture method changed [2]. If you have an established brand voice already in production, re-clone it and listen side by side before replacing the model wholesale in a live pipeline.

Best Practices for Production Voice Pipelines on Eleven v4

A few concrete practices worth adopting if you are moving a real pipeline onto v4 rather than just testing it in the playground:

  • Separate your quality-critical and latency-critical paths. Use standard Eleven v4 for anything pre-recorded (narration, dubbing, audiobooks) and reserve v4 Turbo specifically for live, conversational paths where time to first speech actually affects the user experience. Running Turbo everywhere for consistency's sake gives up quality you are not actually using the latency for.
  • Lock stability higher for brand-critical voices, and lower it deliberately for character work. A customer-facing brand voice benefits from a higher, more repeatable stability setting; an audiobook with multiple expressive character voices benefits from leaning more on explicit audio tags than on low stability to generate variation, since tags are directable and stability-driven variation is not.
  • Version your scripts, not just your audio. Because audio tags are part of the billed text and part of what drives the model's performance, treat a tagged script the same way you would treat code: track changes, review tag placement the way you would review a diff, and do not silently let a script drift from what shipped originally.
  • Re-run your voice cloning QA step whenever you change model_id. Given the documented change in how v4 captures speaker identity, a clone that passed QA on v3 is not automatically validated on v4. Keep a short reference script you regenerate on every model change specifically to catch drift before it reaches production.
  • Watch the Controlled Voice Arena, not just the Provider Voice Arena, if cloning quality is your bottleneck. The two leaderboards measure genuinely different things, and the gap between them is itself useful signal about where a model's real strength lies.

Frequently Asked Questions

Is Eleven v4 actually a new architecture, or just a retrained version of Eleven v3? ElevenLabs describes it as an entirely new architecture rather than a retraining, specifically citing a new method for capturing and preserving speaker identity and broader use of textual context during generation [2]. The company has not published architecture-level technical details, so the specific mechanism cannot be independently confirmed, but the behavioral changes (fixed speaker drift on regeneration, SSML dropped in favor of audio tags, two separate model releases instead of one) are consistent with a real architectural change rather than a minor fine-tune.

Does Eleven v4 really rank #1 on every voice benchmark? No. It ranks first on Artificial Analysis' Provider Voice Arena for September 2026, where Eleven v4 Turbo specifically takes the top spot [4], but it ranks second and third on the separate Controlled Voice Arena, behind Alibaba's Qwen-Audio-3.1-TTS-Plus [5]. Which leaderboard is more relevant depends on whether you care more about overall voice catalog quality or pure model-level cloning fidelity.

How many languages does Eleven v4 actually support? ElevenLabs' marketing describes "more than 90 languages," up from roughly 70 in Eleven v3, while its technical documentation references coverage extending toward 99 languages at varying levels of fluency, noting explicitly that some are handled more fluently than others [1] [2]. Treat "90-plus" as the reliable headline figure and test your specific target language before shipping it in production.

Can I still use SSML with Eleven v4? No. ElevenLabs' documentation confirms SSML is not supported on v4, along with the Style and Speed sliders available for some earlier models [2]. Pacing and emphasis direction needs to move into natural-language audio tags instead.

What is the real difference between Eleven v4 and Eleven v4 Turbo for my use case? Eleven v4 targets higher quality for content you produce once and publish (audiobooks, dubbing, narration), while Eleven v4 Turbo targets low latency for real-time, conversational use like voice agents, with a median time to first speech around 150 milliseconds in ElevenLabs' own benchmark [1] [7]. If a listener is waiting on the other end of a live call, use Turbo; if you are rendering a file nobody is waiting on in real time, use standard v4.

Does switching to Eleven v4 break my existing cloned voices? Not in the sense of deleting them, but ElevenLabs' documented change in speaker-identity capture means a voice that sounded a certain way on v3 is not guaranteed to sound identical when generated with v4 [2]. Re-test any brand-critical cloned voice side by side before replacing the model in a live pipeline.

Is the 10-second voice cloning claim real, or is that a best-case marketing number? It is a real, documented minimum for ElevenLabs' Instant Voice Clone feature on Eleven v4 [8], not a fabricated statistic, but it is a floor, not a recommendation. Clone quality still scales with how clean and representative your reference sample is, and professional work benefits from the larger Professional Voice Clone pipeline rather than the instant minimum.

Conclusion

Eleven v4 and Eleven v4 Turbo are a genuine model swap under one of the most widely used voice AI products on the market, shipped nine days before this post and already live across ElevenLabs' free and paid tiers. The real, verifiable changes are specific: speaker identity now holds steady across regenerations in a way it visibly did not on v3, audio tags are followed more reliably and can be stacked and customized in natural language, voice cloning works from a 10-second sample with Professional Voice Clones restored, language coverage crossed 90, and Eleven v4 Turbo's roughly 150-millisecond time to first speech puts it ahead of Cartesia Sonic 3.6 and OpenAI's GPT-4o mini TTS in ElevenLabs' own streaming benchmark. The one place marketing outruns the full picture is the benchmark claim itself: Eleven v4 genuinely tops Artificial Analysis' Provider Voice Arena, but a second, independently run Controlled Voice Arena tells a more nuanced story where Alibaba's Qwen-Audio-3.1-TTS-Plus currently leads on pure cloning fidelity. Whichever model you build on, test it against your own scripts and your own cloned voices before trusting any single leaderboard number, including the ones in this post.

References

  1. ElevenLabs: Introducing Eleven v4
  2. ElevenLabs Docs: Eleven v4 (Text to Speech overview)
  3. ElevenLabs: Introducing Eleven v3 (Alpha)
  4. Artificial Analysis: Text to Speech Provider Voice Leaderboard
  5. Artificial Analysis: Text to Speech Controlled Voice Leaderboard
  6. ElevenLabs: ElevenLabs Audio Tags List
  7. ElevenLabs: Eleven v4 Turbo in ElevenAgents
  8. ElevenLabs Help Center: What is Eleven v4?
  9. ElevenLabs Pricing
  10. ElevenLabs API Reference: Text to Speech Convert
  11. ElevenLabs: ElevenCreative
  12. Cartesia: Announcing Sonic, a low-latency voice model for lifelike speech
  13. arXiv: A Survey on Audio Diffusion Models: Text To Speech Synthesis and Enhancement in Generative AI (2303.13336)
  14. arXiv: Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers (VALL-E, 2301.02111)
  15. TestingCatalog: ElevenLabs launches Eleven v4 and v4 Turbo voice models