
What is Eleven v4?

Eleven v4 is the fourth generation of ElevenLabs' speech model, from one of the best-known AI voice companies, and the company calls it "the first in a new generation of models" on its Eleven v4 docs page.
The launch post says it runs on an entirely new architecture, one that reads a script like a voice actor would, paying attention to who's speaking and what just happened before deciding how a line should land.
I build AI agents at eesel (the AI teammates platform), and the team here has spent years putting AI on live support queues, so when a launch like this lands I mostly want to know one thing: what changes for someone who has to ship it into production? Short answer, v4 is a real step up, but it's a migration too, not a drop-in swap.
Per the models reference, there are two variants:
- Eleven v4 (
eleven_v4): the quality model. 10,000 character limit (about 10 minutes of audio), multi-speaker dialogue, served through the Text to Dialogue API. - Eleven v4 Turbo (
eleven_v4_turbo): the real-time model, with a median inference latency of ~100ms, served through the Text to Dialogue WebSocket and inside ElevenAgents.
You can use both today in ElevenCreative (the web app) and ElevenAgents, plus the API, according to the September 28 changelog. ElevenLabs also put out a short developer walkthrough, if video is more your thing.

What actually changed from v3?
Expressiveness and cloning accuracy get the headlines. The details you actually plan around sit in the docs, though, so I pulled this side-by-side from the models page and the API pricing page.
| Eleven v3 | Eleven v4 | Eleven v4 Turbo | |
|---|---|---|---|
| Best for | Dramatic performances | Audiobooks, voiceovers, dubbing | Voice agents, interactive apps |
| Languages | 70+ | 90+ | 90+ |
| Character limit per request | 5,000 | 10,000 | Streaming |
| Latency | Not real-time | Not real-time | ~100ms median inference |
| API list price per 1K characters | $0.08 | $0.08 | $0.04 |
| Professional Voice Clones | Not fully optimized | Supported | Supported |
| Style and Speed sliders | Yes | No | No |
| SSML | No | No | No |
| Arena Elo (Artificial Analysis) | 1174 | 1321 | 1334 |
Some of these need more room than a table cell gives them.
Professional Voice Clones are back. According to the v3 best-practices page, PVCs were not fully optimized for v3, which is why teams with an expensive professional clone often stayed on older models. The best practices guide now points those voices straight at v4. There's a caveat, though. The docs say PVC support "is currently rolling out to everyone," so a few accounts may still be waiting.
Instant clones need less audio. The launch post says Instant Voice Clones capture a voice from "just 10 seconds of audio," though the docs still recommend one to two minutes if you want a solid result.
Long-form is steadier. ElevenLabs says request stitching (chaining generations into a longer piece) is "significantly more reliable," and audiobook producers working in Studio are the ones who'll notice that most.
Only two sliders remain. Stability and Similarity stay in v4. Style and Speed don't, so direction now lives in the script itself, as audio tags like [whispers] or [sighs].
Why Turbo is the one most people should pick

This is the part that surprised me. You'd assume the "quality" model wins on quality over the "fast" one. On the Artificial Analysis Provider Voice Arena, where listeners pick the more natural of two clips without knowing which model made them, it's the other way round:
- Eleven v4 Turbo: 1334 Elo across 1,645 votes, $40 per 1M characters.
- Eleven v4: 1321 Elo across 1,930 votes, $80 per 1M characters.
- Qwen-Audio-3.1-TTS-Plus: 1292 Elo, $19.30 per 1M characters.
- Cartesia Sonic 3.6: 1278 Elo, $49 per 1M characters.
- Gemini 3.8 Flash TTS: 1275 Elo, $16.50 per 1M characters.
The two v4 models overlap within their confidence ranges (both sit in the "1-2" rank band), so the fair way to read it is that Turbo matches the full model on short-clip naturalness at half the price. Where the full model earns its price is on things the arena doesn't test, like the 10,000 character limit and keeping scene-level context across a long script. Stitching a whole chapter without drift is another one.

The gap over v3 holds up too, and it isn't marketing rounding. v3 sits at 1174 on the same board, 147 points behind v4, and on an Elo scale that means listeners prefer v4 most of the time. ElevenLabs' own testing claims v4 was preferred by ~75% of listeners against Cartesia Sonic 3.6, Inworld TTS-2, and two Gemini 3.8 TTS models, per the footnotes in the launch post. I'd treat that one as a vendor claim and lean on the arena numbers as the independent check.
My own rule of thumb is Turbo by default, and the full model only when a single request has to carry minutes of continuous narration.
Comparing vendors? My Cartesia Sonic 3 vs ElevenLabs and Gemini 3.8 Flash TTS review cover the next two names on that board.
The Gemini 3.8 Flash TTS pricing post explains why Google undercuts everyone on cost. There's a wider shortlist in ElevenLabs alternatives.
Which Eleven model fits your project?
Pick what you're building. Prices are API list rates per 1,000 characters.
How native accent handling changes multilingual voices
The docs single this out as one of the biggest changes in v4. It's also the one most likely to catch an existing customer off guard.
A cloned voice on v3 took its accent with it into other languages. v4 splits that into two cases. If the output language matches the voice's own language, the accent is preserved; if it doesn't, v4 produces fluent speech with a native accent for the target language instead. The docs give this example: clone a Korean speaker, generate English, and you get natural English, not English with a Korean accent.

That's a big upgrade for dubbing, and for agents serving callers in several languages, since one brand voice can now sound local everywhere. If your voice is the accent, though, it's a regression. ElevenLabs is upfront that a toggle is "still a research project" with no timeline.
Not every early user agrees it works as described. On the official subreddit thread, one commenter hit the opposite problem:
"In English it's good but the American accent bleeds terribly into other languages."
Someone else wanted accent control set at the voice level instead of per line. For anyone running a long series, that's a fair ask:
"I want accent/dialects to be encoded at the voice level, not have to prompt for it in every single line. It is really impossible to keep a stable idiolect that way and generic [Scottish] or whatever can still come out a million ways."
So, practically: if you run multilingual customer support through a voice agent, generate a test script in each of your top languages before you switch. According to the turbo-in-agents post, the biggest gains landed in Japanese, Brazilian Portuguese, Mandarin, and Cantonese, which makes them good ones to try first.
Will my cloned voices sound the same?

Probably not, and ElevenLabs says as much. v4 reproduces the source voice more accurately (timbre, cadence, loudness, even recording flaws), so a v4 clone "may sound quite different from its Eleven v3 version," per the v4 FAQ. There's also a line in the docs I'd want every audio lead to read twice: accuracy and personal preference are not always the same.
That has two knock-on effects:
- Your training audio matters more. v4 faithfully copies plosives and harsh sibilance, and volume jumps too. The showcase clips on the docs page are left raw on purpose to show this. If a noisy recording is where your clone came from, re-record it before blaming the model.
- Library voices can shift under you. Users on ElevenReader (the listening app) noticed voices changing mid-book. Some loved it and some didn't.
"In ElevenReader it genuinely blew me away. I didn't even know the update was coming, but suddenly my narrator sounded like an actual real human was reading my book to me."
Creators who had tuned a voice for a specific, restrained delivery saw it the other way:
"I'm actually having big issues with v4 its super expressive but for my niche its too damn theatrical and you cannot make certain voices be straight/monotone while speaking..."
There's a quieter asterisk for Voice Design voices (the ones generated from a text description). The docs say they work on v4 but "may not be as performative or sound as good as with earlier models."

How do audio tags and prompting work in v4?
Since the Style and Speed sliders are gone, direction has to come from the script. You write tags inline, like [whispers], [sighs], [laughs harder], or sound cues like [door slams] and [light rain], and v4 follows them "more reliably than v3," per the v4 landing page. There's also an ElevenLabs audio tags list with 50+ examples.
Some notes from the prompting guide that can save you a re-generation or two:
- No SSML. Neither v4 nor v3 supports SSML break tags, so that's a change if you're coming from an older speech API. Use ellipses and line breaks instead, or tags like
[short pause]or[long pause]. - Tags can be misread as sound effects. v4 generates sound effects as well as voices, so a vague tag can come out as a noise. The docs suggest something more descriptive, such as
[low, gravelly voice]. - Tags work best inside the voice's range. A voice that never whispered in its training data can still follow
[whispering], but "the result may not be optimal." - IPA pronunciation is native now. You wrap IPA symbols in forward slashes right in the text, and no XML phoneme tags are needed. That's handy for drug names, product names, or place names a support voice agent has to get right.
Of all the complaints I saw, the missing Speed slider came up most. One creator wrote that v4 "does seem to be speaking much faster (there isn't a slider in the voice settings to control that)" in the same Reddit thread, and the subreddit moderator passed it along as feedback. Until that changes, pacing is something you handle in the writing, with shorter sentences and explicit pause tags.
How much does Eleven v4 cost?

At list price on the API, Eleven v4 costs the same as v3, and Turbo matches v3 Conversational. For now there's also a 72% launch discount until October 12, 2026, per the API pricing page.
| Model | List price per 1K characters | Launch price (until Oct 12) | Latency | Character limit |
|---|---|---|---|---|
| Eleven v4 | $0.08 | $0.022 | Not real-time | 10,000 |
| Eleven v4 Turbo | $0.04 | $0.011 | ~100ms | Streaming |
| Eleven v3 | $0.08 | n/a | Not real-time | 5,000 |
| Eleven v3 Conversational | $0.04 | n/a | ~280ms | Streaming |
| Flash v2.5 | $0.04 | n/a | ~75ms | 40,000 |
In the web app, v4 draws from the same monthly credit pool as everything else. Current plans on the ElevenCreative pricing page are Free ($0, 10,000 credits), Starter ($6, 30,000), Creator ($22, 121,000), Pro ($99, 600,000), Scale ($299, 1.8M), and Business ($990, 6M). The pricing FAQ puts Text to Speech at roughly 1 credit per character. There's also a launch promo for Creator plans and above where v4 usage up to 2x your monthly TTS credits doesn't count against your balance in the web and mobile apps. Oddly, the banner and the promo section on that page disagree on whether it's 2x or 3x, so I'd check your own balance before planning around it. My ElevenLabs pricing guide walks through every tier, and the ElevenLabs reviews post covers what paying users say about the credit system.
A worked example. A Reddit user asked what an 8-hour audiobook would cost, so here's the math. The docs peg 10,000 characters at about 10 minutes of audio, which puts 8 hours at roughly 480,000 characters:
- Eleven v4 on the API at list: 480 × $0.08 = $38.40
- Eleven v4 at the launch price: 480 × $0.022 = $10.56
- Eleven v4 Turbo at list: 480 × $0.04 = $19.20
- In the app: about 480,000 credits, which fits inside the Pro plan's 600,000 at $99/month.
Voice agents aren't billed per character. ElevenAgents pricing works per call minute instead, with plans from Free (15 minutes) up to Business ($990 for 12,375 minutes), extra minutes at $0.08 each, and the LLM charged separately. I unpacked the same math in the best AI voice agents roundup.
Should you put Eleven v4 Turbo on your support line?

ElevenLabs is clearly aiming Turbo at customer service. The ElevenAgents launch post talks about agents that sound "apologetic about a missed delivery, brisk when someone just wants a confirmation," and Salesforce's Agentforce Voice team is quoted on the v4 page praising Turbo's "faster, more natural responses."
It's landing in a crowded month, too: OpenAI's GPT-Live-1, xAI's Grok Voice Think Fast 2, and Decagon Voice 3 all target the same support calls.
Salesforce has its own AI voice agent story too. For callers, that's a real improvement, since a monotone bot reading a refund policy is part of why people mash zero to reach a human. What a TTS launch can't fix is the answer itself. The voice reads whatever the agent produced, and a beautifully apologetic voice confirming the wrong refund window is still wrong, only now it sounds more convincing.
Years of watching AI on real support queues point to the same lesson. Accuracy comes from what the agent knows (your help center, macros, and past tickets) and from testing it before customers ever do, which is why every eesel rollout gets simulated against historical tickets first. If you're shortlisting voice platforms, my guides to AI call center agents, Retell AI, and Vapi compare how each handles knowledge and handoff, plus helpdesk integrations.
My voice support roundup lines them all up side by side.
Voice is also a smaller slice of support than it feels like. Most teams still get the bulk of their volume through tickets and chat, so that's where answer quality problems show up first. If you're already on Zendesk, its voice AI agents and an AI customer support layer for the ticket queue can share the same knowledge.
Try eesel for the answers behind the voice
eesel is an AI teammate platform, and the AI helpdesk teammate is the one built for this exact problem. It joins your existing Zendesk, Freshdesk, or Gorgias queue and learns from your past tickets and help center, then drafts or sends replies in your team's tone. Before it touches a live customer, you can run it in simulation over your historical tickets to see what it would have said.
To be honest about scope, eesel is text-first today: it works tickets and chat, not phone audio. One support team using eesel told me voice-recording support "would significantly increase our usage," which is why it's on the radar. If you're pairing a voice platform with Eleven v4 Turbo, eesel covers the ticket side of the same knowledge, and the eesel CLI (more in this AI agent CLI guide) lets a script or a coding agent like Claude Code update that knowledge and test answers from the terminal, so both channels stay in sync. Plans start free with 100 credits on the pricing page, and every ticket or chat handled is one credit.
Try eesel and see how it handles last month's tickets before you decide anything.
Frequently Asked Questions
What is Eleven v4?
eleven_v4 for content and eleven_v4_turbo for real-time use, covers 90+ languages, and adds back Professional Voice Clone support that v3 lacked.How much does Eleven v4 cost?
Is Eleven v4 better than Eleven v3?
What is the difference between Eleven v4 and Eleven v4 Turbo?
Can I use Eleven v4 for a customer support voice agent?
Does Eleven v4 keep a cloned voice's accent in other languages?
Does Eleven v4 support SSML?
What are the best Eleven v4 alternatives?

Article by
Kira
Kira is a writer at eesel AI with a Computer Science background and over a year of hands-on experience evaluating AI-powered customer service tools. She focuses on breaking down how helpdesk platforms and AI agents actually work so that support teams can make better buying decisions.








