Eleven v4 vs v3: Side-by-Side Test with Audio

Eleven v4 vs v3, measured: same scripts and voices, 42 generations. v4 rendered 1.6× faster with the same accuracy. Hear every pair, then compare both free.

Last updated: 2026-10-05

Short answer: use Eleven v4. We sent the same 7 scripts to both models with the same voices, 3 times each. v4 generated the audio 1.6× as fast as v3, read the scripts just as accurately, and performed every audio tag without speaking it, as v3 did. It also takes twice as much text per request (10,000 characters against 5,000), adds 11 languages, Cantonese among them, and costs the same.

Compare them on your own text. Generate once with each model and both results stay on screen side by side:

Try Eleven v4 now: free, no sign-up

2 free generations a day, up to 200 characters each. Tags in [brackets] are performed, not read aloud.

ModelGenerate once with each to compare them side by side.
Voice
Tags98 / 200
Examples
2 free left today

Eleven v4 vs v3 at a glance

Eleven v3Eleven v4
Released2025September 28, 2026
Model IDeleven_v3eleven_v4
Languages (ElevenLabs)70+90+
Language codes in the API model list7485
Max characters per request5,00010,000
Audio tagsYesYes, stackable
Low-latency siblingv3 Conversationalv4 Turbo (~100 ms median inference, per ElevenLabs)
PriceStandard TTS rateSame as v3
Measured: seconds to generate 1 s of audio (median)0.370.23
Measured: time to first audio byte (median)0.81 s0.70 s
Measured: transcription error vs script (mean)1.9%1.8%
Measured: tags read aloud0 of 21 runs0 of 21 runs

How we tested

  • When: October 5, 2026, through the public ElevenLabs API.
  • Scripts: 7 short pieces, 8 to 19 seconds each when spoken: English narration, English with four audio tags, Mandarin narration, Mandarin with three tags, Japanese, Brazilian Portuguese, and a two-character Mandarin scene sent to text-to-dialogue.
  • Fairness: each pair used the same voice, the same text and default settings. Every script ran 3 times per model, with the order alternating (v3 first, then v4 first) so a slow minute on the network could not favor either model. That makes 42 generations and 531 seconds of audio.
  • Speed: timed on the streaming endpoint from request to first audio byte and to last byte, from our network in Asia. The network round trip is included, so compare the ratio, not the absolute seconds.
  • Accuracy: every output was transcribed with ElevenLabs Scribe v2 and compared with the script with tags removed: character error rate for Chinese and Japanese, word error rate otherwise.

Results by script

Median of 3 runs. Lower is better in every column.

Scriptv3: s per s of audiov4: s per s of audiov3 errorv4 error
English narration0.400.230%0%
English with tags0.550.253.8%3.8%
Mandarin narration0.370.201.2%0%
Mandarin with tags0.330.214.3%4.3%
Japanese0.390.240%0%
Brazilian Portuguese0.330.260%0%
Two-character scene (Mandarin)0.350.193.9%4.4%

What the error column is picking up:

  • The tagged and dialogue scripts score 4–5% on both models because a performed laugh or sigh gets written down by the transcriber as a word ("哎", "哈哈"). That is the tag working, not a misread.
  • English with tags loses one word on both models: Scribe hears "heard" as "hurt".
  • The only genuine difference is in Mandarin narration: in two of v3's three runs, 眼镜 (glasses) came back as 眼睛 (eyes). v4 got it right all three times.

Listen side by side

Run 1 of each pair, unedited. v3 first, then v4.

English narration

Voice: British ButlerEnglishModel: eleven_v314 s
Transcript:The house had been empty for years, yet every clock in it still told the right time. Margaret noticed it the moment she stepped inside: the soft, patient ticking, as if someone had been winding them every night, waiting for her to come home.
Voice: British ButlerEnglishModel: eleven_v415 s
Transcript:The house had been empty for years, yet every clock in it still told the right time. Margaret noticed it the moment she stepped inside: the soft, patient ticking, as if someone had been winding them every night, waiting for her to come home.

English with audio tags

In our runs the transcriber picked up the laugh and the sigh as sound events in all three v4 takes and in two of the three v3 takes.

Voice: Sweet GirlEnglishModel: eleven_v310 s
Transcript:[whispers] Okay, don't move. I think it's still in the kitchen. [nervous] What if it heard us? [laughs] Relax, it's just the cat. [sighs] You scared me half to death.
Voice: Sweet GirlEnglishModel: eleven_v411 s
Transcript:[whispers] Okay, don't move. I think it's still in the kitchen. [nervous] What if it heard us? [laughs] Relax, it's just the cat. [sighs] You scared me half to death.

Mandarin narration

Voice: News AnchorChineseModel: eleven_v314 s
Transcript:天还没亮,菜市场已经亮起了灯。卖豆腐的老陈把第一板豆腐抬上案子,热气一下子冒起来,模糊了他的眼镜。这么多年,他每天都是第一个到。
Voice: News AnchorChineseModel: eleven_v414 s
Transcript:天还没亮,菜市场已经亮起了灯。卖豆腐的老陈把第一板豆腐抬上案子,热气一下子冒起来,模糊了他的眼镜。这么多年,他每天都是第一个到。

Mandarin with audio tags

Voice: Ice QueenChineseModel: eleven_v313 s
Transcript:[laughs] 你以为我会就这么算了?[whispers] 合同在我手里,三天之内,你会求着来找我。[sighs] 可惜啊,我们本来可以是朋友。
Voice: Ice QueenChineseModel: eleven_v412 s
Transcript:[laughs] 你以为我会就这么算了?[whispers] 合同在我手里,三天之内,你会求着来找我。[sighs] 可惜啊,我们本来可以是朋友。

Japanese

Voice: Anime GirlJapaneseModel: eleven_v39 s
Transcript:おはようございます!今日は朝から雨ですが、新しい傘を買ったので、ちょっとだけ楽しみです。駅まで一緒に歩きませんか?
Voice: Anime GirlJapaneseModel: eleven_v48 s
Transcript:おはようございます!今日は朝から雨ですが、新しい傘を買ったので、ちょっとだけ楽しみです。駅まで一緒に歩きませんか?

Brazilian Portuguese

Voice: ThiagoPortugueseModel: eleven_v310 s
Transcript:Bom dia! Hoje vamos aprender a fazer um café coado perfeito, sem pressa e sem segredo. Primeiro, aqueça a água sem deixar ferver.
Voice: ThiagoPortugueseModel: eleven_v410 s
Transcript:Bom dia! Hoje vamos aprender a fazer um café coado perfeito, sem pressa e sem segredo. Primeiro, aqueça a água sem deixar ferver.

Two characters in one request

Voice: Old Storyteller / Monkey KingChineseModel: eleven_v319 s
Transcript:[curious] 你这猴头,又从哪里闯了祸回来? [laughs] 师父莫急,俺老孙不过是借了他三根毫毛! [sighs] 借?只怕人家此刻正满山找你呢。 [excited] 找便找,俺一个筋斗,早到了十万八千里外!
Voice: Old Storyteller / Monkey KingChineseModel: eleven_v417 s
Transcript:[curious] 你这猴头,又从哪里闯了祸回来? [laughs] 师父莫急,俺老孙不过是借了他三根毫毛! [sighs] 借?只怕人家此刻正满山找你呢。 [excited] 找便找,俺一个筋斗,早到了十万八千里外!

What changed, and what didn't

  • Speed is the clearest win. v4 was faster on all 7 scripts, by 1.3× to 2.2×. On short clips, most of the wait is the model finishing rather than starting: time to first byte improved less, from 0.81 s to 0.70 s.
  • Accuracy was already high on v3 and stays there. Both models read every script in full; neither skipped, repeated or invented words.
  • Tags carry over unchanged. Every tag we used on v3 worked on v4, and neither model read one aloud. v4 adds stacking and sequencing, per ElevenLabs.
  • More room per request. 10,000 characters means a typical chapter in one call instead of two, with fewer seams.
  • More languages. v4's API model list adds Cantonese, Burmese, Mongolian, Uzbek, Yoruba, Maltese, Māori, Odia, Tajik, Asturian and Occitan. v3 has no language that v4 lacks.

One caveat from production: we rendered the 76 episodes of our audio drama theater on v4. In long Chinese text-to-dialogue requests of about 1,700 characters or more, v4 occasionally dropped the last few lines without an error. Requests under 1,200 characters for Chinese, Japanese and Korean, plus a check that every line has timestamps, solved it. We did not run v3 at that length.

Which one should you use?

  • New voiceover, audiobook or drama: Eleven v4.
  • Live voice agent: Eleven v4 Turbo, the low-latency sibling.
  • A project already half-finished on v3: finish it on v3 if new lines must match takes you have approved, then move on.
  • Lowest cost per minute: Google's Gemini 3.8 TTS costs about $0.017 a minute. On the same scripts it was 2.4× slower than v4 and added filler words in two takes.

On VoiceSmiths you can pick either model in the voice studio at the same price, and the audio drama studio renders on v4. More about the model itself is on our Eleven v4 page.

FAQ

Is Eleven v4 better than Eleven v3? For new work, yes. In our test v4 generated audio 1.6 times as fast as v3 with the same accuracy, and it doubles the per-request limit to 10,000 characters and adds 11 languages, at the same price.

Do v3 audio tags work in Eleven v4? Yes. [whispers], [laughs], [sighs], [excited], [nervous] and [curious] worked on both models in our test, and neither model read a tag aloud in 42 generations.

Does Eleven v4 cost more than v3? No. ElevenLabs bills v4 at the same credit rate as its other TTS models, and VoiceSmiths charges the same for v4 as for v3.

Is Eleven v4 more accurate than v3? About the same. Scored against the script with speech-to-text, v4 averaged 1.8% and v3 1.9%; nearly all of the gap is laughs and sighs being written down as words. The one real slip was v3 in Mandarin, where 眼镜 (glasses) was heard as 眼睛 (eyes) in two of three runs.

Should I switch my v3 projects to v4? Start new projects on v4. For a project half-finished on v3, keep v3 until it is done if you need re-renders to match takes you have already approved.

Eleven v3, Eleven v4 and ElevenLabs are trademarks of ElevenLabs. VoiceSmiths is an independent studio built on the ElevenLabs API and is not affiliated with ElevenLabs.