
A good AI avatar video can still fall flat on one thing: a voice that reads the script instead of performing it. The lip sync is right, the framing is right, but the delivery sounds the same in the first line and the last.
That is the gap ElevenLabs set out to close with Eleven v4, released on September 28, 2026. AI Studios now runs on it. Below is what the model changes, what it means for videos you make in AI Studios, and how to write scripts that get the most out of it.
What is ElevenLabs Eleven v4?
Eleven v4 is ElevenLabs' newest text-to-speech model. Its focus is performance: reading tone, pacing, emotion, and character from the script, not just pronouncing the words.
It ships in two versions. Eleven v4 is the full-featured model for produced content. Eleven v4 Turbo shares the same architecture but is tuned for speed, aimed at real-time voice agents.
A few numbers ElevenLabs published at launch:
- Ranked first on the Artificial Analysis text-to-speech leaderboard (September 2026)
- Preferred by about 75% of listeners in blind head-to-head tests against competing models
- Around 150 ms median time to first audible speech on v4 Turbo
One practical change for script writers: v4 moves away from SSML-style markup. Direction is written inline as plain-language tags, such as [laughs], [whispers], or [light rain].
Five things that got better in v4
1. A wider range of emotion tags
v4 understands more directions and follows them more reliably than v3, especially when several tags appear in a row. You can go beyond basic emotions to delivery cues like [said angrily in a French accent] or sound cues like [door slams].
2. 90+ languages
Language support now covers more than 90 languages. The same voice can speak each one with a native-sounding accent, and rhythm and emotion carry over instead of flattening in translation.
3. The same speaker sounds like the same speaker
Earlier models could drift: a voice that sounded slightly different from one paragraph, or one video, to the next. v4 keeps a speaker's voice stable without distortion, including in multi-speaker dialogue.
4. Long scripts stay connected
v4 reads a script as a whole, not as isolated sentences. Context stitching keeps tone and pacing consistent across long-form audio, so a 10-minute training module doesn't sound like 40 separate clips joined together.
5. 150 ms to first speech
On v4 Turbo, median time to first audible speech is about 150 ms, with roughly 100 ms of inference latency. ElevenLabs reports competitors at 262 to 814 ms. That gap is what makes natural back-and-forth possible in conversational use.
What changes in AI Studios
Two things change for AI Studios users.
Emotion tags now render on v4. Until now, scripts with emotion tags in AI Studios were voiced with Eleven v3. They now run on v4, so you get the wider tag range and more reliable tag-following without changing how you write.
Voices stay consistent. When the same avatar and voice appear across a series, whether that's a course, a product walkthrough, or weekly updates, the voice holds steady from scene to scene and video to video. Viewers hear one presenter, not a close approximation of one.
How to write a script with emotion tags
Tags work best as direction to an actor, not decoration. Place them right before the line they should affect, and use them where the delivery should change.
A plain script:
Welcome back. Last week we covered onboarding. Today we're looking at the one mistake almost every new team makes.
The same script with direction:
[warmly] Welcome back. Last week we covered onboarding.
[pauses] [lowers voice] Today we're looking at the one mistake almost every new team makes.
A few habits that help:
- Tag the change, not every line. If the whole script is warm, one tag at the start is enough.
- Be specific. [sighs, relieved] gives the model more to work with than [happy].
- Keep scripts whole. Because v4 reads context across the script, pasting the full text gives better pacing than splitting it into short scenes.
- Listen once, adjust once. Preview the first scene before generating the full video.
Where the difference shows
FAQ
What is ElevenLabs v4?
Eleven v4 is ElevenLabs' latest text-to-speech model, released September 28, 2026. It focuses on expressive delivery, supports 90+ languages, and comes with a low-latency Turbo version.
Does AI Studios use ElevenLabs v4?
Yes. Scripts that use emotion tags in AI Studios are now voiced with Eleven v4 instead of v3.
Do I need to change my existing scripts?
No. Tags you already use still work. v4 follows them more reliably and supports a wider range.
What are emotion tags?
Short directions in square brackets, like [laughs] or [whispers], placed in the script to shape how a line is delivered.
How many languages does Eleven v4 support?
More than 90, with the same voice able to speak each one in a native-sounding accent.
Try it on your next script
The easiest way to hear the difference is to take a script you've already produced, add two or three tags where the delivery should shift, and generate it again. Open AI Studios and start with your next video.
Sources
- Eleven v4: Our most expressive text-to-speech AI model yet (ElevenLabs blog)
- Eleven v4 product page (ElevenLabs)
- ElevenLabs Launches Eleven V4 With Low-Latency Turbo Variant (Unite.AI)

