ElevenLabs v4 clones a voice from 10 seconds of audio, in 90+ languages
Eleven v4 and the faster Eleven v4 Turbo support more than 90 languages and take inline tags like [laughs] to direct the delivery.
At a glance
- What
- Eleven v4 and Eleven v4 Turbo, from ElevenLabs
- Languages
- More than 90
- Voice cloning
- From 10 seconds of audio
- Direction
- Inline tags in the script, like [laughs]
- Speed
- v4 about 100 ms; v4 Turbo about 150 ms to first speech
- Where
- ElevenAgents, ElevenCreative, ElevenAPI
- Released
- September 28, 2026
Why it matters
Voiceover for ads and explainers gets easier to localize: one script, dozens of languages, the same voice.
ElevenLabs has released Eleven v4, its new text-to-speech model, and Eleven v4 Turbo, a low-latency version for voice agents. Both support more than 90 languages, and Instant Voice Clones now need only 10 seconds of audio.
You direct the read with inline tags in the script, such as [laughs] or a described tone, and can add sound cues the same way. ElevenLabs puts v4’s median inference latency at about 100 ms, and v4 Turbo’s median time to first speech at about 150 ms. Both are in ElevenAgents, ElevenCreative and the API.
What to check
- If you localize voiceover, test one script across your languages with a single cloned voice.
Questions
How many languages does ElevenLabs v4 support?
More than 90, in both Eleven v4 and Eleven v4 Turbo.
How much audio does a voice clone need?
ElevenLabs says Instant Voice Clones need 10 seconds of audio.
Sources
ElevenLabsSeptember 28, 2026
Eleven v4: Our most expressive text-to-speech AI model yetBacks up: Languages, voice cloning, inline tags, latency and where it’s available
Written in our own words from these sources. AI News is independent: the companies we cover don’t sponsor, review or endorse it. Spotted a mistake? See our editorial standards or email zev@esy.com.
Follow ElevenLabs v4 by email
The week’s AI news in The Marketing Engineer
