A voice agent can have the right answer and still sound wrong. A delayed reply breaks the rhythm of a call; a flat one can make an apology sound like a menu prompt. ElevenLabs is addressing both problems with two models announced on September 28, 2026: Eleven v4 for expressive speech production and Eleven v4 Turbo for faster, streaming conversations.
The release restores Professional Voice Clone support, which was absent from Eleven v3, and expands both models to more than 90 languages. Creators will want to know whether v4 can keep a voice and its performance consistent across a long project. Developers need to find out whether Turbo’s reported speed holds up inside a complete voice-agent system.
Eleven v4 Puts More Direction Into the Script
According to ElevenLabs’ launch announcement, v4 uses a new architecture designed to interpret tone, pacing, character and conversational context. The company says it handles exchanges between speakers more naturally, allowing a character’s line to reflect what another speaker has just said.
Creators can direct delivery with inline tags and natural-language instructions. A script can ask for a laugh, a whispered phrase or a change in emotional tone; tags can also specify sounds such as rain or a phone buzzing. ElevenLabs introduced expression tags in v3. For v4, it claims more reliable handling of those directions, including sequences of tags.
In production, an awkward line can mean regenerating an entire scene. ElevenLabs says v4 better preserves speaker identity across dialogue, regenerated lines and longer passages, while improved request stitching helps join generations for audiobooks and other long-form work. Those are company claims, not a guarantee that a full book will sound like one uninterrupted take.
The controls require some adjustment. The Eleven v4 documentation lists Stability and Similarity settings for both variants, but says the Style and Speed sliders are unavailable and SSML is unsupported. Teams that rely on those older controls should test how well tags, punctuation and the remaining settings reproduce their existing direction. ElevenLabs acknowledges that tag following is not perfect.
Professional Voice Clones Return With a Migration Catch
Professional Voice Clones, or PVCs, are supported in Eleven v4 after being unavailable in v3. Both v4 variants also support Instant Voice Clones. ElevenLabs says an Instant Voice Clone can be made from as little as 10 seconds of audio. That short-sample claim should not be confused with the training material or fidelity expected from a Professional Voice Clone.
For a publisher with an established narrator, or a studio maintaining character voices across episodes, the key change is compatibility with v4’s expressive controls. ElevenLabs says its updated approach captures speaker identity more faithfully. A more faithful clone, however, will not necessarily sound like the output a team approved in v3.
ElevenLabs warns in its documentation that v4 can sound substantially different from v3 and recommends comparing both models with a team’s own content. Reference recording quality matters: noise, distortion and inconsistent audio can become audible characteristics of the generated voice. Its current guidance favors clean training audio in a single speaking style, although the company says that advice may change as it learns more about the model.
Existing clone owners should treat v4 as a new voice approval, not a routine model-ID update. Compare the same passages in v3 and v4, including quiet speech, emphatic lines, names and difficult pronunciations. If a clone carries unwanted recording artifacts or has lost the intended character, revisit the source audio before deciding whether to retrain it. A different result is not automatically a worse one, but it should not reach a finished audiobook, ad or agent without review.
Turbo Is Built for the Agent Loop
Eleven v4 Turbo is intended for interactive speech. ElevenLabs says developers can stream text into it while a large language model (LLM) is still generating an answer, then receive audio before the full sentence is complete. Turbo is available through the API and ElevenAgents, the company’s conversational-agent platform.
A voice agent must receive speech, interpret the request, generate a response and speak it. Waiting for the LLM to finish a full written answer before starting text-to-speech adds an audible pause. Streaming allows speech generation to overlap with the LLM’s output, although developers still need to manage interruptions, incomplete responses and the rest of the conversation loop.
ElevenLabs positions Turbo as carrying v4’s expressive range into those faster interactions. An agent might need to sound different when confirming a booking than when handling a complaint, without changing its recognizable voice. That is a useful design goal for customer support and games. The test is an actual conversation: whether the voice remains intelligible and appropriate as the agent responds, pauses or changes direction.
The 100 ms and 150 ms Claims Measure Different Things
ElevenLabs reports roughly 100 milliseconds of median inference latency for Turbo and roughly 150 milliseconds of median time to first speech. These figures measure different things. The latter measures the interval from a request to audible speech in the company’s test; neither number describes the time from a caller finishing a question to hearing a complete agent reply.
The launch announcement’s benchmark notes say ElevenLabs measured time to first speech in September 2026 using Turbo over WebSocket streaming, identical scripts and default settings for the compared systems. Network latency was measured and removed. The result is useful as a vendor-run speech-model comparison, though a live call also crosses a real network and passes through speech recognition, an LLM and application logic. A median says nothing about unusually slow turns.
ElevenLabs makes a separate quality claim: it says v4 was preferred by about 75% of listeners in blind, head-to-head tests against four named competing models. Its description says listeners compared the same line and judged expressiveness and naturalness, with ties counted as half. That does not establish a 75% preference over every speech model, every voice or every language. The company also cites a September 2026 Artificial Analysis voice-arena ranking. Neither figure replaces testing with a product’s own scripts and users.

Voice-agent teams should measure end-to-end response time alongside speech quality. A fast first sound is less valuable if the model mishandles a customer’s name, starts speaking before the answer is settled or delivers lines inconsistently across turns.
Ninety-Plus Languages Change the Accent Calculation
Both variants support more than 90 languages, up from the 70 supported by v3 as TechCrunch reported. ElevenLabs says v4 also improves rhythm, emotion and accent handling when a voice speaks in a language different from its reference recording.
Multilingual teams need to hear one specific behavior before migrating. When the generated language differs from the reference voice’s language, the company’s documentation says v4 aims for a natural accent in the target language rather than preserving the source-language accent. When the languages match, it says the original accent is retained.
That choice could help dubbing or multilingual agents sound more local while keeping a recognizable speaker identity. It could be wrong for a character whose cross-language accent is part of the performance. ElevenLabs describes the behavior as a deliberate change and says an optional toggle remains a research project, with no announced timeline. Teams that need an accent to carry across languages should test that requirement directly.
Language count is only a starting point for localization. Proper names, regional vocabulary, pronunciation and emotional delivery still need review by people who understand the target audience. More supported languages broaden where a model can be tried; they do not certify every output for release.
How to Evaluate the Upgrade
Eleven v4 and v4 Turbo are available in ElevenCreative, ElevenAgents and through ElevenLabs’ API, according to the company. The documented model IDs are eleven_v4 and eleven_v4_turbo, allowing API users to test a different variant by changing the model selection. That technical switch is simpler than approving the resulting audio.
For existing projects, a measured rollout would use a small set of representative material: a multi-speaker scene, a long narration excerpt, regenerated lines and any languages the product actually serves. Compare those outputs with the approved v3 versions, then check voice identity, pronunciation, tag behavior and the source-audio issues that v4 may expose. Teams using Professional Voice Clones should include their real production clones rather than relying on a library demo.
Agent developers need a different test set. Run Turbo through the complete application, including the LLM and any speech-recognition step. Record when users first hear a response, how long a useful answer takes, and what happens during interruptions and longer sessions. ElevenLabs’ 150 ms figure is a benchmark for one part of that experience, not a deployment target guaranteed by switching models.
Final Thoughts
Creators no longer have to choose between v3’s expressive direction and Professional Voice Clone support. For projects built around a particular narrator or character, v4 offers a reason to revisit that workflow, provided the new output passes a fresh voice review.
Turbo’s latency claim is promising but narrower. ElevenLabs has reported how quickly its speech model can start talking under its test conditions. Developers still have to establish how quickly, and how well, their agents can answer.
Frequently Asked Questions
4 questions
1What Is the Difference Between Eleven v4 and v4 Turbo?
Eleven v4 is the variant ElevenLabs recommends when speech quality is the priority, such as audiobooks and character voiceovers. Eleven v4 Turbo is designed for real-time applications, including voice agents, and supports streaming text input and audio output. Both variants support more than 90 languages and work with Instant and Professional Voice Clones.
2How Fast Is Eleven v4 Turbo?
ElevenLabs reports roughly 100 ms median inference latency and roughly 150 ms median time to first speech in its vendor-run tests. The first figure concerns model inference; the second concerns how soon audio becomes audible after a request. Neither includes the full delay of a live voice agent, which also depends on networking, speech recognition and response generation.
3Does Eleven v4 Support Existing Professional Voice Clones?
Eleven v4 supports Professional Voice Clones, a capability that was unavailable in v3. Owners of existing clones should compare v4 output with approved v3 audio before switching production work. ElevenLabs says v4 may reproduce a reference voice differently, including unwanted qualities in noisy or distorted source recordings, so some voices may benefit from better training audio.
4Will Eleven v4 Preserve a Cloned Voice’s Accent in Another Language?
Not necessarily. ElevenLabs says v4 aims to use a natural accent for the target language when that language differs from the voice’s reference recording, while preserving the original accent when the languages match. This may suit dubbing, but teams whose characters must retain a source-language accent across languages should test v4 before migrating.
Sources
- ElevenLabs’ launch announcementelevenlabs.io
- Eleven v4 documentationelevenlabs.io
- ElevenAgentselevenlabs.io
- TechCrunch reportedtechcrunch.com





