Eleven v4 and Eleven v4 Turbo
Eleven v4 and Eleven v4 Turbo are speech synthesis models featuring multi-speaker generation, in-script acoustic and emotion direction, voice cloning, and low-latency bidirectional streaming across 90+ languages.
Eleven v4 and Eleven v4 Turbo are speech synthesis models featuring multi-speaker generation, in-script acoustic and emotion direction, voice cloning, and low-latency bidirectional streaming across 90+ languages.
What the product does and how it is positioned
Eleven v4 and Eleven v4 Turbo are synthetic speech models developed by ElevenLabs to generate expressive, multi-speaker audio across more than 90 languages. Built on an architecture that evaluates dialogue context and speaker identity, the model follows in-text tags for emotional inflections and integrated sound effects.
The release includes Eleven v4 Turbo, an optimized variant engineered for real-time applications and conversational agent pipelines. It supports bidirectional streaming, returning synthesized audio while text continues to stream from an upstream language model with approximately 100 milliseconds median inference latency.
Source-supported ways to use the product
Producers can generate complete audiobooks and spoken-word projects using context stitching to maintain uniform pacing and speaker stability across lengthy scripts.
Developers can integrate Eleven v4 Turbo into voice agent loops to achieve low-latency response times with bidirectional streaming.
Creators can draft scripts containing multiple distinct character voices and inline sound effects for games, podcasts, and video productions.
Eleven v4 introduces script-level controls that allow creators to direct emotional tone and insert environmental audio directly into text. Users can place tags such as whispers, laughter, or physical sound effects inline with dialogue, which the model interprets contextually.
To maintain consistency across complex productions, context stitching preserves speaker delivery and pacing over extended scripts, while voice regeneration maintains vocal stability without drift between takes.
Eleven v4 Turbo is engineered for interactive conversational systems, operating with a reported median inference latency of approximately 100 milliseconds and a median time to first speech of approximately 150 milliseconds.
The model supports bidirectional streaming, allowing text generated incrementally by an external large language model to be submitted while audio streams back before the full sentence is finalized.
Checks to run with your own material and workflow
What was checked and when
Answers based on the source-checked product record
Eleven v4 is designed for emotive speech and long-form context stitching across scripts, whereas Eleven v4 Turbo is optimized for real-time applications and conversational agents, offering a median inference latency of approximately 100 ms and time to first speech of approximately 150 ms.
Eleven v4 supports speech synthesis in more than 90 languages, including fluent generation with native accents for languages such as Japanese, Spanish, and Portuguese.
Yes, Eleven v4 supports instant voice cloning from ten seconds of audio as well as Professional Voice Clones, which carry the model's emotional range across supported languages.
Users can insert bracketed direction tags like laughs, whispers, and sound effects directly into the script, while a pronunciation dictionary allows defining custom phonetics and IPA representations for specific words and names.
Yes, Eleven v4 and Eleven v4 Turbo can be accessed via REST API endpoints, streaming endpoints, and official SDKs for Python and TypeScript, switchable using specific model identifiers.