On September 28, 2026, ElevenLabs released two new TTS models: Eleven v4 for content and long scripts, and Eleven v4 Turbo for voice agents. Both understand more than 90 languages, including Russian, mark emotions with inline tags such as [laughs], [whispers], and [pause] instead of disabled SSML, and, according to the company, respond with first speech in approximately 150 ms. The release has already passed its first live test: the author of the Telegram channel "Love. Death. Transformers." integrated Eleven v4 into the Hugging Face Reachy Mini desktop robot, and the robot's antennas began to synchronously repeat the emotions of the synthesized voice.

image
image

What Happened

ElevenLabs released two variants of the new model under different model_id values: eleven_v4 for content and long scripts, available via the Text to Dialogue API, and eleven_v4_turbo for voice agents, operating via WebSocket with bidirectional streaming designed for agentic cycles. The models support more than 90 languages and generate up to 10,000 characters per request, while "context stitching" between fragments is intended to maintain a unified delivery in long audiobooks. Emotional markup has been moved directly into the text in the form of inline tags [laughs], [whispers], [pause] — SSML is disabled in the new models. Support for Professional Voice Clones has been restored, however, old clones now need to be retrained for the new architecture. According to the official ElevenLabs page, v4 Turbo provides a median time to first speech of about 150 ms and a median inference time of about 100 ms — compared to 262 ms for Cartesia Sonic 3.6 and 814 ms for OpenAI GPT-4o mini TTS.

Context

Eleven v3, despite its notable speech quality, had documented weaknesses: dialogues, emotional markup, cloning, and generation length limits. The fourth version addresses exactly these issues — the release should be read not as an improvement of a single feature, but as a response to accumulated market complaints. The broader picture was as follows: in the realtime segment of voice agents, Cartesia had previously set the tone, while solutions at the level of OpenAI GPT-4o mini TTS remained significantly slower; at the same time, TTS as a whole was moving from the role of "text narration" to the role of a building block for agentic products, in which emotion becomes markup that the language model itself can place. The transition from SSML to inline tags is a direct consequence of this logic: the interface between the LLM and speech synthesis shifts from declarative markup to tags embedded directly in the text generated by the model.

Why This Matters for the Industry

For the industry, the main point is the realtime segment: with a claimed median of about 150 ms to first speech, ElevenLabs officially becomes a competitor to Cartesia in the voice agent niche, where speed was the defining advantage of a single player. The platform also covers long-form content (Text to Dialogue, audio tags, voice cloning) and real-time agents — this is a claim to consolidate the voice layer: for startups, entry into voice products becomes cheaper and faster, while differentiation increasingly shifts to emotions, dialogue management, and API ergonomics, rather than speech synthesis itself. At the same time, migration becomes a manageable obligation for the coming weeks: you need to change the model_id in the API and retrain old Professional Voice Clones — those who do not migrate will remain on Multilingual v2 without emotional markup and realtime latency. If the trajectory from v3 to v4 is maintained, TTS risks becoming a commodity component of the agentic stack, where competition revolves around the combination of "latency plus controllable emotion."

Why This Matters for Users

For readers, the release provides several practical things at once. Russian voice can be generated among the 90+ supported languages, and the promotional rate with tripled credits on Creator+ plans is valid until October 12, 2026 — this window is enough to launch a pilot and measurements for almost free. Those who listen to audiobooks and podcasts will get a more stable delivery of long recordings: generation proceeds in large fragments that are "stitched" by context, without a choppy tempo between chapters. For builders of personal bots and robots, the trick from the demo is particularly illustrative: the author of the post integrated Eleven v4 into the Hugging Face Reachy Mini desktop robot and programmed its antennas to repeat the voice's emotions — the code for this was written by Claude Opus 5.5, which parses v4's emotional markup and synchronizes movements with intonations in real time. The meaning of the trick is broader than a toy: the "emotion" of a voice agent stops being just sound and can control external reactions of a device — a potential channel for embodied interfaces around synthesized speech.

What Is Still Unknown / Limitations

All numerical characteristics — about 150 ms, about 100 ms, and the phrase "the best TTS in the world" — are ElevenLabs' self-report without a published methodology and independent reproductions. Latency is given only as medians: the provided sources do not include p95/p99 percentiles, descriptions of hardware, network, and agentic cycle load, and realtime systems usually break down precisely on the tails of the distribution — therefore, it is too early to judge that the bar for human-like voice agents has already been passed. The latency comparison used a weak baseline of OpenAI GPT-4o mini TTS with 814 ms, which contrasts favorably with the vendor's figures, while the real competitor Cartesia Sonic 3.6 lags by only about 110 ms — a classic sign of selective selection of comparison points. The abandonment of SSML in favor of inline tags remains a design decision without a public assessment: no one has measured how stably the model interprets the tags, especially in languages other than English. Scenarios such as "LLM itself places emotional tags as standard" and "inline tags will displace SSML" are interpretations not confirmed by sources; the same applies to the transformation of TTS into an interface layer between an agent and its "body."

Sources

Author

Look at AI, editorial team