Term
ElevenLabs v3
ElevenLabs v3 is a highly expressive text-to-speech model with inline audio tags for emotion, emphasis and non-verbal sounds. It supports more than 70 languages and multi-speaker dialogue and is seen as a benchmark for realism.
ElevenLabs v3 — explained in more detail
ElevenLabs v3 is the third generation of the text-to-speech model from the provider ElevenLabs, announced on 3 June 2025 first as an alpha and now generally available. The model targets conversational realism: it mirrors the rhythm and emotional depth of human speech, shifts tone mid-sentence and moves seamlessly between multiple speakers.
The central innovation is audio tags written directly into the text, such as [whispers], [excited], [laughs] or [sighs], along with sound cues. These let users steer emphasis, pacing and non-verbal reactions. A dialogue mode handles natural interruptions and tone shifts across several voices. The model supports more than 70 languages and is offered as a proprietary, closed product via website and API; there are no open weights.
Example / In practice
To voice an audio drama or a commercial with several roles, a script can be enriched with audio tags so that one character whispers, another laughs and the tone fits each scene, without having to combine multiple separate recordings.
Distinction from similar terms
Unlike open models such as Chatterbox or Fish Audio S2, ElevenLabs v3 cannot be self-hosted and ties usage to the provider platform. As a speech and voice model it generates speech from text (TTS), setting it apart from speech recognition (speech-to-text); its distinguishing feature is the focus on expressive, emotion-controlled output.