Term
Chatterbox
Chatterbox is an open-source text-to-speech model from Resemble AI under the MIT license. In blind tests most listeners preferred it over ElevenLabs; it offers emotion control, voice cloning and low latency.
Chatterbox — explained in more detail
Chatterbox is an open-source text-to-speech model from the provider Resemble AI, released in 2025 under the permissive MIT license. This allows the weights and code to be used, adapted and self-hosted freely. In blind tests a majority of listeners preferred Chatterbox over the commercial ElevenLabs, praising naturalness, clarity and emotional range.
Its features include emotion control, voice cloning from a few seconds of reference audio, multilingual output and low latency in the range of a few hundred milliseconds. Because the model is openly available, it suits operation on your own infrastructure without vendor lock-in; it reached wide adoption on GitHub and Hugging Face.
Example / In practice
A team that wants to run speech output in its own data center in a privacy-compliant way can set up Chatterbox locally, clone a company voice from a short recording and use it to voice announcements or product videos without sending text to an external service.
Distinction from similar terms
Unlike the closed ElevenLabs v3, Chatterbox is freely licensed as an open-weight model and can be self-hosted. Like other speech and voice models it generates speech from text (TTS) and is distinct from speech recognition; its profile centers on open availability while remaining competitive in quality.