Term
Kokoro
Kokoro is an open text-to-speech model (v1.0, January 2025) with just 82M parameters. It produces natural-sounding speech in 8 languages with 54 voices, runs on around 1 GB of VRAM and ships under the Apache-2.0 license.
Kokoro — explained in more detail
Kokoro (Kokoro-82M) is a text-to-speech (TTS) model whose version 1.0 was released on January 27, 2025 under the permissive Apache-2.0 license. What stands out most is its small size: with only 82 million parameters it is a fraction of many competing speech-synthesis models, yet it delivers comparable quality — at significantly higher speed and lower cost. The weights occupy roughly 1 GB of VRAM, so Kokoro runs even on modest hardware or in the browser.
The v1.0 release ships 54 voices across 8 languages. The model was deliberately trained only on permissive, non-copyrighted audio data with IPA phoneme labels — which is what makes the permissive licensing possible in the first place. Kokoro is therefore the TTS counterpart to speech-to-text models like Whisper: instead of turning audio into text, it turns text into audible speech.
Example / In practice
Kokoro is well suited for voiceovers, screen readers, in-app speech output or narrating articles and podcasts — locally and without API costs. Because it is small and Apache-licensed, it can be embedded into your own products without being tied to a cloud provider.
Distinction from similar terms
Kokoro generates speech (TTS), Whisper transcribes it (STT) — both belong to the Speech & Voice use case. Compared to commercial services such as ElevenLabs, Kokoro wins on open weights and self-hosting, while those offer broader voice variety and voice cloning.