Speech & Voice

in AI Models

Voice AI in both directions — speech recognition/transcription (STT) and speech synthesis (TTS). A cross-vendor category, e.g. Whisper, Kokoro, and ElevenLabs.

Blog Posts

Glossary

AssemblyAI Universal-2 Speech & Voice

Universal-2 is AssemblyAIs closed speech-to-text model with industry-leading accuracy (mean WER around 6 percent) and strong formatting, punctuation and proper-noun recognition across 99 languages.

ElevenLabs v3 Speech & Voice

ElevenLabs v3 is a highly expressive text-to-speech model with inline audio tags for emotion, emphasis and non-verbal sounds. It supports more than 70 languages and multi-speaker dialogue and is seen as a benchmark for realism.

Chatterbox Speech & Voice Open-Weight

Chatterbox is an open-source text-to-speech model from Resemble AI under the MIT license. In blind tests most listeners preferred it over ElevenLabs; it offers emotion control, voice cloning and low latency.

Deepgram Nova-3 Speech & Voice

Nova-3 is Deepgrams real-time speech-to-text model (released February 2025) with very low streaming latency of around 200 to 300 milliseconds, keyterm prompting and support for more than 30 languages.

Fish Audio S2 Speech & Voice Open-Weight

Fish Audio S2 is an open-source text-to-speech model from the Fish Speech family. With inline tags it controls emotion and emphasis at word level, supports more than 80 languages and delivers very low latency.

Kokoro Speech & Voice Open-Weight

Kokoro is an open text-to-speech model (v1.0, January 2025) with just 82M parameters. It produces natural-sounding speech in 8 languages with 54 voices, runs on around 1 GB of VRAM and ships under the Apache-2.0 license.

NVIDIA Canary Speech & Voice Open-Weight

Canary is NVIDIAs open ASR and speech-translation model family (CC-BY-4.0) that tops the Open ASR Leaderboard on English accuracy ahead of OpenAI Whisper and covers 25 European languages.

Whisper LLM-Grundlagen Speech & Voice

Whisper is OpenAI's open speech-to-text model from 2022 — a multilingual encoder-decoder transformer that turns audio into text and ships in several sizes under an MIT license.

Related Topics