Speech & Voice
Sprach-KI in beide Richtungen — Spracherkennung/Transkription (STT) und Sprachsynthese (TTS), anbieterübergreifender Einsatzzweck, z. B. Whisper, Kokoro, ElevenLabs.
- AssemblyAI Universal-2 Speech & Voice
Universal-2 is AssemblyAIs closed speech-to-text model with industry-leading accuracy (mean WER around 6 percent) and strong formatting, punctuation and proper-noun recognition across 99 languages.
- Chatterbox Speech & Voice Open-Weight
Chatterbox is an open-source text-to-speech model from Resemble AI under the MIT license. In blind tests most listeners preferred it over ElevenLabs; it offers emotion control, voice cloning and low latency.
- Deepgram Nova-3 Speech & Voice
Nova-3 is Deepgrams real-time speech-to-text model (released February 2025) with very low streaming latency of around 200 to 300 milliseconds, keyterm prompting and support for more than 30 languages.
- ElevenLabs v3 Speech & Voice
ElevenLabs v3 is a highly expressive text-to-speech model with inline audio tags for emotion, emphasis and non-verbal sounds. It supports more than 70 languages and multi-speaker dialogue and is seen as a benchmark for realism.
- Fish Audio S2 Speech & Voice Open-Weight
Fish Audio S2 is an open-source text-to-speech model from the Fish Speech family. With inline tags it controls emotion and emphasis at word level, supports more than 80 languages and delivers very low latency.
- Kokoro Speech & Voice Open-Weight
Kokoro is an open text-to-speech model (v1.0, January 2025) with just 82M parameters. It produces natural-sounding speech in 8 languages with 54 voices, runs on around 1 GB of VRAM and ships under the Apache-2.0 license.
- NVIDIA Canary Speech & Voice Open-Weight
Canary is NVIDIAs open ASR and speech-translation model family (CC-BY-4.0) that tops the Open ASR Leaderboard on English accuracy ahead of OpenAI Whisper and covers 25 European languages.
- Whisper LLM-Grundlagen Speech & Voice
Whisper is OpenAI's open speech-to-text model from 2022 — a multilingual encoder-decoder transformer that turns audio into text and ships in several sizes under an MIT license.