AssemblyAI Universal-2
Universal-2 is AssemblyAIs closed speech-to-text model with industry-leading accuracy (mean WER around 6 percent) and strong formatting, punctuation and proper-noun recognition across 99 languages.
Sprach-KI in beide Richtungen — Spracherkennung/Transkription (STT) und Sprachsynthese (TTS), anbieterübergreifender Einsatzzweck, z. B. Whisper, Kokoro, ElevenLabs.
Universal-2 is AssemblyAIs closed speech-to-text model with industry-leading accuracy (mean WER around 6 percent) and strong formatting, punctuation and proper-noun recognition across 99 languages.
Chatterbox is an open-source text-to-speech model from Resemble AI under the MIT license. In blind tests most listeners preferred it over ElevenLabs; it offers emotion control, voice cloning and low latency.
Nova-3 is Deepgrams real-time speech-to-text model (released February 2025) with very low streaming latency of around 200 to 300 milliseconds, keyterm prompting and support for more than 30 languages.
ElevenLabs v3 is a highly expressive text-to-speech model with inline audio tags for emotion, emphasis and non-verbal sounds. It supports more than 70 languages and multi-speaker dialogue and is seen as a benchmark for realism.
Fish Audio S2 is an open-source text-to-speech model from the Fish Speech family. With inline tags it controls emotion and emphasis at word level, supports more than 80 languages and delivers very low latency.
Kokoro is an open text-to-speech model (v1.0, January 2025) with just 82M parameters. It produces natural-sounding speech in 8 languages with 54 voices, runs on around 1 GB of VRAM and ships under the Apache-2.0 license.
Canary is NVIDIAs open ASR and speech-translation model family (CC-BY-4.0) that tops the Open ASR Leaderboard on English accuracy ahead of OpenAI Whisper and covers 25 European languages.
Whisper is OpenAI's open speech-to-text model from 2022 — a multilingual encoder-decoder transformer that turns audio into text and ships in several sizes under an MIT license.