Term
AssemblyAI Universal-2
Universal-2 is AssemblyAIs closed speech-to-text model with industry-leading accuracy (mean WER around 6 percent) and strong formatting, punctuation and proper-noun recognition across 99 languages.
AssemblyAI Universal-2 — explained in more detail
Universal-2 is the commercial speech-to-text (ASR) model from the US provider AssemblyAI and the successor to Universal-1. It is a closed model available only through AssemblyAIs cloud API; the weights are not public. The model covers around 99 languages and, according to the provider, reaches a mean word error rate (WER) of roughly 6 percent across many datasets. Compared with Universal-1 it lowers WER by about 3 percent relative and surpasses the nearest external model by about 15 percent relative.
Universal-2 focuses not only on raw word recognition but on production-ready output: the architecture includes an all-neural text-formatting component (Universal-2-TF). According to AssemblyAI this improves proper-noun recognition by about 24 percent, formatting and casing by about 15 percent, and timestamp accuracy substantially. This combination positions Universal-2 as an accuracy leader among commercial transcription services.
Example / In practice
A typical use is automatic transcription of podcasts, meetings or support calls: audio is sent to the API, and Universal-2 returns already formatted text with correct punctuation, casing and precise timestamps. Especially for technical terms, product and person names, the improved proper-noun recognition reduces manual post-editing effort.
Distinction from similar terms
Universal-2 sits as a closed, accuracy-optimized batch model between streaming-oriented services such as Deepgram Nova-3 (latency leader for real time) and open models such as NVIDIA Canary or OpenAI Whisper, which can be self-hosted. While open models offer control and local operation, Universal-2 emphasizes maximum output quality and ready-made formatting as a managed service.
Discover more
Better speech recognition in boostN-CLI: Audio normalization + a flexible Whisper model
Why dictated text got swallowed, how SoX normalization and a model switcher fixed it — and what is actually happening under the hood.
GlossaryElevenLabs v3
ElevenLabs v3 is a highly expressive text-to-speech model with inline audio tags for emotion, emphasis and non-verbal sounds. It supports more than 70 languages and multi-speaker dialogue and is seen as a benchmark for realism.
EncyclopediaComparing AI models — who builds what and how to choose
The major model families in 2026 at a glance. Who builds Claude, GPT, Gemini, Llama, Mistral, DeepSeek, Qwen — and which model to pick when.