Back to glossary

Term

Deepgram Nova-3

Nova-3 is Deepgrams real-time speech-to-text model (released February 2025) with very low streaming latency of around 200 to 300 milliseconds, keyterm prompting and support for more than 30 languages.

Deepgram Nova-3 — explained in more detail

Nova-3 is the current speech-to-text model from the US provider Deepgram, released in February 2025. It is a closed model delivered through Deepgrams cloud API and via marketplaces; the weights are not public. Its focus is real-time transcription: Nova-3 streams audio over a WebSocket endpoint with end-to-end latency of around 200 to 300 milliseconds under good conditions, returning interim transcripts that may be revised, then finalized transcripts that are locked as more audio arrives.

On accuracy, Deepgram reports a median streaming WER of about 6.8 percent and roughly 5.3 percent in batch mode. The model supports multilingual real-time transcription across more than 30 languages, offers self-serve keyterm prompting (up to 100 terms) for domain vocabulary, and real-time redaction of sensitive data. This combination of low latency and customization makes Nova-3 a leader for real-time and streaming among commercial ASR services.

Example / In practice

A typical use case is voice assistants and live captions: in a voice agent or call-center system, microphone audio is streamed continuously to Nova-3, which returns the spoken text with minimal delay. This lets an agent react while the speaker is still talking, and domain-specific terms can be recognized reliably via keyterm prompting.

Distinction from similar terms

While AssemblyAI Universal-2 targets maximum accuracy and ready-made formatting in batch mode, Nova-3 is optimized for low latency in real-time streaming — serving the use-case segment of voice agents and live transcription. Unlike open models such as NVIDIA Canary, Nova-3 is a closed managed service without self-hosting of the weights.

See everything in one place:Speech & Voice