Term
DeepSeek V3
DeepSeek V3 is an open model from the Chinese provider DeepSeek, released in December 2024. It uses a Mixture-of-Experts architecture with 671 billion parameters (37 billion active per token) and a context window of around 128,000 tokens.
DeepSeek V3 — explained in detail
DeepSeek V3 is an open-weight language model from the Chinese company DeepSeek, released in December 2024. The code is under the MIT licence, the weights are freely available under a dedicated model agreement and commercially usable. This means the model can be run locally or on your own infrastructure.
Architecturally, DeepSeek V3 is a Mixture-of-Experts (MoE) model with a total of 671 billion parameters, of which only around 37 billion are activated per token. This substantially lowers the compute cost per request. It uses Multi-head Latent Attention (MLA) and the DeepSeekMoE architecture; in addition, the model applies an auxiliary-loss-free load-balancing strategy and a multi-token prediction objective during training. The context window spans around 128,000 tokens.
DeepSeek V3 was trained on 14.8 trillion mostly English and Chinese tokens over about 55 days on 2,048 NVIDIA H800 GPUs, at an estimated training cost of roughly 5.6 million US dollars. This cost efficiency attracted particular attention across the industry.
Example / practical relevance
DeepSeek V3 is used where strong performance at low operating cost is needed — for example in self-hosting, in fine-tuning projects or in applications that require data sovereignty. Thanks to the MoE architecture, the compute per token stays moderate despite the large total size. The model is supported across numerous inference backends such as SGLang, vLLM and TensorRT-LLM.
Distinction
As an open-weight model, DeepSeek V3 sets itself apart from purely proprietary models (such as GPT or Claude). Within the DeepSeek line, V3 is the general base model; reasoning models built on top of it (such as the R series) are separately specialised in step-by-step reasoning. Against other open families (Llama, Qwen), V3 stands out through its MoE efficiency and low training cost.