Term
DeepSeek V3.2
DeepSeek V3.2 (December 1, 2025) is an open MoE model with 685B parameters under an MIT license — especially efficient on long contexts thanks to DeepSeek Sparse Attention.
DeepSeek V3.2 — explained in more detail
DeepSeek released V3.2 on December 1, 2025 as an open model with freely available weights under an MIT license. Unlike proprietary models such as GPT or Claude, V3.2 can be downloaded and self-hosted. Technically it is a very large Mixture-of-Experts transformer with roughly 685 billion parameters (about a 690 GB download). At the heart of version 3.2 are two efficiency mechanisms: DeepSeek Sparse Attention (DSA) and Multi-Head Latent Attention (MLA). DSA markedly reduces compute, above all on long contexts, without a large drop in quality.
Alongside V3.2 came the V3.2-Speciale variant, trained exclusively on deep reasoning tasks, which reached gold-level results at competitions such as IMO, CMO, ICPC World Finals and IOI 2025.
Example / Practical use
For users, sparse attention concretely means lower costs on long inputs — for example when analyzing large codebases or document collections. DeepSeek also cites a new method for synthesizing agent training data across 1,800+ environments and 85,000+ complex instructions; V3.2 integrates thinking directly into tool use (thinking-in-tool-use). Because the weights are open, the model suits organizations with data-sovereignty requirements that want to run inference on their own infrastructure — rather than being tied to a proprietary API.
Distinction
DeepSeek V3.2 belongs to the class of open-weight models and thus contrasts with the closed APIs of OpenAI, Anthropic or Google. Within the DeepSeek V line, 3.2 is the successor to V3.1 and sharpens above all the efficiency on long context. The Speciale variant is not a general-purpose model but a reasoning-specialized offshoot.