Term
DeepSeek V4-Flash
DeepSeek V4-Flash (July 2026) is an open MoE model with 284B parameters (13B active), a 1M-token context and an MIT license — freely downloadable on Hugging Face.
DeepSeek V4-Flash — explained in more detail
DeepSeek released the production version DeepSeek-V4-Flash (deepseek-ai/DeepSeek-V4-Flash-0731) on July 31, 2026 as an open-source model under the MIT license. It is a Mixture-of-Experts (MoE) model with 284B total parameters — 304B including the attached DSpark speculative-decoding draft module — of which about 13B are activated per token. Technical features include dynamic reasoning controls and advanced sparse-attention mechanisms.
Example / Practical use
The open weights are on Hugging Face (166.9 GB across 48 shards, BF16 tensors in the 0731 GA build); the API runs in public beta. According to the release, the model fits a full 1M-token context in a single 128GB box. Because the weights are freely available under the MIT license, V4-Flash can be run on-premises — relevant for teams with data-protection or cost requirements that do not want to depend on proprietary APIs. Notably, the smaller Flash model reportedly outperformed parts of its own flagship.
Distinction
V4-Flash is the lean, efficient variant of the DeepSeek V4 range below the larger V4-Pro models (up to 1.6T MoE parameters). Unlike proprietary models from OpenAI, Anthropic or Google, DeepSeek stands for open weights: access is not limited to an API but comes through free download and self-hosting of the weights.