Back to glossary

Term

Mixtral 8x7B

Mixtral 8x7B (December 2023) is Mistral AIs open Sparse-MoE model with 8 experts — 46.7B parameters, 12.9B active per token, 32K context, Apache 2.0 license.

Mixtral 8x7B — explained in more detail

Mistral AI released Mixtral 8x7B on December 11, 2023 as an open-source model under the Apache 2.0 license. It is a Sparse Mixture-of-Experts (SMoE) model with eight expert networks: a learned router selects two of the eight experts per token at each layer and combines their outputs. The model has 46.7B total parameters, of which only about 12.9B are activated per forward pass — hence its high efficiency. The context window spans 32,768 tokens.

Example / Practical use

At release, Mixtral 8x7B outperformed Llama 2 70B on most benchmarks at roughly six times faster inference, and matched or beat GPT-3.5 on common standard benchmarks. Because only a fraction of the parameters is active per token, the model delivers the quality of a much larger model at the compute cost of a smaller one. As an Apache 2.0 model with open weights, it can be freely downloaded, adapted and self-hosted — an important building block of the early open-weight wave.

Distinction

Within the Mistral family, Mixtral stands for the MoE architecture — unlike the dense base model Mistral 7B, where all parameters are active. Compared with later, larger models (such as Mistral Large), Mixtral 8x7B is the early, efficiency-oriented MoE model. The name “8x7B” refers to eight experts based on the 7B architecture, but due to shared components it amounts to 46.7B total parameters rather than 56B.

See everything in one place:Mistral