Why does Mixtral 8x7B total 46.7B parameters instead of 56B?
answer
- the name is a label, not multiplication
- not every part of a layer is duplicated
- two numbers per sparse model, not one
- one sizes storage, the other sizes compute
- 46.7B total, ~12.9B active
basics
~20 sThe 8x7B name is a naming convention, not multiplication. Only the feed-forward blocks are replicated into 8 experts; attention and embedding weights are shared across them, so Mixtral 8x7B holds 46.7B total parameters and uses about 12.9B per token.
solid answer
~50 sMistral's `8x7B` label says "eight experts, roughly a 7B-class base", and people mis-read it as 8 x 7 = 56B. Only the expert (feed-forward) portion of each layer is duplicated eight times; the attention weights, embeddings and normalisation layers are shared by all experts and counted once. That gives Mixtral 8x7B **46.7B total parameters**, of which roughly **12.9B are active per token** because the router selects two experts per token per layer. The larger sibling follows the same rule: Mixtral 8x22B is **141B total, about 39B active**, not 176B. Both numbers matter for different reasons — the total is what you must load and store, the active count is what drives per-token compute and therefore latency. Quoting the wrong one is how capacity plans go wrong. Both models are Apache 2.0 and were the open MoE flagships of the 2023-24 generation.
go deeper
Recall that Mixtral 8x7B is 46.7B parameters in total with about 12.9B used per token, and that the name is a convention rather than a multiplication you perform.
Explain why only part of each layer is replicated across experts, and state clearly which of the two numbers describes storage and which describes per-token compute.
Demonstrate that you quote total parameters when planning residency and active parameters when reasoning about throughput and latency, and that you read both from the model card rather than the name.
Frame it as a catalogue-currency judgment: sparse totals complicate capacity planning, and by mid-2026 a dense Small-class model is often the better default than the Mixtral generation it replaced.
## Reading Mistral's model names Mistral names its sparse models `NxMB`, where N is the number of experts per MoE layer and M gestures at the size class of the base dense model. It is a label, not an arithmetic expression. Candidates who multiply get 56B for Mixtral 8x7B and 176B for Mixtral 8x22B, and both are wrong. The published figures are: - **Mixtral 8x7B** — 46.7B total parameters, ~12.9B active per token, 32k context, Apache 2.0, released December 2023. - **Mixtral 8x22B** — 141B total parameters, ~39B active per token, 64k context, Apache 2.0, released April 2024. ## Why the totals are smaller than the multiplication A transformer layer is not one undivided block. In Mistral's sparse models only the feed-forward sublayer is replaced by a bank of eight expert feed-forward networks; the attention projections, the token embeddings, the output head and the normalisation parameters exist once per layer and are used by every token regardless of which experts fire. Duplicating only part of each layer eight times therefore multiplies only part of the parameter budget. That is the entire arithmetic reason the total lands at 46.7B rather than 56B. ## Why the active count is smaller still For each token at each layer a lightweight router picks a small subset of the experts — two of the eight in Mixtral's case. The other six experts contribute nothing to that token's forward pass. So the FLOPs per token correspond to a much smaller model than the parameter count suggests: about 12.9B active for 8x7B, about 39B for 8x22B. This is the whole selling point of the design — near-dense-large quality at near-dense-small compute. ## Which number you quote, and when This is where the practical judgment sits, and where interviewers push. - **Storage and memory footprint follow the total.** Every expert has to be resident and reachable, because the next token may route to any of them. A 46.7B model is a 46.7B download and a 46.7B residency problem, and that does not shrink because only two experts fire. - **Compute and per-token latency follow the active count.** Throughput per accelerator, and the cost of a token when you are billed by compute, track the ~12.9B figure. - **Benchmark comparisons need both.** Comparing Mixtral 8x7B against a dense 13B on quality alone flatters the dense model's memory profile; comparing against a dense 47B on quality alone flatters Mixtral's speed. State both axes. ## Where Mixtral sits in the line-up as of mid-2026 Mixtral was the model that made open MoE mainstream, and both variants remain Apache-2.0 and freely deployable. But the catalogue has moved on: the dense Mistral Small line (24B class, Apache 2.0 from Small 3 onward) generally matches or beats Mixtral 8x7B on quality while being far simpler to serve and far smaller to hold in memory, and it added multimodal input in later revisions. For a new self-hosting project, Mixtral is now usually the wrong default — you reach for it when you specifically want the MoE throughput profile at 8x22B scale, or when you have existing fine-tunes on it. Knowing that the flagship of eighteen months ago is no longer the default is itself part of the answer. ## Related naming traps The same misreading appears across the industry whenever a vendor publishes a sparse model with a total/active split, and vendors have not standardised which number goes in the name — some name the total, some the active count, some use the `NxM` style. The safe habit is to ignore the name and read the two published figures from the model card: total parameters and active parameters per token. If a model card gives only one number, find out which one it is before you size anything. ## Common wrong answers "It's 56B but they quantised it" — no, quantisation changes bytes per parameter, not parameter count. "The experts are shared between layers" — no, each MoE layer has its own bank of experts. "Only two experts are stored" — no, all eight are stored; only two run.
- What are the equivalent figures for Mixtral 8x22B, and does the same reasoning hold?Mixtral 8x22B is 141B total parameters with roughly 39B active per token, and a 64k context window — not the 176B that multiplying would suggest. The reasoning is identical: only the feed-forward experts are replicated eight times while attention, embeddings and normalisation are shared, and the router activates two experts per token per layer.
- Would you still choose Mixtral 8x7B for a new self-hosted deployment in 2026?Usually not. The dense Mistral Small line at 24B class, Apache 2.0 since Small 3, generally matches or beats Mixtral 8x7B on quality while being roughly half the weights to hold and much simpler to serve, with multimodal input in later revisions. Mixtral stays relevant at the 8x22B scale for throughput-oriented workloads, or where you already own fine-tunes on it.
- Does quantising Mixtral change either of those two parameter numbers?No. Quantisation changes bytes per parameter, not the count of parameters. A 4-bit Mixtral 8x7B still has 46.7B parameters and still activates about 12.9B per token; what shrinks is the memory footprint per parameter and, with it, the total bytes you must hold. Reporting a quantised model as "smaller" in parameter terms is a category error.
saying these in an interview costs you the question
- Multiplying 8 by 7B to get 56B parameters
- Claiming only the selected experts need to be stored
- Saying quantisation explains the 46.7B figure
- Quoting active parameters when sizing memory footprint
- Assuming all eight experts share one bank across layers