skip to content

An MoE model has 670B total and 37B active parameters — which number sizes GPU memory?

level: middleimportance: must knowfreq 66%

answer

  1. two numbers, two different budgets
  2. any token can reach any expert
  3. one number sets the entry ticket
  4. the other sets per-token arithmetic
  5. big to hold, small to run

basics

~20 s

Total parameters size memory. Every expert's weights must be resident because any token may route to any expert. Active parameters size compute — the arithmetic and latency per token. A sparse model is therefore memory-hungry like a huge model but computes like a small one.

solid answer

~50 s

The two numbers govern different resources, and they pull in opposite directions. **Total parameters** (670B) determine how much accelerator memory the weights occupy, because routing is decided token by token: any expert may be needed at any moment, so all of them must be loaded. That sets how many GPUs you need before you serve a single request. **Active parameters** (37B) determine the arithmetic per token — the FLOPs, and therefore the compute cost and much of the latency of generating each token. So you buy hardware sized like a 670B model and get per-token compute closer to a 37B one. That is the whole bargain: cheap cost *per token*, expensive cost *per GPU-hour of residency*. It also means utilization matters — a lightly used deployment pays the full memory bill to serve very few tokens.

go deeper

for a junior

Remember the pair: total parameters say how much memory the model needs, active parameters say how much work each token costs. Say plainly that all experts must be loaded even though few of them run.

for a middle

Explain why all experts must be resident — routing is decided per token per layer, so the needed expert is unknowable in advance — and connect active parameters to FLOPs, latency and per-token cost.

for a senior

Turn it into a sizing decision: fit is a total-parameter question, throughput is an active-parameter question, and expert sharding moves memory between devices without reducing the aggregate. Mention the all-to-all traffic that sharding introduces.

for a principal

Own the economics: sparse models make cost per token cheap and cost per resident GPU-hour expensive, so their value depends on utilization. Be ready to say when that bargain is wrong for a given traffic profile.

## Two numbers, two budgets Every sparse model is quoted with a pair: total parameters and active (or activated) parameters per token. Reading them as one number is the classic mistake, in both directions — assuming you only need memory for the active slice, or assuming a 670B model must be slow. **Total parameters -> memory.** The weights of every expert must be resident on the accelerators before inference starts. There is no way around this in the normal case, because the router decides per token, per layer: the expert needed for the next token is not known until the previous layers have run. Nothing can be predictably left out. So memory footprint, and therefore the number of accelerators, tracks the total count. **Active parameters -> compute.** For any single token, only the router-selected experts multiply. The matrix multiplications performed per token — and the energy, FLOPs and a large part of the generation latency — track the active count, plus the always-dense parts (attention, embeddings, any shared expert). "Active parameters" is normally quoted inclusive of those dense parts, so 37B active means roughly the arithmetic of a dense 37B model. ## What this makes cheap and what it makes expensive - **Cheap:** cost per generated token, given that the machine is already up and busy. You are paying 37B-model arithmetic for a model that knows what a 670B parameter budget can hold. - **Expensive:** the entry ticket. You cannot serve the model at all until you have the memory for all 670B parameters, and that hardware is sitting there whether you send it one request per minute or thousands per second. This is why sparse models suit high-volume serving and suit low-volume self-hosting badly. It also explains a pricing observation people find confusing: providers can offer very large sparse models at per-token prices closer to those of far smaller dense ones, because per-token cost is set by active parameters while their capital cost is set by total. ## Where the memory actually goes Across a fleet, expert weights are usually **sharded** rather than replicated: with expert parallelism, each device holds a slice of the experts for each layer, and tokens are exchanged between devices — an all-to-all dispatch and combine per MoE layer — so each token reaches the device holding its expert. That divides the memory per device, but the *aggregate* memory still equals the total parameter count, plus per-device runtime overheads. Sharding changes who holds the bytes, not how many bytes exist. ## Does an MoE with 37B active behave like a dense 37B model? On speed, roughly yes. On quality, no — it is clearly stronger, because it has far more total capacity to store knowledge and specialized behaviour. Equally, it is weaker than a *dense* model of the same 670B total would be, if anyone could afford to train and serve one. The honest framing is that sparsity buys you a better quality-per-FLOP curve and a worse quality-per-byte-of-memory curve than dense. ## How to reason about a serving decision Ask two separate questions and do not let one answer the other: 1. *Can I fit it?* Driven by total parameters, plus room for runtime state. If the answer is no, nothing else matters. 2. *Can I serve my traffic on what I fit?* Driven by active parameters and how much work the hardware can do per second. A team that sizes by active parameters discovers at deploy time that the model does not load. A team that sizes by total parameters and then assumes matching slowness over-provisions compute and under-uses the accelerators they bought. ## The one-sentence version VRAM is sized by total parameters; FLOPs, latency and cost per token are sized by active parameters — the sparse model's whole point is that those two numbers are allowed to be far apart.

  • If you shard experts across eight GPUs, does that reduce the model's total memory requirement?
    No — it divides it. With expert parallelism each device holds a slice of each layer's experts, so per-device memory drops roughly eightfold, but the aggregate is still the full total parameter count. What sharding adds is traffic: each MoE layer needs an all-to-all exchange to send tokens to the device holding their expert and bring the results back.
  • Would a dense model with the same 37B parameters give the same quality?
    No, it would be meaningfully weaker. The sparse model has roughly the same arithmetic per token but far more total capacity to store knowledge, and that extra capacity shows up in quality. The other direction also holds: a dense 670B model would beat the sparse one, which is why sparsity is best described as trading quality-per-byte for quality-per-FLOP.
  • Why can a provider price a very large sparse model close to a much smaller dense one?
    Because per-token price tracks per-token compute, which tracks active parameters. The provider's capital cost tracks total parameters, but that cost is amortized across very high utilization, so at scale the marginal token really is cheap. The same arithmetic works badly for a low-traffic private deployment, where the idle memory has nothing to amortize against.

saying these in an interview costs you the question

  • Assumes only the active parameters need to be loaded
  • Expects a 670B sparse model to be as slow as dense 670B
  • Thinks sharding experts reduces total memory needed
  • Says active parameters can be predicted and preloaded ahead of time
  • Treats the sparse model as equal in quality to a dense model of the same total size

context