Three of 64 experts take most tokens in your MoE — what breaks and how do you fix it?
answer
- a feedback loop, not a bug
- cold experts still cost memory
- synchronous collectives wait for the slowest
- balance without a competing loss term
- watch max-to-mean load per layer
basics
~20 sThat is routing collapse. Most of the model's capacity goes unused while the hot experts saturate; in a deployment that shards experts across devices, the devices holding them become stragglers and throughput drops to their speed. Fixes push load back toward the idle experts.
solid answer
~60 sSkewed routing costs you on two axes. **Quality**: the parameters in the 61 cold experts are effectively wasted, so you are paying to hold a 64-expert model and getting the capacity of a handful — and during training the collapse is self-reinforcing, since experts that win tokens improve fastest and win more. **Throughput**: with expert parallelism, tokens must be shipped to the device that owns their expert, so an all-to-all exchange per layer finishes only when the busiest device finishes. Three hot experts means a few pinned devices and a lot of idle silicon. In training frameworks with a fixed per-expert capacity, the overflow tokens are simply dropped, silently degrading those positions. The mainstream fixes as of mid-2026 are a per-expert **routing bias** nudged up for underloaded experts and down for overloaded ones between steps, which steers selection without adding a term that fights the language-modelling objective; the older approach is an explicit auxiliary load-balancing loss. At serving time you can also rebalance or replicate hot experts across devices.
code
python · 17 linesimport numpy as np
n_experts, top_k, update_rate = 8, 2, 1e-3
bias = np.zeros(n_experts)
rng = np.random.default_rng(0)
def select(scores, bias, top_k):
# bias steers selection only; gating weights still use raw scores
return np.argsort(-(scores + bias))[:top_k]
counts = np.zeros(n_experts)
for _ in range(512): # tokens in one training step
scores = rng.normal(size=n_experts) + np.array([2, 0, 0, 0, 0, 0, 0, 0])
counts[select(scores, bias, top_k)] += 1
bias += update_rate * np.sign(counts.mean() - counts) # starve -> up, hot -> down
print(counts, counts.max() / counts.mean(), bias)go deeper
Know that MoE routing can concentrate on a few experts, that this wastes the rest of the model, and that training uses explicit mechanisms to keep expert load roughly even.
Explain the feedback loop that causes collapse, and describe both balancing mechanisms: an auxiliary loss added to training, and a per-expert routing bias adjusted between steps that steers selection only.
Show the operational half: synchronous all-to-all makes the busiest device set the pace, dropped tokens degrade output silently, and you diagnose with per-expert token counts and a max-to-mean load ratio before latency alerts fire.
Frame balance as a permanent tax on sparsity and decide where to pay it — training-time objective pressure, structural routing schemes, or serving-time replication and placement — and set the telemetry and thresholds the organization runs on.
## What collapse is The router is trained jointly with everything else and gets no intrinsic reward for spreading tokens around. Early in training a few experts are marginally better by luck; they receive more tokens, so they get more gradient, so they get better, so the router prefers them more. Left unchecked this is a positive feedback loop that ends with a small clique of experts absorbing most traffic. The standard name is **routing collapse**, and it is the defining failure mode of sparse models. ## The two costs **Capacity you paid for and cannot use.** Memory is sized by *total* parameters, so cold experts occupy accelerator memory whether or not they see tokens. A 64-expert layer where three experts take most of the traffic is, functionally, a much smaller model with a very large memory bill. Quality plateaus below what the parameter count suggests. **Throughput collapse under expert parallelism.** In a sharded deployment, each MoE layer performs an all-to-all: send each token to the device holding its chosen expert, compute, and send results back. Collective operations are synchronous, so the layer takes as long as the slowest participant. When three experts sit on one or two devices, those devices do most of the work and everyone else waits. Aggregate utilization falls even though every GPU is powered on. This is the incident shape people actually hit: latency and tokens-per-second crater, GPU utilization graphs go bimodal, and nothing in the model's output obviously changed. **Dropped tokens.** Many training implementations give each expert a fixed **capacity** — a per-batch token limit derived from a capacity factor. Tokens that arrive at a full expert are dropped: they skip the expert entirely and pass through on the residual connection. That is silent quality loss concentrated on exactly the tokens the router thought needed a specific expert. ## How it is prevented **Auxiliary load-balancing loss (the classic method).** Add a term to the training loss that penalizes the correlation between the fraction of tokens sent to each expert and the router's mean probability for it, so the optimizer is pushed toward uniform assignment. It works, and it has a well-known drawback: it is a second objective competing with the language-modelling objective, and turning it up enough to guarantee balance measurably costs quality. Related stabilizers such as a router z-loss on the logit magnitudes were added to keep gating numerically well behaved. **Auxiliary-loss-free balancing via routing bias (the mid-2026 default).** Keep a per-expert scalar bias. Add it to the affinity scores *only for the top-k selection*, not to the gating weights that scale the expert outputs. After each step, compare each expert's observed token count against the average and nudge its bias up if it is starved, down if it is saturated. Balance is achieved by an outside control loop rather than by a gradient term, so it never trades against the language objective. Popularized by DeepSeek-V3 and now widely copied; often paired with a very small sequence-level balancing term to stop any single sequence from being pathological. **Routing schemes that make imbalance structurally impossible.** Expert-choice-style routing inverts the assignment — each expert selects its own quota of tokens — which guarantees perfect balance by construction, at the cost of some tokens getting more experts than others and of being awkward for causal decoding. ## What you do at serving time At inference the weights are frozen; you cannot retrain the router in the middle of an incident. Your levers are placement and replication: measure per-expert token counts, then re-shard so that hot experts do not co-locate, or replicate the hottest experts across several devices so their traffic can be split. Some serving stacks do this adaptively from observed traffic. If the skew is severe and structural, it is a training-time defect that placement can only paper over. ## How you would know Instrument it. Per-expert token counts per layer, and a **max-to-mean load ratio** as a single scalar per layer, are the direct signals; a ratio near 1 is healthy, and a ratio in the tens means collapse. Secondary signals: bimodal GPU utilization across the expert-parallel group, rising all-to-all wait time, and — during training — a dropped-token rate above zero. Alert on the load ratio and on dropped tokens, not on end-to-end latency alone, because latency tells you something is wrong long after the routing telemetry did.
- Why is a per-expert routing bias preferred over an auxiliary balancing loss?Because the auxiliary loss is a second objective added to the training loss, and it competes with the language-modelling objective — strong enough to force balance is usually strong enough to cost quality. The bias is applied only to top-k selection and adjusted by an outside control loop between steps, so it steers routing without contributing gradient to the model's objective. It is the mainstream default as of mid-2026.
- What is a capacity factor, and what happens to the tokens that exceed it?Capacity factor sets a per-expert token quota per batch, sized as some multiple of the even share, so implementations can use fixed-shape buffers. Tokens routed to a full expert are dropped — they skip the expert and continue on the residual path with no expert contribution. Raising the factor reduces dropping but wastes memory and compute on padding, so it is a direct memory-versus-fidelity dial.
- You are mid-incident on a served model with three hot experts. What can you actually change?Not the router — the weights are frozen. You change placement: measure per-expert token counts, re-shard so the hot experts do not sit on the same devices, and replicate the hottest experts so their traffic can be split across several. That restores utilization. If the skew is inherent to the checkpoint, placement only bounds the damage and the real fix is at training time.
saying these in an interview costs you the question
- Says imbalance only affects quality, not throughput
- Thinks the router self-balances because it is trained
- Believes dropped tokens raise an error rather than silently degrading output
- Proposes retraining the router live during a serving incident
- Applies the balancing bias to the gating weights as well as to selection