Why do modern MoE models pair many fine-grained experts with an always-on shared expert?
answer
- same FLOPs, far more combinations
- narrow experts can afford to be narrow
- the common case shouldn't be duplicated
- one expert nobody has to choose
- smaller matmuls, busier interconnect
basics
~20 sSplitting the feed-forward budget into many small experts multiplies the routing combinations available at the same active-parameter cost, letting each expert specialize narrowly. An always-on shared expert absorbs the general knowledge every token needs, so the routed experts do not each relearn it.
solid answer
~50 sTwo ideas, introduced together in the DeepSeekMoE line and now the copied default. **Fine-grained segmentation**: instead of 8 large experts with top-2 routing, use 64 or 256 experts each roughly an eighth the size, with top-8 routing. Active parameters stay about the same, but the number of distinct expert combinations per token grows enormously, so the model can express far more precisely targeted computations and each expert can specialize instead of being a jack-of-all-trades. **Shared expert isolation**: keep one or two experts always active for every token. Common patterns — the general-purpose transformations every token needs — otherwise have to be duplicated inside every routed expert, wasting capacity through redundancy. Pulling that common knowledge into a shared expert frees the routed experts to hold only what is distinctive. The costs are real: many small matrix multiplications are less efficient on accelerators than a few large ones, top-8 routing means more tokens in flight across devices, and the shared expert's FLOPs are paid on every token.
go deeper
Know that current MoE models use many small routed experts plus one or two experts that run for every token, and that the always-on expert covers what all tokens need.
Explain the mechanism: splitting the same active budget into more, smaller experts multiplies the possible expert combinations per token, and a shared expert removes the duplication of common knowledge across routed experts.
Argue the costs as well: smaller matrix multiplications use accelerators less efficiently, top-8 routing widens per-layer all-to-all traffic, and the shared expert spends active-parameter budget on every token unconditionally.
Own the open part — optimal granularity is hardware- and interconnect-dependent and not settled, and shared experts are a strong default rather than a law. Be able to say what evidence would move you off the default.
## The two design moves Early sparse layers used a handful of large experts and top-1 or top-2 routing. The current mainstream shape — many small routed experts, top-8, plus one or two always-on shared experts — comes from two separate observations that happen to combine well. ### Fine-grained expert segmentation Fix the active-parameter budget. You can spend it as 2 experts of size S, or as 8 experts of size S/4. The arithmetic per token is the same; the *expressiveness of the routing decision* is not. With 8 experts choosing 2, there are 28 possible combinations per token per layer. With 64 experts choosing 8, there are over four billion. That combinatorial space is what lets a token's computation be assembled from narrow, reusable pieces rather than selected from a short menu of general-purpose blocks. The practical consequence is that each expert can afford to be narrow. A large expert that receives a wide variety of tokens must be competent at all of them; a small expert that only ever sees a thin slice of the input distribution can dedicate all its parameters to that slice. Empirically, finer granularity improves quality at equal active parameters — which is why the field moved that way and kept going. ### Shared expert isolation Routed experts are selected per token, so anything *every* token needs must be learned independently by every expert that sees a meaningful share of traffic. That duplication is pure waste: the same general-purpose transformation stored dozens of times, occupying memory that could hold something distinctive. The fix is to designate one or two experts as **shared** — always active, never routed, applied to every token in addition to the top-k routed experts. They become the natural home for the common substrate, and the routed experts' capacity is freed for specialization. This is why the two moves pair so well: fine-grained experts specialize best when they are not each obliged to carry the general case, and the shared expert is what relieves them of it. ## What it costs **Hardware efficiency.** Accelerators like big matrix multiplications. Splitting one large expert FFN into eight smaller ones produces eight smaller GEMMs with lower arithmetic intensity, and the kernel and launch overheads are amortized over less work. Grouped and fused MoE kernels recover much of this, but fine granularity is not free at the hardware level — the FLOP count is unchanged while the achieved efficiency drops. **Communication.** With expert parallelism, top-8 routing means each token may need to reach up to eight devices per layer instead of two. The all-to-all payload and the number of distinct destinations both grow, making the interconnect a more serious constraint. Designs often cap how many devices a single token's experts may span, precisely to bound this. **Always-on FLOPs.** A shared expert runs for every token, so its parameters are active parameters by definition. It is a deliberate spend of the active budget on something that is never skipped, justified only if the redundancy it eliminates is larger than the compute it consumes. **Routing overhead and stability.** More experts means a wider router output, more skew to manage, and more sensitivity to load imbalance — a hundred-way distribution is easier to make lopsided than an eight-way one. ## Is this settled? Mostly, but not universally. Fine-grained routing with shared experts is the copied default across frontier sparse models as of mid-2026, and the direction of travel has been consistently toward more, smaller experts. What remains genuinely open is *how many* — the optimal granularity clearly interacts with hardware, interconnect topology and training scale, and published choices span a wide range rather than converging on a number. Some strong models also ship without a shared expert. Treat the pattern as a well-supported default, not a proof. ## The interview version "Fine-grained experts buy combinatorial routing freedom at constant active parameters, so specialization gets sharper. A shared expert stops every routed expert from having to relearn the common case. You pay for both in kernel efficiency and cross-device traffic." That is the whole argument, and it shows you understand the design as a set of trades rather than a recipe.
- If finer granularity is better, why not use thousands of tiny experts per layer?Because the costs are not in the FLOP count. Smaller matrix multiplications have lower arithmetic intensity and worse accelerator utilization, top-k over a very wide router raises routing overhead, load balancing gets harder as the distribution widens, and with expert parallelism a token's experts may span more devices, inflating all-to-all traffic. The optimum is set by hardware and interconnect, and published choices vary rather than converging.
- Are the shared expert's parameters counted as active parameters?Yes. Active parameters mean everything that computes for a given token, which includes attention, embeddings, the routed top-k experts and any always-on shared expert. That is why a shared expert is a deliberate spend rather than a free addition — it consumes part of the active budget on every single token, and only pays for itself if the redundancy it removes from the routed experts is worth more.
- Does a shared expert make routing collapse less likely?It helps at the margin but does not solve it. Because the common case has a guaranteed home, the router is under less pressure to funnel generic tokens toward whichever routed expert learned generic behaviour best. But the reinforcing loop between traffic and expert quality among the routed experts is untouched, so explicit balancing — typically a per-expert routing bias adjusted between steps — is still required.
saying these in an interview costs you the question
- Claims fine-grained experts increase active parameters per token
- Says the shared expert is chosen by the router like the others
- Assumes more experts always improves throughput
- Treats each expert as owning a human-readable domain
- Ignores that smaller experts mean less efficient matrix multiplications