In a transformer, what does a Mixture-of-Experts layer replace, and how does routing work?
answer
- only one sublayer becomes sparse
- attention is untouched
- a tiny scorer picks a few
- decided per token, per layer
- chosen scores become blend weights
basics
~20 sA Mixture-of-Experts layer swaps a transformer block's feed-forward sublayer for many parallel feed-forward experts plus a small router. The router scores experts for each token and activates only the top few, so most expert weights compute nothing for that token.
solid answer
~50 sSparsity in a Mixture-of-Experts (MoE) model lives in the **feed-forward sublayer**, which normally holds the majority of a block's parameters. Instead of one shared feed-forward network, the block holds N independent expert feed-forward networks and a tiny router — usually a single linear projection from the hidden state to N affinity scores. For every token, at every MoE layer, the router takes the top-k experts (k is small, commonly 2 or 8 out of tens or hundreds), normalizes their scores into gating weights, runs the token through only those experts, and returns the gate-weighted sum of their outputs. Everything else stays dense: attention projections, embeddings, normalization layers. Routing is per token and per layer, not per request, so two tokens in the same sentence usually take different experts, and the same token takes different experts at different depths.
code
python · 14 linesimport numpy as np
def moe_layer(x, router_w, expert_ws, top_k=2):
scores = x @ router_w # one affinity score per expert
idx = np.argsort(-scores)[:top_k] # top-k selection (discrete)
gates = np.exp(scores[idx] - scores[idx].max())
gates /= gates.sum() # renormalize over chosen experts
return sum(g * (x @ expert_ws[e]) for g, e in zip(gates, idx))
rng = np.random.default_rng(0)
x = rng.normal(size=4)
router_w = rng.normal(size=(4, 8)) # 8 experts
expert_ws = rng.normal(size=(8, 4, 4))
print(moe_layer(x, router_w, expert_ws)) # 2 of 8 experts contributedgo deeper
Be able to say that an MoE model contains many feed-forward experts and a router that picks a few per token, so the model has far more parameters than it uses on any single token.
Explain the mechanics precisely: the FFN sublayer is replaced, the router emits affinity scores, top-k experts run, their outputs are combined with normalized gate weights, and the decision is remade for every token at every MoE layer.
Show you know what this costs to run: all experts occupy memory regardless of routing, routing skew wastes capacity, and sharding experts across devices adds an all-to-all exchange per layer that makes interconnect a real constraint.
Own the framing that MoE decouples capacity from per-token compute, and be ready to argue where that decoupling is worth its memory and operational bill for your workload rather than treating sparsity as free quality.
## The problem MoE is solving In a dense transformer, every parameter participates in every token's forward pass. Quality improves with parameter count, but so does the cost of each token, one-for-one. Mixture-of-Experts (MoE) breaks that coupling: it adds parameters that only *some* tokens use, so total capacity grows much faster than per-token compute. ## What actually gets replaced A standard decoder block has two sublayers: attention, then a position-wise feed-forward network (FFN). The FFN — two or three large matrices with a nonlinearity between them — typically holds most of the block's parameters. An MoE block keeps attention untouched and replaces that single FFN with: - **N experts**, each an independent FFN with its own weights, usually all the same shape; - **a router (gate)**, a small linear map from the token's hidden vector to N affinity scores. Some architectures make only every other block an MoE block and leave the rest dense; early layers are sometimes kept dense too, because routing there is noisier. ## How a token flows through 1. The hidden state for one token position enters the layer. 2. The router produces N scores (a softmax or sigmoid over expert affinities). 3. **Top-k selection** picks the k highest-scoring experts. k is small and fixed — top-1 in the earliest sparse models, top-2 in the Mixtral generation, top-8 in the fine-grained designs that dominate now. 4. The k chosen scores are re-normalized into **gating weights** that sum to 1. 5. The token is run through those k expert FFNs only. 6. The layer output is the gate-weighted sum of the k expert outputs, added back to the residual stream. The unit of routing is **one token at one layer**. This is the single most common misunderstanding: MoE does not route a prompt, a request, or a sequence to "the coding expert". A 20-token prompt in a 60-layer model with top-2 routing makes 20 x 60 independent routing decisions, each selecting 2 of N experts. ## What stays dense Attention (query/key/value/output projections), token embeddings and the output head, normalization layers, and the router itself all run for every token. Designs that include an always-on **shared expert** also run that one for every token. So a "sparse" model is only sparse in one sublayer — but that sublayer is where most of the parameters are, which is why the saving is large. ## Is the router trained? Yes, jointly with everything else. The top-k *selection* is a discrete, non-differentiable operation, but the gating weights multiply the expert outputs, so gradients flow back into the router through those weights: an expert that produced a useful output pushes its gate up. This is also why MoE training needs explicit load-balancing pressure — left alone, the router happily concentrates traffic on a handful of experts that got a good start. ## What do experts specialize in? Rarely anything a human would name. Analyses of trained MoE models find routing correlates far more with token identity, syntax and surface form than with topics like "medicine" or "French". Treat the panel-of-consultants picture as an intuition pump about *cost*, not a claim about interpretability: it is honest to say the router consults 2 of 64 specialists per token, and dishonest to say expert 17 is the SQL expert. ## The consequences worth naming - **Compute per token** is set by the experts that actually ran (the *active* parameters), which is why an MoE serves at roughly small-model speed. - **Memory** is set by all experts, because any token can route anywhere; all weights must be resident. - **Balance** becomes an engineering problem: uneven routing wastes capacity and, in deployments that shard experts across devices, turns one overloaded device into everyone's straggler. - **Communication**: when experts live on different accelerators, each MoE layer needs an all-to-all dispatch of tokens and an all-to-all combine of results, so interconnect quality starts to matter. ## Saying it well in an interview Lead with the sublayer ("the FFN, not attention"), then the granularity ("per token, per layer"), then the payoff ("total capacity grows, per-token FLOPs don't"), then the bill ("but every expert still occupies memory"). That sequence covers what an interviewer is listening for.
- If the router sent every token to all N experts, what would you have built?A dense model with extra steps. Routing to all experts makes the layer a gate-weighted average of N feed-forward networks, so you pay the FLOPs of all N while gaining nothing over a single wider FFN. The entire benefit comes from k being much smaller than N; sparsity is the product, not a side effect.
- The top-k selection is not differentiable — so how does the router learn anything?Gradients reach the router through the gating weights, not through the selection. The chosen experts' outputs are multiplied by their gate values, so if an expert's output reduced the loss, the gradient pushes that gate up and the router's affinity for similar tokens rises. Unselected experts get no signal that step, which is exactly why training needs explicit load-balancing pressure.
- Does every transformer block in an MoE model have to be an MoE block?No. Many designs interleave dense feed-forward blocks with MoE blocks, and it is common to keep the first layers dense because routing decisions on barely-contextualized token representations are noisy and unstable. The mix is an architecture choice: more MoE layers means more total capacity per unit of active compute, and more routing and communication overhead.
saying these in an interview costs you the question
- Says MoE routes a whole prompt or request to one expert
- Claims experts map to human topics like math or French
- Thinks attention is what gets sparsified
- Describes MoE as an ensemble that votes on outputs
- Believes only the selected experts need to be in memory