Which layers should LoRA adapters target, and why is attention-only a weak default?
answer
- ask where the parameters actually live
- the feed-forward block dwarfs attention
- compare at matched parameter count
- experts hold the mass in a sparse model
- embeddings only when the vocabulary changes
basics
~20 sTarget every linear projection, not just attention. Adapting the feed-forward up, gate and down projections alongside query, key, value and output matters because the MLP holds most of a transformer's parameters. Attention-only adapters underperform even at matched trainable-parameter count.
solid answer
~50 sThe original formulation adapted only the query and value projections, and that choice stuck around as a default long after the evidence turned against it. Current guidance is to adapt all linear layers: query, key, value and output projections in attention, plus the feed-forward up, gate and down projections — and, in a sparse mixture-of-experts model, the expert matrices too. The reason is where the parameters live. The feed-forward block holds roughly two-thirds of a dense transformer's weights, and a far larger share in an MoE model, so an attention-only adapter leaves most of the model's transformation capacity untouched no matter what rank you pick. The comparison that settles it is at matched trainable-parameter count: spreading the same budget across all linear layers beats concentrating it in attention at higher rank. Embeddings and the output head are usually left alone unless the vocabulary changes; normalization layers and biases are not linear projections and are typically excluded.
go deeper
Know that a LoRA config chooses which weight matrices get adapters, and that the modern default is to cover all the linear projections rather than just the attention ones.
Be able to name the candidate matrices — query, key, value, output, and the feed-forward up, gate and down — and explain that the feed-forward block holds most of the block's parameters.
Make the matched-parameter-count argument explicitly, and reason about mixture-of-experts models where the parameter mass sits in the experts. Say which modules you deliberately exclude and why.
Own the inherited-default problem: recipes copied from years-old references carry choices that have since been overturned, and part of the job is deciding which defaults get re-tested before a training budget is committed.
## The question behind the knob A LoRA configuration has to say *which* weight matrices get adapters. Every transformer block contains several candidate matrices: the query, key, value and output projections of attention, and the feed-forward block's projections — typically an up projection, a gate projection and a down projection in a gated design. Around the blocks sit the token embedding table, the output head, and normalization layers. The original low-rank adaptation work adapted the query and value projections only. That was a reasonable exploration choice at the time, and it propagated into defaults and tutorials for years. It is no longer the right default. ## Where the parameters actually are In a dense transformer, attention's four projections are each roughly hidden-size squared. The feed-forward block's projections are hidden-size by intermediate-size, and the intermediate size is conventionally two to four times the hidden size. Summed across the block, the feed-forward matrices hold roughly two-thirds of the layer's parameters, and attention roughly one-third. That imbalance is the core of the argument. An adapter can only modify the transformation performed by a matrix it is attached to. Restricting adapters to attention means the majority of the model's per-layer transformation capacity is untouched, permanently, at any rank. Rank controls how much you can change a matrix; target modules control *which* matrices you are allowed to change at all — and no amount of the former substitutes for the latter. ## The matched-budget comparison The naive objection is that you can compensate by raising rank on the attention matrices, buying the same number of trainable parameters. The empirical answer is that you cannot: at matched trainable-parameter count, spreading the budget across all linear layers outperforms concentrating it in attention at higher rank. The feed-forward block is where much of a transformer's learned feature transformation and factual association lives, and adapting it is qualitatively different from turning the attention dials harder. This is the single most consequential correction to older LoRA folklore, and it has been absorbed into mainstream adapter tooling as an explicit "all linear layers" target setting rather than a hand-written module list. ## Mixture-of-experts models make it starker In a sparse mixture-of-experts model, the feed-forward block is replaced by a bank of expert feed-forward networks with a router selecting a small number per token. The parameter mass shifts even further away from attention: the experts hold the overwhelming majority of the model's weights. An attention-only adapter on such a model is adapting a small minority of the parameters and leaving the part of the network that actually does the specialized transformation entirely frozen. A common misunderstanding here is that because only a few experts are active per token, adapting all of them is wasteful. Over a training batch the router distributes tokens across the full expert bank, so the adapters do receive gradient; sparsity per token is not sparsity per batch. The practical caveat is engineering, not principle — the expert matrices are numerous, so adapting them all inflates adapter size and can complicate the training stack. ## What to leave alone **Embeddings and the output head.** These are usually excluded. They are enormous — vocabulary size times hidden size — so adapting them inflates the adapter substantially, and for most tasks the token representations do not need to move. The exception is when you add tokens: new special tokens, a new chat template's control tokens, or a domain vocabulary extension. Then the embedding rows for the new tokens must be trainable or the model has no way to learn what they mean. **Normalization layers and biases.** These are not linear projections, so LoRA's factorization does not apply to them. Some recipes train them directly as full parameters alongside the adapters; that is a separate decision from the adapter target list and adds only a negligible parameter count. **The router in an MoE model.** Adapting the routing network is generally avoided, since destabilizing routing during a short fine-tune can shift load across experts in ways that are hard to evaluate. ## The practical rule Use the all-linear target as the default and deviate only for a stated reason: a memory ceiling that forces you to drop the largest matrices, a vocabulary change that forces you to add the embedding rows, or an ablation you are deliberately running. If you inherit a configuration that lists only query and value projections, treat that as a legacy default to be re-tested, not as a considered choice.
- Why not just raise the rank on the attention projections instead of adapting the MLP?Because the comparison has been run at matched trainable-parameter count, and spreading the same budget across all linear layers wins. Rank controls how far you can move a matrix; target modules control which matrices can move at all. The feed-forward block carries most of the parameters and much of the learned feature transformation, so leaving it frozen caps what the adapter can express no matter how much capacity you pile onto attention.
- When would you include the embedding table and output head in the target list?When the token inventory changes — adding special tokens, control tokens for a new chat template, or a domain vocabulary extension. New embedding rows start as noise and must be trainable for the model to learn what those tokens mean. Otherwise leave them out: they are among the largest matrices in the model, and adapting them inflates the adapter checkpoint for little benefit on most tasks.
- Does per-token expert sparsity in an MoE model mean most expert adapters never train?No. The router sends different tokens to different experts, so across a training batch the whole expert bank receives gradient even though only a few experts fire per token. The real caveat is engineering rather than learning: expert matrices are numerous, so adapting all of them enlarges the adapter and can strain the training stack. Skipping them entirely leaves most of the model's parameter mass frozen.
saying these in an interview costs you the question
- Keeping query-and-value-only because the original paper used it
- Assuming higher rank compensates for narrow targeting
- Believing attention is where most transformer parameters live
- Adapting embeddings by default without a vocabulary change
- Skipping expert matrices in a sparse model to save space