Why would a long-context model leave some head dimensions unrotated by RoPE?
answer
- rotation is not required everywhere
- some dimensions carry no position
- distance-blind content matching
- retrieval heads fight the locality prior
- split ratio is contested, not settled
basics
~20 sRotating a dimension makes its contribution to the attention score oscillate and fall away with distance. Leaving part of the representation unrotated gives the model a position-agnostic subspace that can match purely on content, at any distance.
solid answer
~50 sApplying rotation to every dimension is a choice, not a requirement — early rotary models rotated only a quarter of each head's dimensions, and by mid-2026 partial-RoPE and NoPE subspace splits are a live design question in frontier stacks. The argument is that rotation couples matching to distance: summed across many rotating planes, a query-key score for generic vectors tends to oscillate and decay as the offset grows, which is a useful locality prior for language but a liability for a head whose job is to find one relevant fact 400,000 tokens back. Reserving a subspace with no rotation gives such a head a channel where similarity is unaffected by how far apart the tokens are. The same idea appears at layer granularity, interleaving position-agnostic layers with rotary ones. The cost is real: strip too much rotation and the model loses precise local ordering, which the causal mask alone supplies only weakly. There is no settled recipe — the split ratio is an empirical decision, and stacks disagree.
go deeper
Know that rotary encoding does not have to cover every dimension — models have shipped rotating only a fraction of each head, with the rest left position-free.
Explain why rotation makes a query-key score vary and fade with distance, and why that is a helpful locality prior for most heads but an obstacle for a retrieval head.
Distinguish the variants — dimension split, layer split, full NoPE — and name the cost: weaker exact local ordering, which the causal mask covers only weakly.
Own that this is unsettled ground. Frame it as a hypothesis to test with two separate evaluations, and resist quoting a split ratio as best practice when the field has not converged.
## Partial rotation is older than the current debate It is easy to assume RoPE applies to a whole head. It never had to. Some of the first models to adopt rotary encoding rotated only about a quarter of each head's dimensions and left the remainder alone, and it worked. What has changed by mid-2026 is that the split stopped being an implementation detail and became an explicit architectural lever, discussed alongside NoPE — running with no positional encoding whatsoever and relying on the causal mask's implicit signal. ## The mechanism: rotation entangles matching with distance A rotated dimension pair contributes a term to the query-key dot product that varies with the angle between the two tokens' rotations — that is, with their offset. Sum many such planes across the fast and slow frequency bands and, for generic vectors, the aggregate contribution tends to oscillate and attenuate as the gap grows. For most of language this is a desirable prior: nearby tokens usually matter more, and the model gets that for free. But it is a prior imposed on *every* head in *every* layer. Some heads exist to do long-range retrieval — locate the single stack frame, clause, or record that matters, wherever it is. For those, distance-dependent attenuation is noise fighting the objective, and the head must spend capacity learning to overcome its own positional machinery. An unrotated subspace removes that fight. In those dimensions the dot product is a pure content match: a query and a key that agree score the same whether they are 12 tokens apart or 120,000. The model gets both behaviours in one head — locality where rotation applies, distance-blind matching where it does not — instead of one compromise. ## The variants you will hear named **Dimension splits.** Rotate a fraction of each head's dimensions; leave the rest position-agnostic. The fraction is a hyperparameter with no consensus value. **Layer splits.** Keep some layers fully rotary and make others position-free, interleaved through the stack, so early layers can establish local structure while designated layers do distance-blind retrieval. **Full NoPE.** No positional encoding at all, relying on the causal mask giving each position a differently sized visible prefix. Research showed decoder-only models can learn order this way and can generalize to longer inputs better than encoded ones on some tasks. It is fragile where exact local ordering matters, so pure NoPE at frontier scale remains contested. Separately, some compressed-KV attention designs are *forced* into a split for cache reasons: when keys are reconstructed from a compressed latent, rotation cannot be applied to the compressed form, so a small dedicated set of dimensions carries the position signal and the rest carries content. That is a mechanical constraint, not the retrieval argument above — worth distinguishing, because they arrive at superficially similar architectures for entirely different reasons. ## The costs, stated honestly Withdrawing rotation withdraws the model's sharpest signal for exact local ordering. The causal mask supplies order only weakly and indirectly, and tasks that hinge on precise adjacency — parsing nested structure, exact sequence reproduction, code with strict token order — are where an over-aggressive split shows up. It also complicates reasoning about the model: two subspaces with different positional semantics is harder to analyse and to debug than one uniform rule. And there is no formula. The split ratio, the layer pattern, and whether to do it at all are empirical calls made per architecture, and mid-2026 stacks disagree with each other. Anyone presenting a specific number as settled best practice is overstating the state of the field. ## How you would actually decide Treat it as a hypothesis with two competing evaluations, not one. Measure a long-range retrieval task where the target sits far from the query, and separately measure a task that hinges on exact local order. A good split improves the first without moving the second; a bad one buys retrieval by quietly losing precision that only shows up on structured output. Averaging both into a single score hides exactly the tradeoff you are trying to observe. ## What a principal-level answer sounds like Name the mechanism (rotation couples similarity to distance), name the benefit (a distance-blind subspace for retrieval heads), name the cost (weaker exact local ordering), name the variants (dimension split, layer split, full NoPE), and say plainly that the field has not converged. Confidence about the right ratio is the tell that someone is reciting rather than reasoning.
- If the causal mask already leaks position, why not drop RoPE entirely?Some architectures do, and NoPE research shows decoder-only models can learn order from the mask's varying prefix size, sometimes generalizing better to longer inputs. But that signal is weak and indirect, and tasks needing exact local ordering — nested structure, strict token sequences — are where it shows its limits. At frontier scale, pure NoPE is a live but unsettled bet rather than an established default.
- How would you evaluate whether a partial-rotation split actually helped?With two evaluations, not one. Run a long-range task where the needed information sits far from the query, and separately run a task that depends on exact local order. A good split lifts the first without degrading the second. Collapsing both into one average is precisely how you miss the regression — the mechanism trades one capability for another, so the metric has to separate them.
- Some compressed-KV attention designs also separate position-carrying dimensions. Is that the same idea?No — same shape, different cause. When keys are reconstructed from a compressed latent, rotation cannot be applied to the compressed form, so a small dedicated set of dimensions carries position while the rest carries content. That is a mechanical constraint from the cache design, not a deliberate choice to give retrieval heads a distance-blind subspace. Worth distinguishing when reading an architecture description.
saying these in an interview costs you the question
- Assumes RoPE must be applied to every dimension of every head
- Says an unrotated subspace loses no capability at all
- Claims the causal mask fully substitutes for a positional encoding
- Presents one split ratio as the settled industry best practice
- Confuses cache-driven position/content separation with the retrieval argument