How does RoPE make attention depend on relative offset using absolute positions?
answer
- rotate, don't add
- pairs of dimensions as 2-D planes
- angle proportional to absolute position
- rotations compose, absolutes cancel
- queries and keys, every layer
basics
~20 sRoPE rotates each consecutive pair of dimensions in a query and a key by an angle proportional to that token's absolute position. Rotations compose, so the resulting dot product depends only on the difference between the two positions.
solid answer
~50 sRotary Position Embedding treats a query or key vector as a stack of 2-D pairs and rotates pair *i* by an angle of position × θᵢ, where the θᵢ are spaced geometrically from fast to slow across the dimensions. Nothing is added to the token embedding; the rotation is applied to queries and keys after projection, in every layer, and never to values. The payoff is an algebraic identity: rotating a query by mθ and a key by nθ and taking their dot product is the same as rotating one of them by (n − m)θ — the score depends on the offset alone. Concretely, the pair at positions 5 and 8 produces the same score contribution as the same vectors at positions 500 and 503. That gives relative behaviour from a purely absolute, parameter-free operation, applied per layer rather than surviving from the embedding, and it preserves vector norms because rotations are orthogonal. It is the mid-2026 baseline in essentially every decoder-only LLM.
code
python · 17 linesimport numpy as np
def rope(x, pos, base=10000.0):
d = x.shape[-1]
i = np.arange(d // 2)
theta = pos / (base ** (2 * i / d))
cos, sin = np.cos(theta), np.sin(theta)
even, odd = x[0::2], x[1::2]
out = np.empty_like(x)
out[0::2] = even * cos - odd * sin
out[1::2] = even * sin + odd * cos
return out
rng = np.random.default_rng(0)
q, k = rng.normal(size=64), rng.normal(size=64)
print(rope(q, 5) @ rope(k, 8)) # offset 3
print(rope(q, 500) @ rope(k, 503)) # same offset, same scorego deeper
Recall the one-line summary: RoPE rotates query and key vectors by an angle that grows with position, so attention ends up depending on how far apart two tokens are.
Be able to walk the geometry — dimension pairs as 2-D planes, geometrically spaced frequencies, rotations composing so absolute positions cancel — and to state that values are untouched.
Explain the frequency spectrum as a multi-resolution ruler, why RoPE displaced absolute schemes and ALiBi, and why relative scores still do not buy free extrapolation past the trained range.
Own the architectural consequence: the frequency schedule you pick at pretraining time determines how far the model can later be stretched and what a stretch will cost you in local resolution.
## The construction After a head's query and key projections, RoPE views each vector of dimension *d* as *d/2* independent 2-D planes, formed from consecutive dimension pairs. For a token at absolute position *m*, plane *i* is rotated by the angle *m·θᵢ*, where the frequencies are geometrically spaced — conventionally θᵢ = base^(−2i/d) with the base commonly 10,000. Plane 0 rotates fast (a short wavelength, a few tokens per cycle); the last planes rotate extremely slowly, with wavelengths far longer than a typical sequence. The operation is a rotation, hence orthogonal: it changes direction, never magnitude. It adds **no parameters**, since the angles are a fixed function of the index and the frequency schedule. It is applied to queries and keys only. Values are left alone — RoPE modifies *how strongly* tokens attend, not *what* they transmit. ## Why relative offset falls out Rotations in a plane compose by adding angles, and the dot product of two rotated vectors equals the dot product of one rotated by the *difference*. If R(m) is the block-diagonal rotation for position *m*, then ⟨R(m)q, R(n)k⟩ = ⟨q, R(n − m)k⟩. The absolute positions cancel; only the offset survives. This is the whole trick. Each token is tagged absolutely — you can rotate a key once when you cache it, without knowing which query will later look at it — yet every score the model computes is a function of distance. An absolute-embedding scheme has to *learn* that positions 5 and 8 should behave like 500 and 503. RoPE makes it an identity. Work the example. Take a random query and key vector. Rotate the query for position 5 and the key for position 8; record the dot product. Now rotate the same query for position 500 and the same key for position 503. The two numbers agree to floating-point precision. Nothing was trained to make that true. ## The frequency spectrum, and what it buys The geometric spread of θᵢ is not decoration. Fast planes complete many cycles inside a short window, so they encode fine local distinctions — adjacent versus two tokens back. Slow planes barely move across thousands of tokens, so they encode coarse long-range separation. The model gets a multi-resolution ruler in a single mechanism, and different heads learn to lean on different bands. A side effect of summing many rotating planes is that, for generic vectors, the score contribution tends to oscillate and decay as the offset grows. This is often described as a helpful inductive bias toward local attention, and it is also the reason some 2026 architectures deliberately leave part of the representation unrotated, so that a subspace can match content at any distance without that decay. ## Why it displaced the alternatives Absolute sinusoidal and learned encodings inject position once, at the embedding, and it must then survive dozens of residual updates while every layer re-derives distance from two absolute tags. RoPE re-injects position at every layer, costs no parameters, and gives relative dependence structurally. ALiBi, which subtracts a penalty proportional to distance from the attention logits, was the other serious relative answer and extrapolated better out of the box, but it imposes a fixed monotone locality bias that limits precise long-range retrieval; it belongs to an earlier model generation. As of mid-2026, RoPE — with RMSNorm, SwiGLU and grouped-query attention — is the conservative baseline. ## The limit, stated honestly RoPE does not extrapolate for free. A model trained to length L only ever observes the slow planes across a small arc of their cycle; at positions well past L, those planes present angles the network has never seen, and quality degrades even though generation continues fluently. That is why a family of frequency-rescaling schemes exists at all. Saying "RoPE is relative, so it generalizes to any length" is the classic overclaim — relative dependence is a property of the *scores*, not a guarantee about *unseen angles*. ## What to say in an interview Lead with the geometry: pairs of dimensions, rotated by an angle proportional to position, at a spread of frequencies. Then the identity: rotations compose, so the dot product depends only on the offset. Then the practical facts: queries and keys only, every layer, no parameters, norm-preserving. Then the honest limit: it still has a trained range.
- Why is RoPE applied to queries and keys but not to values?Because the point is to make the attention *score* a function of distance, and the score comes only from the query-key dot product. Values carry the content that gets mixed; rotating them would distort the transmitted information without adding any positional structure to the weighting. Keeping values untouched also means a cached key can be rotated once at write time while its value stays as-is.
- How does RoPE compare with ALiBi, which subtracts a penalty proportional to distance?Both make attention relative without adding parameters. ALiBi bakes in a fixed monotone decay with distance, which extrapolated well past the training length but caps precise long-range retrieval — a distant token is always penalized regardless of relevance. RoPE lets heads learn distance-dependent behaviour per band rather than imposing one. RoPE became the standard; ALiBi is a previous-generation choice.
- If the scores depend only on offsets, why does a RoPE model still degrade past its training length?Because the low-frequency planes have wavelengths longer than the training window, so during training they only ever swept a fraction of a full rotation. Positions well beyond that window place them at angles the weights have never seen — valid geometry, unfamiliar input. Output stays fluent while precise reference degrades, which is exactly why frequency-rescaling schemes were invented.
- Does RoPE add parameters or change the vector norms?Neither. The angles come from the token index and a fixed geometric frequency schedule, so there is nothing to train, and a rotation is orthogonal so it preserves length exactly. That is part of why it was easy to adopt: it slots into an existing attention implementation as a transformation on queries and keys with no change to parameter count or normalization behaviour.
Each dimension pair is a clock hand ticking at its own speed. Comparing two tokens compares the angle between their hands, and that angle depends only on how much time elapsed between them — not on what the clock read when either one started.
saying these in an interview costs you the question
- Says RoPE is added to the token embeddings like sinusoidal encoding
- Claims RoPE is applied to values as well as queries and keys
- Thinks relative dependence means unlimited length generalization
- Says RoPE introduces learned position parameters
- Describes RoPE as scaling vector magnitudes rather than rotating them