skip to content

How does RoPE make attention depend on relative offset using absolute positions?

level: middleimportance: must knowfreq 68%

answer

  1. rotate, don't add
  2. pairs of dimensions as 2-D planes
  3. angle proportional to absolute position
  4. rotations compose, absolutes cancel
  5. queries and keys, every layer

basics

~20 s

RoPE rotates each consecutive pair of dimensions in a query and a key by an angle proportional to that token's absolute position. Rotations compose, so the resulting dot product depends only on the difference between the two positions.

solid answer

~50 s

Rotary Position Embedding treats a query or key vector as a stack of 2-D pairs and rotates pair *i* by an angle of position × θᵢ, where the θᵢ are spaced geometrically from fast to slow across the dimensions. Nothing is added to the token embedding; the rotation is applied to queries and keys after projection, in every layer, and never to values. The payoff is an algebraic identity: rotating a query by mθ and a key by nθ and taking their dot product is the same as rotating one of them by (n − m)θ — the score depends on the offset alone. Concretely, the pair at positions 5 and 8 produces the same score contribution as the same vectors at positions 500 and 503. That gives relative behaviour from a purely absolute, parameter-free operation, applied per layer rather than surviving from the embedding, and it preserves vector norms because rotations are orthogonal. It is the mid-2026 baseline in essentially every decoder-only LLM.

code

python · 17 lines
python
import numpy as np

def rope(x, pos, base=10000.0):
    d = x.shape[-1]
    i = np.arange(d // 2)
    theta = pos / (base ** (2 * i / d))
    cos, sin = np.cos(theta), np.sin(theta)
    even, odd = x[0::2], x[1::2]
    out = np.empty_like(x)
    out[0::2] = even * cos - odd * sin
    out[1::2] = even * sin + odd * cos
    return out

rng = np.random.default_rng(0)
q, k = rng.normal(size=64), rng.normal(size=64)
print(rope(q, 5) @ rope(k, 8))       # offset 3
print(rope(q, 500) @ rope(k, 503))   # same offset, same score

go deeper

for a junior

Recall the one-line summary: RoPE rotates query and key vectors by an angle that grows with position, so attention ends up depending on how far apart two tokens are.

for a middle

Be able to walk the geometry — dimension pairs as 2-D planes, geometrically spaced frequencies, rotations composing so absolute positions cancel — and to state that values are untouched.

for a senior

Explain the frequency spectrum as a multi-resolution ruler, why RoPE displaced absolute schemes and ALiBi, and why relative scores still do not buy free extrapolation past the trained range.

for a principal

Own the architectural consequence: the frequency schedule you pick at pretraining time determines how far the model can later be stretched and what a stretch will cost you in local resolution.

## The construction After a head's query and key projections, RoPE views each vector of dimension *d* as *d/2* independent 2-D planes, formed from consecutive dimension pairs. For a token at absolute position *m*, plane *i* is rotated by the angle *m·θᵢ*, where the frequencies are geometrically spaced — conventionally θᵢ = base^(−2i/d) with the base commonly 10,000. Plane 0 rotates fast (a short wavelength, a few tokens per cycle); the last planes rotate extremely slowly, with wavelengths far longer than a typical sequence. The operation is a rotation, hence orthogonal: it changes direction, never magnitude. It adds **no parameters**, since the angles are a fixed function of the index and the frequency schedule. It is applied to queries and keys only. Values are left alone — RoPE modifies *how strongly* tokens attend, not *what* they transmit. ## Why relative offset falls out Rotations in a plane compose by adding angles, and the dot product of two rotated vectors equals the dot product of one rotated by the *difference*. If R(m) is the block-diagonal rotation for position *m*, then ⟨R(m)q, R(n)k⟩ = ⟨q, R(n − m)k⟩. The absolute positions cancel; only the offset survives. This is the whole trick. Each token is tagged absolutely — you can rotate a key once when you cache it, without knowing which query will later look at it — yet every score the model computes is a function of distance. An absolute-embedding scheme has to *learn* that positions 5 and 8 should behave like 500 and 503. RoPE makes it an identity. Work the example. Take a random query and key vector. Rotate the query for position 5 and the key for position 8; record the dot product. Now rotate the same query for position 500 and the same key for position 503. The two numbers agree to floating-point precision. Nothing was trained to make that true. ## The frequency spectrum, and what it buys The geometric spread of θᵢ is not decoration. Fast planes complete many cycles inside a short window, so they encode fine local distinctions — adjacent versus two tokens back. Slow planes barely move across thousands of tokens, so they encode coarse long-range separation. The model gets a multi-resolution ruler in a single mechanism, and different heads learn to lean on different bands. A side effect of summing many rotating planes is that, for generic vectors, the score contribution tends to oscillate and decay as the offset grows. This is often described as a helpful inductive bias toward local attention, and it is also the reason some 2026 architectures deliberately leave part of the representation unrotated, so that a subspace can match content at any distance without that decay. ## Why it displaced the alternatives Absolute sinusoidal and learned encodings inject position once, at the embedding, and it must then survive dozens of residual updates while every layer re-derives distance from two absolute tags. RoPE re-injects position at every layer, costs no parameters, and gives relative dependence structurally. ALiBi, which subtracts a penalty proportional to distance from the attention logits, was the other serious relative answer and extrapolated better out of the box, but it imposes a fixed monotone locality bias that limits precise long-range retrieval; it belongs to an earlier model generation. As of mid-2026, RoPE — with RMSNorm, SwiGLU and grouped-query attention — is the conservative baseline. ## The limit, stated honestly RoPE does not extrapolate for free. A model trained to length L only ever observes the slow planes across a small arc of their cycle; at positions well past L, those planes present angles the network has never seen, and quality degrades even though generation continues fluently. That is why a family of frequency-rescaling schemes exists at all. Saying "RoPE is relative, so it generalizes to any length" is the classic overclaim — relative dependence is a property of the *scores*, not a guarantee about *unseen angles*. ## What to say in an interview Lead with the geometry: pairs of dimensions, rotated by an angle proportional to position, at a spread of frequencies. Then the identity: rotations compose, so the dot product depends only on the offset. Then the practical facts: queries and keys only, every layer, no parameters, norm-preserving. Then the honest limit: it still has a trained range.

  • Why is RoPE applied to queries and keys but not to values?
    Because the point is to make the attention *score* a function of distance, and the score comes only from the query-key dot product. Values carry the content that gets mixed; rotating them would distort the transmitted information without adding any positional structure to the weighting. Keeping values untouched also means a cached key can be rotated once at write time while its value stays as-is.
  • How does RoPE compare with ALiBi, which subtracts a penalty proportional to distance?
    Both make attention relative without adding parameters. ALiBi bakes in a fixed monotone decay with distance, which extrapolated well past the training length but caps precise long-range retrieval — a distant token is always penalized regardless of relevance. RoPE lets heads learn distance-dependent behaviour per band rather than imposing one. RoPE became the standard; ALiBi is a previous-generation choice.
  • If the scores depend only on offsets, why does a RoPE model still degrade past its training length?
    Because the low-frequency planes have wavelengths longer than the training window, so during training they only ever swept a fraction of a full rotation. Positions well beyond that window place them at angles the weights have never seen — valid geometry, unfamiliar input. Output stays fluent while precise reference degrades, which is exactly why frequency-rescaling schemes were invented.
  • Does RoPE add parameters or change the vector norms?
    Neither. The angles come from the token index and a fixed geometric frequency schedule, so there is nothing to train, and a rotation is orthogonal so it preserves length exactly. That is part of why it was easy to adopt: it slots into an existing attention implementation as a transformation on queries and keys with no change to parameter count or normalization behaviour.

Each dimension pair is a clock hand ticking at its own speed. Comparing two tokens compares the angle between their hands, and that angle depends only on how much time elapsed between them — not on what the clock read when either one started.

saying these in an interview costs you the question

  • Says RoPE is added to the token embeddings like sinusoidal encoding
  • Claims RoPE is applied to values as well as queries and keys
  • Thinks relative dependence means unlimited length generalization
  • Says RoPE introduces learned position parameters
  • Describes RoPE as scaling vector magnitudes rather than rotating them

context