skip to content

How do NTK-aware scaling, YaRN and LongRoPE rescale RoPE frequencies differently?

level: seniorimportance: should knowfreq 42%

answer

  1. slow dimensions never finished a cycle
  2. uniform squeeze crowds local detail
  3. raise the base, not the index
  4. classify dimensions by wavelength band
  5. search the factors per dimension

basics

~20 s

All three keep far-out positions inside the angle range a model trained on. NTK-aware scaling raises the rotary base so slow dimensions absorb most of the compression; YaRN treats dimensions differently by wavelength band and adds an attention-temperature correction; LongRoPE searches a per-dimension factor instead of deriving one.

solid answer

~60 s

Start from why RoPE breaks past its trained length: the low-frequency dimensions have wavelengths longer than the training window, so they only ever swept part of a cycle, and beyond that window they present unseen angles. **Plain position interpolation** divides every index by a constant so the maximum lands back in range — uniform, and it crowds the fast dimensions that carried local distinctions. **NTK-aware scaling** instead increases the rotary base, which barely touches the fastest dimensions and applies most of the compression to the slowest ones, trading a little long-range precision for preserved local resolution; a dynamic variant adjusts the factor with the current sequence length. **YaRN** makes that explicit: dimensions whose wavelength is far shorter than the training context are left alone, dimensions whose wavelength exceeds it are fully interpolated, and a ramp blends the middle — plus a temperature correction on attention logits to offset the entropy shift from a longer sequence. **LongRoPE** abandons closed form entirely, running an evolutionary search for a separate rescale factor per rotary dimension against an evaluation objective, and keeping the first few positions unscaled. Formula, band-wise rule, searched vector — increasing generality at increasing cost.

go deeper

for a junior

Know the shared goal: all of these compress far-out positions back into the range of rotation angles the model actually saw during training.

for a middle

Explain why low-frequency dimensions are the out-of-distribution ones, and state the difference between dividing the position index and raising the rotary base.

for a senior

Compare the three by how non-uniformly they treat the spectrum, name YaRN's separate attention-temperature term, and be explicit that every scheme trades local resolution for range.

for a principal

Own the judgment about which resolution you can afford to lose for your workload, and be clear that a stretched window is not the same asset as a natively trained one when you plan capability.

## Why RoPE needs rescaling at all RoPE gives each dimension pair a wavelength, spread geometrically from a few tokens to far longer than any training sequence. During training at length L, a fast dimension completes many full cycles and the network sees every angle it can take. A slow dimension whose wavelength exceeds L never completes a cycle: it only ever sweeps an arc. Push the model to position 4L and those slow dimensions occupy angles no weight ever encountered. The geometry is still valid; the *input distribution* is not. This is the mechanism behind the familiar observation that a model past its trained length keeps producing fluent text while its accuracy on anything requiring precise reference collapses. Every scheme below is an answer to the same question: how do you map far-out positions back onto angles the model already understands, and which resolution do you sacrifice doing it? ## Position interpolation — the uniform baseline The simplest fix divides every position index by a scale factor s, so a model trained to L can accept sL tokens with all angles landing inside the familiar arc. It is uniform: every dimension is compressed identically. The cost is concentrated where you can least afford it. The fast dimensions were the ones distinguishing "one token back" from "three tokens back"; after dividing by 8, those distinctions are squeezed into an eighth of the angular space they had. Local resolution — the model's sharpest positional signal — degrades to buy long-range range. ## NTK-aware scaling — spend the compression where it is cheap The NTK-aware family reframes the problem: rather than scaling positions, scale the frequency schedule itself by increasing the rotary base. Because frequencies are geometric in the dimension index, raising the base leaves the fastest dimensions nearly unchanged while stretching the slowest ones substantially. High-frequency detail is preserved and the interpolation burden is pushed onto the low-frequency dimensions, which were the out-of-distribution ones to begin with. A dynamic variant computes the scale from the *current* sequence length rather than a fixed maximum, so short prompts are processed with the original schedule and only long ones are rescaled. That avoids paying an extension penalty on requests that never needed it. ## YaRN — make the band split explicit YaRN turns the NTK intuition into an explicit rule, sometimes called NTK-by-parts. For each dimension it computes how its wavelength compares to the training context: a wavelength far shorter than the context means the model saw the full cycle and can safely be left alone, so that dimension is *not* interpolated. A wavelength exceeding the context means the dimension is genuinely out of distribution past L, so it is interpolated fully. In between, a ramp blends the two treatments so there is no discontinuity across the spectrum. YaRN adds a second, separate ingredient: a temperature correction applied to the attention logits, scaling the extension factor logarithmically. The motivation is that spreading attention over a much longer sequence changes the entropy of the softmax distribution, and correcting for that recovers accuracy independently of the frequency treatment. Practically, YaRN is reported to need far less continued training than uniform interpolation to reach the same quality — that efficiency, not just the final score, is much of why it was adopted. ## LongRoPE — search instead of derive LongRoPE observes that no closed-form rule is guaranteed to be optimal per dimension, and replaces derivation with search: an evolutionary algorithm looks for a separate rescale factor for every rotary dimension, scored against an evaluation objective rather than a formula. It also treats the first handful of positions specially, leaving them unscaled, since the earliest tokens play an outsized role in attention. Combined with progressive, staged extension it reached window sizes far beyond what uniform interpolation supported. Later work in this line replaced perplexity-style objectives with retrieval-driven ones, because perplexity turned out to be a weak proxy for whether long-range reference actually works. The tradeoff is plain: search costs compute and produces a vector of numbers with no interpretable rule, but it is not bound by anyone's assumption about which dimensions matter. ## How to reason about the choice The axis running through all four is *how non-uniformly you treat the frequency spectrum*: uniform (interpolation), formula-driven non-uniform (NTK-aware), band-wise non-uniform with an attention correction (YaRN), and freely searched per dimension (LongRoPE). Generality rises, so does the cost of arriving at the coefficients. In mid-2026 these remain the standard vocabulary in open-weight stacks; frontier models increasingly bake long-context handling into a dedicated mid-training stage instead of retrofitting a scaling rule. One caveat to state out loud: none of these are free. Every one trades some positional resolution for range, and a model whose window was stretched by a scaling rule is not equivalent to a model trained natively at that length.

  • Why are the low-frequency rotary dimensions the ones that go out of distribution first?
    Their wavelengths are longer than the training window, so across training they only ever swept part of a single rotation. Every angle they took was drawn from that arc. Past the training length they occupy angles no weight has seen. Fast dimensions, by contrast, cycled many times within the window, so their full angular range is already familiar and they extrapolate comparatively well.
  • What does YaRN's attention-temperature correction address that frequency rescaling does not?
    Softmax entropy. Attending over a much longer sequence spreads probability mass differently than the model was trained for, which degrades accuracy independently of whether the angles are in range. Scaling the attention logits by a factor that grows with the extension counteracts that shift. It is a separate lever from the frequency treatment, which is why YaRN is described as two ingredients rather than one.
  • When would you prefer a searched per-dimension factor over a formula?
    When you can afford the search compute and you have an evaluation objective you trust more than the formula's assumptions — especially at large extension ratios, where formula-derived schedules leave measurable quality on the table. The catch is picking the objective: perplexity is a weak proxy for long-range reference, so later work moved to retrieval-style objectives instead.
  • Does any of this make a stretched model equivalent to one trained natively at that length?
    No, and claiming so is the common overreach. Every scheme trades positional resolution for range, and a stretched model typically shows weaker precision at both the fine-grained local scale and the far end of its new window than a model that saw those lengths in training. Treat a scaling rule as a way to reach a length, not as a substitute for training at it.

Uniform interpolation is resizing a photograph — everything shrinks, including the fine detail you cared about. The later schemes are content-aware resizing: compress the empty sky, leave the faces alone.

saying these in an interview costs you the question

  • Says the schemes are interchangeable and differ only in constants
  • Claims uniform position interpolation costs nothing
  • Thinks high-frequency dimensions are the ones out of distribution
  • Describes YaRN as pure frequency scaling with no attention correction
  • Assumes a rescaled model matches one natively trained at that length

context