skip to content

How does Llama 3.1 reach a 128K context when Llama 3 was trained at 8K?

level: seniorimportance: should knowfreq 40%

answer

  1. Positions past training are unseen phases
  2. Not every frequency is treated alike
  3. Fast dimensions untouched, slow ones divided
  4. Config records factor 8 over 8192
  5. Scaling plus continued long-context pretraining

basics

~20 s

Llama 3.1 rescales the rotary position embeddings so positions past the original 8,192 fall back into a range the model has seen, then continues pretraining on progressively longer sequences up to 131,072 tokens. The config records this as a rope_scaling block with rope_type llama3 and factor 8.

solid answer

~50 s

Llama encodes position with **rotary position embeddings (RoPE)**: query and key vectors are rotated by an angle proportional to the token's position, with each embedding dimension rotating at its own frequency. Feed a position far beyond the training range and the low-frequency dimensions land at phases the model never saw, and quality collapses. Llama 3.1 fixes that with frequency-dependent scaling rather than uniform interpolation: long-wavelength dimensions are divided down by a factor of 8, short-wavelength ones are left alone so local ordering stays sharp, and a smooth ramp covers the middle. In the Hugging Face config that appears as `rope_scaling` with `rope_type: "llama3"`, `factor: 8.0`, `low_freq_factor: 1.0`, `high_freq_factor: 4.0` and `original_max_position_embeddings: 8192`. Scaling alone is not enough — Meta then continued pretraining on long-sequence data in stages up to 128K. Operationally, remember that the window is advertised, not free: prefill is quadratic and the KV cache grows linearly, so a full window is expensive long before it is inaccurate.

code

json · 11 lines
json
{
  "max_position_embeddings": 131072,
  "rope_theta": 500000.0,
  "rope_scaling": {
    "rope_type": "llama3",
    "factor": 8.0,
    "low_freq_factor": 1.0,
    "high_freq_factor": 4.0,
    "original_max_position_embeddings": 8192
  }
}

go deeper

for a junior

Know that Llama encodes position by rotating query and key vectors, and that reaching a longer window required rescaling those rotations plus more training — not simply a config edit.

for a middle

Explain why naive extrapolation fails: slow-rotating dimensions never complete a cycle within the original window, so long positions land at phases the model never saw during training.

for a senior

Demonstrate the operational side — that the advertised window costs quadratic prefill and linear KV cache, that quality thins toward the tail, and that you measure your task at several lengths before promising long-context behaviour.

for a principal

Own the policy on long context: when to pay for a full window versus retrieve-and-chunk, how context length interacts with batch size and cost per token, and what evidence must exist before a product feature depends on the tail of the window.

## What RoPE does, briefly Llama has no learned absolute position table. Instead it uses **rotary position embeddings**: before attention, the query and key vectors for a token at position `m` are rotated by an angle `m x theta_i`, where each pair of embedding dimensions `i` gets its own frequency `theta_i`. The frequencies form a geometric series from fast to slow, controlled by a base value (`rope_theta`, 500000.0 in Llama 3 versus 10000.0 in Llama 2). Because attention compares a query at position `m` with a key at position `n`, and the rotations combine, the dot product ends up depending on `m - n` — relative position — which is exactly what you want. ## Why the window does not just extend for free High-frequency dimensions complete many rotations within a few thousand tokens, so they generalise: the model has seen every phase. Low-frequency dimensions rotate slowly, and within an 8K training window they never complete even one cycle. Push a token to position 100,000 and those dimensions sit at angles the model has never encountered. The result is not graceful degradation; attention scores go haywire and output quality falls off a cliff. ## The three families of fix 1. **Position interpolation** — divide all positions by a factor so the new maximum maps back onto the trained range. Simple, but it squashes the high-frequency dimensions too, blurring the model's sense of nearby token order. 2. **NTK-aware / base scaling** — raise `rope_theta` instead, stretching wavelengths non-uniformly. 3. **Frequency-dependent scaling (what Llama 3.1 does)** — treat the dimensions differently by wavelength. Dimensions whose wavelength is short relative to the original context are untouched, so local ordering stays crisp. Dimensions whose wavelength exceeds the original context are divided by the scaling factor, so they stay inside seen phases. A smooth ramp interpolates in between rather than switching abruptly. In the Llama 3.1 `config.json` the block reads roughly: `rope_type` `"llama3"`, `factor` 8.0, `low_freq_factor` 1.0, `high_freq_factor` 4.0, `original_max_position_embeddings` 8192, alongside `max_position_embeddings` 131072. The factor of 8 is the 8K-to-64K stretch; the boundary parameters decide which dimensions count as low- versus high-frequency by comparing their wavelength against the original 8,192-token window. ## Scaling alone is not the whole answer This is the part candidates most often miss. Rescaling the frequencies makes long positions *representable*; it does not make the model *good* at using them. Meta paired the change with continued pretraining on long-context data, increasing the sequence length in stages up to 128K, so the model actually learns to retrieve and reason across the extended range. A model with rope scaling bolted on at load time and no long-context training will pass a trivial retrieval probe and fail at anything requiring synthesis over the window. ## What this means operationally - **The config is the contract.** If your serving stack is told the window is 8,192 — an old config, a truncated maximum length, a wrapper that clamps — you get the short window regardless of the checkpoint. Conversely, forcing a larger maximum than the config's scaling supports produces silently degraded output rather than an error. - **Cost is not flat across the window.** Prefill attention is quadratic in prompt length, so a 100K-token prompt is not twelve times a 8K one — it is far worse. The KV cache grows linearly, and on Llama 3.1 8B a full 128K context is roughly 16 GiB of cache per request in bf16, which usually collapses your batch size to one or two. - **Quality degrades before the limit does.** Advertised windows are the ceiling, not the recommended operating point. Measure your own task at 8K, 32K and 128K rather than assuming the last token is as usable as the first. - **Do not confuse rope scaling with other context-extension schemes.** YaRN, for instance, is a different published method used by other model families; naming it as Llama 3.1's mechanism is a factual slip an interviewer will catch. ## Answering well Start from why extrapolation fails — unseen rotation phases in the slow dimensions — then describe the frequency-selective fix and name the config fields, then add the two things that separate a real practitioner from a paper-reader: that continued long-context pretraining was required alongside it, and that the cost curve, not the config, is what actually limits how much of the window you use in production.

  • Why not just interpolate all positions uniformly instead of scaling by frequency?
    Uniform interpolation divides every position, including the fast-rotating dimensions that already generalise. Those dimensions carry local ordering, so squashing them blurs the model's sense of which nearby token came first, and short-context quality regresses. Frequency-dependent scaling leaves short-wavelength dimensions alone and only compresses the long-wavelength ones that would otherwise land at unseen phases, preserving local precision while extending reach.
  • If a model's config declares 128K, can I safely send 128K-token prompts in production?
    Declaring it and being good at it are different claims. Prefill attention cost grows quadratically with prompt length and the KV cache grows linearly, so a full window typically drops your batch size to one and adds seconds of time-to-first-token. Quality also tends to thin out toward the tail. Treat the advertised window as a ceiling, measure your task at several lengths, and pick an operating point rather than the maximum.
  • What breaks if you edit the config to claim a bigger window than the checkpoint was scaled and trained for?
    Nothing raises an error — that is the danger. The model happily accepts longer inputs and produces degraded, often incoherent output past the range its position encoding covers, because the slow-rotating dimensions are again at unseen phases. There is no runtime signal; you only see it in evaluation. Context length is a property of the trained checkpoint, not a serving-time setting to raise at will.

saying these in an interview costs you the question

  • Saying the window was extended purely by fine-tuning on longer text
  • Claiming all RoPE frequencies are scaled by the same factor
  • Assuming a declared context length is free to use fully
  • Believing you can raise max_position_embeddings at load time
  • Naming an unrelated extension method as Llama 3.1's mechanism

context