How do you extend a shipped 8K-context LLM to 128K without pretraining from scratch?
answer
- reuse the checkpoint, don't retrain it
- the position axis gets compressed, not extended
- uniform squeeze versus frequency-aware squeeze
- then a short stage on genuinely long data
- NTK-aware, YaRN, LongRoPE, mid-training
basics
~20 sRescale the model's rotary position encodings so 128K positions fall inside the range it was trained on — uniform position interpolation, or frequency-aware variants such as NTK-aware scaling, YaRN and LongRoPE — then continue training briefly on long sequences so the model adapts.
solid answer
~50 sPost-hoc extension is a two-part recipe you run on an already-trained checkpoint. First you change how positions are fed in: a model trained on 8K has never seen position 100,000, and rotary encodings extrapolate badly, so instead of extrapolating you **compress** the position axis. Plain position interpolation divides every index by the extension factor (16x for 8K→128K) so the model only ever sees positions it was trained on. Frequency-aware schemes — NTK-aware scaling, YaRN, LongRoPE — do this unevenly, leaving the fast-rotating dimensions that carry short-range, token-adjacent distinctions nearly untouched and stretching mainly the slow ones, which preserves local resolution that a uniform squeeze destroys. Second, you run a short **continued-training** (mid-training) stage on genuinely long sequences — long logs, repositories, books, packed documents — usually a tiny fraction of pretraining compute. Rescaling alone typically leaves a model that is nominally 128K but incoherent well before it; the long-data stage is what makes the stretched positions usable.
code
python · 8 linesdef interpolated_position(pos, trained_len=8192, target_len=131072):
"""Uniform position interpolation: squeeze indices into the trained range."""
scale = target_len / trained_len # 16.0
return pos / scale
print(interpolated_position(0)) # 0.0
print(interpolated_position(100_000)) # 6250.0 -> inside the trained 8K rangego deeper
Know that a context window comes from how the model was trained, not just a setting, and that extending it means rescaling position information plus extra training. Being able to say why simply raising the limit does not work is enough here.
Explain the two halves of the recipe: compress the position axis so long indices land inside the trained range, then continue training on long sequences. Be ready to contrast uniform interpolation with frequency-aware scaling and say what each costs.
Show you have run one. Talk about how much long data the stage needed, what long data you actually had, how you validated the result beyond the fact that long prompts stopped erroring, and what regressed.
Own the decision of whether post-hoc extension is the right instrument at all. Weigh a cheap stretch of an existing checkpoint against selecting or training a natively long model, and be clear that extension changes what the model accepts, never what serving a long prompt costs.
## The problem extension solves A transformer's context window is not a hard architectural wall in the way a fixed-size buffer is. It is a **training-distribution boundary**: the model saw sequences up to some length during pretraining, and the position information it receives at inference was only ever exercised over that range. Feed it a position far outside that range and the position signal becomes out-of-distribution — attention scores go haywire and output collapses into repetition or nonsense, often far short of the nominal limit. Extension techniques are the family of tricks that push that boundary out on a checkpoint you already have, rather than paying for a fresh pretraining run at the longer length. ## Step one: change the position mapping Modern open models use rotary position encoding, where each position rotates parts of the query and key vectors by an angle proportional to the index. The rotation angles are periodic, and the model has only learned to interpret the range of angles it saw. Three deployment recipes, in increasing sophistication: **Extrapolation (the null recipe).** Just raise the configured maximum length and feed longer inputs. The model sees rotation angles it never trained on. In practice quality degrades sharply not far past the original limit. This is why "I set the max length to 128K" is not a context extension. **Uniform position interpolation.** Divide every position index by the extension factor. Stretching 8K to 128K means dividing by 16, so token 100,000 is presented as position 6,250. Every angle the model sees is now inside its trained range — nothing is out of distribution. The cost is **resolution**: adjacent tokens that used to be one unit apart are now 1/16 of a unit apart, so the fine-grained "which of these two neighbouring tokens is which" signal is compressed into a much narrower band. Interpolation is far safer than extrapolation but blunts local precision. **Frequency-aware scaling (NTK-aware, YaRN, LongRoPE).** The rotary dimensions are not interchangeable: some rotate fast and encode near-neighbour relationships, others rotate slowly and encode coarse, document-scale position. Uniform interpolation squeezes both equally, which wastes the compression budget on exactly the dimensions that could least afford it. The frequency-aware family scales the slow dimensions hard (they have huge unused range), leaves the fast ones close to untouched, and blends in between. YaRN adds an attention-temperature correction to keep attention entropy in the range the model expects. LongRoPE searches for a per-dimension scaling schedule rather than deriving one from a formula, and applies it in stages. These recipes reach the same nominal length with markedly less short-range damage and much less fine-tuning. A practically important refinement is **dynamic scaling**: pick the scale factor from the actual sequence length at inference, so a 2,000-token prompt sees essentially the original geometry and only genuinely long inputs are compressed. This directly limits the short-context tax. ## Step two: continued long-sequence training Rescaling makes long positions *legal*; it does not make them *understood*. The model still has to learn to attend across distances it never practised — to carry a definition from token 3,000 to token 90,000, to keep track of which of forty files it is editing. That is what a **continued-training / mid-training** stage buys. In the modern pipeline this is a named stage sitting between pretraining and post-training, and long-context extension is one of its main jobs. The stage is comparatively cheap — often well under a percent of pretraining tokens — but the data is the hard part. You need sequences that genuinely require long-range dependency: long build logs and firmware traces, whole repositories, book-length documents, packed multi-document samples with cross-document questions. Padding short documents together teaches nothing, because nothing in the second half depends on the first. After the frequency-aware recipes, some models need only a few hundred million tokens of long data; after plain interpolation, considerably more. ## What extension does not buy Three honest caveats to state in an interview. First, **nominal is not effective**. A checkpoint that accepts 128K may be reliable over far less; a nominal window is a ceiling, not a promise. Second, **there is a quality tax on short inputs**. Both the rescaling and the long-data training shift the model away from the short-prompt behaviour it was tuned for, and teams routinely measure a few points of regression on their original short-context evals. Third, **extension does not fix cost**. The compute and memory of attending over a long prompt are unchanged by how you numbered the positions. Making a long window *affordable* is an architecture question — sliding-window spans, sparse or hybrid attention stacks — decided when the model is designed, not a knob you turn on a shipped checkpoint. That is why the current generation reaches very long contexts natively rather than purely by post-hoc stretching.
- Why is packing many short documents together a poor substitute for long training data in that continued-training stage?Because nothing in the later half of a packed sample depends on the earlier half. The model can predict every token with purely local attention, so it never gets gradient pressure to build long-range dependencies. Genuinely long data — a full repository, a book, a multi-day log — forces the model to carry information across distance, which is exactly the skill the stage is supposed to teach.
- What is dynamic scaling, and what problem does it address?Dynamic scaling picks the position-rescaling factor from the actual input length rather than fixing it at the maximum. A 1K prompt is served with essentially the original, uncompressed geometry, and only genuinely long inputs get squeezed. It exists because a fixed maximum-factor rescale imposes its resolution loss on every request, including the short ones the model already handled well.
- Someone raises the model's configured maximum length and reports it now supports 128K. What is wrong with that claim?Nothing has changed except a configuration bound. The model is now being fed position signals far outside its training distribution, which typically produces incoherence not far past the original limit. Accepting a long input is not the same as using it. A real extension changes the position mapping and follows it with long-sequence training, and is validated on long-context evals rather than on the fact that the request did not error.
saying these in an interview costs you the question
- Says raising the configured max length is the extension
- Claims interpolation alone needs no further training
- Thinks extension makes long prompts cheaper to serve
- Assumes the nominal window is fully usable after extension
- Confuses window extension with chunking documents for retrieval