For a 1M-token window, do you pretrain natively long or extend a checkpoint post-hoc?
answer
- three budgets: compute, data, serving
- pretrain short, extend in a named stage
- million-token documents barely exist
- renumbering positions does not change cost
- and ask whether length was the real requirement
basics
~20 sAlmost nobody does either purely. Pretraining every step at 1M is prohibitively expensive and there is not enough naturally long data; a pure post-hoc stretch of a dense-attention checkpoint yields a nominal window that is neither usable nor affordable. The practical answer is architecture chosen up front plus a dedicated long-context training stage.
solid answer
~50 sFrame it as three costs, not two options. **Training cost**: attention work grows superlinearly with sequence length, so pretraining at 1M throughout would spend most of the budget on a length that most tokens do not need — which is why the field settled on pretraining mostly short and adding a named long-context stage before post-training. **Data**: genuinely 1M-token documents barely exist, so long samples must be constructed by packing related material and generating synthetic long-dependency tasks; without that, long positions get seen but never exercised. **Architecture and serving**: making 1M affordable at inference is decided when the model is designed — the attention design has to support it — and cannot be retrofitted by renumbering positions. So the honest recommendation as of mid-2026 is: choose an architecture that can serve the length, pretrain short, then run a real long-context stage; a pure post-hoc stretch is the right instrument only for a modest multiple, not a hundredfold.
go deeper
Know that a very long context window comes from a deliberate training stage and an architecture chosen for it, not from a setting, and that longer prompts cost more to serve.
Explain the staged pipeline: pretrain mostly short, then a dedicated long-context stage with position rescaling. Be able to say why full-length pretraining throughout would waste most of the compute budget.
Weigh the three budgets concretely — training compute, availability of genuinely long data, and per-request serving cost — and state that renumbering positions changes what the model accepts but never what a long prompt costs to run.
Own the decision and challenge the requirement. Recommend a path based on how far the target is from the current window and whether you control training, and be ready to argue that a retrieval-and-compaction system beats a maximal window on cost, latency and often accuracy.
## Why this is a judgment question, not a lookup Interviewers ask it because it forces you to hold three separate budgets in mind at once — training compute, data availability, and serving economics — and because the tempting answer ("just apply a scaling recipe, it worked for 8K to 32K") fails at this magnitude for reasons that are not obvious until you cost it out. ## Budget one: training compute Attention cost grows superlinearly with sequence length, so the per-step cost of training at 1M is enormous relative to training at a few thousand tokens. Meanwhile the overwhelming majority of what a model needs to learn — grammar, world knowledge, reasoning patterns, code idiom — is learnable from short samples. Spending the pretraining budget at maximum length would buy almost nothing extra for most of the tokens and would slash the total token count you could afford. This is why the standard pipeline now has a distinct **mid-training** stage between pretraining and post-training: pretrain the bulk at short-to-moderate lengths, then run a comparatively small stage specifically on long sequences, applying position rescaling as part of it. It captures most of the long-context benefit for a small fraction of the compute. "Natively long" in practice almost always means this staged approach with the architecture designed for length from the start — not full-length training throughout. ## Budget two: data The more binding constraint at 1M is that the data does not exist. Genuinely coherent million-token documents are vanishingly rare — even large repositories and long books fall far short, and the ones that qualify are heavily skewed in domain. So long training samples have to be **manufactured**: pack related documents together so that later parts genuinely depend on earlier ones, concatenate a repository with its history and its issue tracker, synthesize tasks whose answer requires combining facts placed far apart. If you skip this and pad with unrelated text, the model sees long positions but never needs to attend across them, so it learns nothing about long-range dependency and you get a nominal window with no usable capability behind it. That failure mode — high advertised length, weak actual performance — is common and is exactly what modern multi-hop long-context evaluations expose. ## Budget three: architecture and serving This is where the post-hoc option really breaks down. Renumbering positions changes what the model *accepts*; it does nothing to what a long prompt *costs*. A dense-attention checkpoint stretched to 1M still pays full attention cost over a million tokens on every request, plus the per-token state that must be held throughout. For most products the resulting per-request economics are simply not viable, regardless of quality. Making very long contexts affordable is an architectural decision taken at design time: how the attention is structured, whether spans are local with periodic global mixing, whether the per-token state is compressed. Those choices are baked into the weights. You cannot retrofit them onto a shipped checkpoint with a scaling recipe. ## Budget four, usually forgotten: evaluation Whoever ships a 1M window owns proving it works. Single-fact retrieval tests are close to worthless at this scale — they are passed by models that fall apart on anything requiring several linked facts. You need multi-hop and multi-needle evaluation at your target length, and you must expect the honest result: performance at 1M typically sits far below performance at a fraction of that length. Budget the eval work as part of the decision, because a window you cannot certify is a liability. ## How to actually decide A defensible framing: - **Modest extension (a few times the original), existing checkpoint, dense attention.** Post-hoc scaling plus a short long-data stage is the right instrument. Cheap, well-understood, effective. Measure the short-context regression and decide whether to accept it. - **Very long target (tens to a hundred times), you control training.** Design the architecture for the length, pretrain short, run a real long-context mid-training stage on constructed long data. This is what shipping a long-context model actually looks like. - **Very long target, you do not control training.** Do not stretch your way there. Adopt a model built long, and spend the saved effort on the system layer instead — retrieval, compaction, offloading content by reference, sub-agent isolation — which is usually where the win is anyway. ## The strategic point worth making out loud Length is rarely the actual requirement. "We need 1M tokens" almost always decomposes into "we need the model to act on a large body of material", and a system that retrieves and compacts the relevant slice usually beats one that pastes everything in — on cost, on latency, and frequently on accuracy, since a long prompt full of marginally relevant material degrades reasoning. A principal-level answer names the extension tradeoff correctly *and* challenges whether the length target was the right requirement to begin with.
- Why is data, rather than compute, often the binding constraint at very long target lengths?Because coherent documents of that length barely exist. Compute can be bought; a corpus of genuine million-token documents cannot. Long samples must be constructed — packing related material so later parts truly depend on earlier ones, synthesizing tasks whose answers require combining distant facts. Skip that and the model sees long positions without ever needing to attend across them, producing a nominal window with no capability behind it.
- A vendor advertises a 1M-token window. What would you ask before believing it is usable?How it was evaluated. Single-fact retrieval at length is close to meaningless — models pass it while failing anything requiring several linked facts. Ask for multi-hop, multi-needle results measured at the full advertised length and compared against the same tasks at a fraction of it. Expect a substantial gap, and treat the advertised number as a ceiling on input size rather than a statement of usable capability.
- Your product lead wants a 1M-token window to load an entire customer archive. How do you push back?Reframe the requirement. The need is to act on a large body of material, not to place all of it in one prompt. A retrieval-and-compaction layer that supplies the relevant slice is usually cheaper, faster, and often more accurate, because a prompt padded with marginally relevant content degrades reasoning. Reserve the very long window for cases with genuinely irreducible whole-corpus dependencies, and cost both paths per request before choosing.
saying these in an interview costs you the question
- Assumes a scaling recipe alone can take a checkpoint to 1M
- Thinks position rescaling reduces the cost of long prompts
- Believes frontier models pretrain at full length throughout
- Ignores that long training data must largely be constructed
- Treats a 1M advertised window as 1M of usable capability