Why is a model's advertised 1M-token window not 1M usable tokens?
answer
- big window, smaller usable band
- acceptance limit, not a quality promise
- quality falls as the window fills
- the curve has a cliff, not a slope
- measure the band per model and task
basics
~20 sAdvertised length is an acceptance limit, not a quality guarantee. Retrieval accuracy and instruction-following degrade as the window fills — the effect commonly called context rot — so the usable ceiling is typically a fraction of the maximum, and it must be measured.
solid answer
~50 sThe advertised number tells you what the serving stack will accept without an error. It says nothing about whether the model can still find, combine and act on everything you sent. Across the long-context benchmark families used in 2026 — RULER-style multi-needle suites, MRCR, NoLiMa, LongBench-v2, HELMET — the consistent finding is that quality holds for some band of utilization and then falls off, with reported effective ranges commonly landing well under the advertised maximum and decay that is cliff-shaped rather than gently linear. A model advertising a 1M-token window can score near-perfectly on a single planted fact and still lose several of twelve interdependent facts once the window is genuinely full. Practically, that means you treat the advertised length as a hard error boundary and your own measured band as the design limit, keep utilization inside it, and re-measure per model rather than assuming the ceiling transfers.
go deeper
Know that the advertised window is the maximum a request may be, not a promise that the model uses it all well, and that shorter, more relevant prompts usually beat longer ones.
Explain context rot concretely: the same task, held fixed, scores worse as surrounding tokens grow; decay is cliff-shaped rather than linear; and aggregation across facts breaks before single-fact lookup does.
Show production judgment — enforce a utilization cap in code, log prompt tokens against the window alongside quality signals, and re-measure the band whenever you change model, prompt structure or retrieval depth.
Own the tradeoff across the fleet: how much window each product surface is allowed to spend given cost and latency, when a larger-window model is worth its price, and how migrations are gated so a bigger advertised number never silently raises utilization.
## Advertised length is an acceptance limit When a provider says a model has a 1M-token window, that is a statement about the serving stack: requests up to that size will be accepted rather than rejected. It is not a claim that the model attends to a million tokens with uniform competence, and it is not backed by a per-task quality guarantee. Confusing the two is the single most common context-budgeting mistake, and it is expensive: you pay for every token you send, latency scales with prompt size, and past a certain fill level you are paying more for a worse answer. ## What "context rot" names The term describes a measured phenomenon: as the number of tokens in the prompt grows, model performance on the same underlying task declines. The task is held constant — same question, same target facts, same grading — and only the surrounding volume changes. The result is that a question the model answers reliably at 20k tokens becomes unreliable at 400k, even though nothing about the question got harder. It is not a bug in a particular model; it shows up across vendors and across architectures, and it is why "just send everything, the window is huge" is not a strategy. Degradation shows up in more than one dimension. Retrieval of specific facts weakens. Instruction-following weakens — constraints stated in a long prompt are more often ignored. Aggregation across many pieces of evidence weakens fastest of all, because it requires holding several retrieved items in play at once rather than surfacing one. ## Why the decay is cliff-shaped The practically important property is that the curve is not a gentle slope you can interpolate. Measured curves typically hold roughly flat through a band and then drop, sometimes sharply, past a threshold that varies by model and by task type. Two consequences follow. First, you cannot infer behaviour at 80% utilization from a spot check at 30% — the interesting region is exactly the one people skip testing. Second, a small increase in retrieved context can produce a large, discontinuous quality regression, which is what makes "we added one more document to the prompt and accuracy fell off" a real and confusing production incident. Task difficulty moves the cliff. Simple lookup of a distinctive string survives far more fill than multi-hop reasoning over several facts that must be reconciled. So there is no single effective-context number for a model — there is a number per task shape, which is precisely why vendor headline figures do not transfer to your workload. ## Where the numbers come from The published picture as of mid-2026 comes from a family of long-context suites rather than any one test. Synthetic scaling suites in the RULER style vary both length and task type — multi-key retrieval, variable tracing, aggregation — so you can watch which capability breaks first. MRCR-style evaluations stress distinguishing many similar earlier turns. NoLiMa deliberately removes literal word overlap between the query and the planted fact, so the model cannot succeed by keyword matching. LongBench-v2 and HELMET use realistic documents and heterogeneous task types. Read together, they agree on the shape: effective context is materially smaller than advertised context, the gap widens as tasks require combining information, and headline window sizes have grown faster than the usable band inside them. Treat any specific percentage you have read as a rough prior, not a constant. The ratio moves with model generation, with task type, and with how adversarial the surrounding content is. ## What to do with this The operational stance is straightforward. **Budget below the ceiling, not to it.** Pick a utilization target you have evidence for and enforce it in code — cap retrieved chunks, cap history, cap tool output — rather than filling whatever space happens to be free. **Prefer signal density over volume.** Since every token spends a finite attention budget, twenty well-chosen chunks routinely beat two hundred loosely relevant ones, and the shorter prompt is cheaper and faster as well. "It fits" is not a reason to include something. **Watch what fills the window silently.** Tool definitions, accumulated tool results and long histories grow without anyone deciding to grow them, and they consume the same budget as evidence you deliberately retrieved. **Measure per model.** A new model with a larger advertised window has its own curve. Migrating and simultaneously raising utilization because "the window is bigger now" is how teams ship a regression they cannot explain. **Instrument utilization in production.** Log prompt tokens as a fraction of the window alongside quality signals, so you can see whether failures cluster in the high-fill band. Without that, high-utilization degradation is invisible — nothing errors, the answers are just quietly worse. ## The honest summary for an interview Advertised context is a capacity number; effective context is a quality number. Only the second one should shape your design, and only you can measure it for your workload.
- Two models both advertise 1M tokens. Can you assume their usable ceilings are similar?No. The effective band varies by model generation and by task type, and vendors optimize for different long-context profiles. One may hold up on single-fact retrieval far past where it collapses on multi-fact aggregation, while another shows the opposite. The ceiling is a property of the model-plus-task pair, so it has to be measured on your own tasks for each candidate model before you set a budget.
- Your accuracy dropped after you raised the retrieved-chunk count from 20 to 60. What is the likely explanation?You probably crossed into the degraded band and added competing material at the same time. More chunks means higher utilization and more near-miss passages the model must discriminate against, and both push in the same direction. The diagnostic is to re-run the same eval at both chunk counts while holding questions and grading fixed, and to check whether the extra chunks ever contain the answer — often they add tokens and no new evidence.
- If effective context is well below advertised, why do vendors keep raising the advertised number?Because acceptance capacity is genuinely useful and genuinely hard to build: it lets you put a whole codebase, transcript or contract set in front of the model without an upstream selection step failing silently. The number is real, it is just a capacity spec rather than a quality spec. The mistake is on the consumer side when a capacity spec is read as permission to fill the window.
saying these in an interview costs you the question
- Treats the advertised window as fully usable capacity
- Assumes quality declines smoothly and predictably with length
- Spot-checks at low fill and generalizes to full windows
- Believes a bigger window removes the need for retrieval
- Carries one model's effective ceiling over to another model