Why does a model's effective context length fall short of its advertised token limit?
answer
- spec sheet versus measured behaviour
- the ceiling is not the cliff
- accuracy curve breaks well before the limit
- sweep length and depth, plot accuracy
- clean early, ragged band later
basics
~20 sAdvertised length is the maximum input the model accepts without erroring; effective length is how far accuracy actually holds. Attention spreads thinner, long-range dependencies are rare in training data, and more text means more plausible-but-wrong material to attend to.
solid answer
~50 sAdvertised context length is a hard input ceiling — the point at which the request is rejected. Effective context length is the point at which task accuracy starts falling apart, and it is usually far lower. The gap comes from attention having to spread over many more tokens, from genuinely long-range dependencies being rare in training data even after a long-context mid-training stage, and from a longer prompt containing more plausible-but-wrong material competing for the model's attention. Treat effective length as an empirical property of *your* model on *your* task, not a spec-sheet number: run a grid that varies input length and the position of the fact you need, plot accuracy, and find where the curve breaks. A 200K-advertised model often shows clean recall out to roughly 40K and a ragged, task-dependent band beyond it. As of mid-2026, models advertising 1M routinely score far below their 200K-range numbers on published long-context suites.
code
python · 9 linesdef build_haystack(filler: str, needle: str, total_tokens: int, depth: float) -> str:
approx_chars = total_tokens * 4
body = (filler * (approx_chars // len(filler) + 1))[:approx_chars]
cut = int(len(body) * depth)
return body[:cut] + "\n" + needle + "\n" + body[cut:]
grid = [(n, d) for n in (8000, 32000, 128000) for d in (0.0, 0.25, 0.5, 0.75, 1.0)]
print(len(grid), "cells")
print(build_haystack("protocol filler text. ", "Exclusion: eGFR below 30.", 100, 0.5)[:80])go deeper
Know that the advertised context limit is only the point where a request is rejected, and that answer quality can drop long before you reach it. Say plainly that you would test with real prompt sizes rather than trust the number.
Be ready to explain the mechanics: attention mass spread across more tokens, scarce long-range examples in training, and more confusable material in a longer prompt. Describe a length-by-depth sweep and what its accuracy curve looks like.
Show that you have measured this on a real system. Talk about repeating each cell because failures are stochastic, using filler that matches your corpus, and turning the break point into a prompt-size budget that the pipeline enforces.
Own the framing that effective length is a per-task, per-model property that must be re-measured on every model change, and that vendor window increases do not transfer into product quality. Decide when the answer is retrieval and decomposition rather than a longer prompt.
## Two different numbers Every model ships with a context limit — the number of tokens it will accept in one request. That number is a **hard ceiling**: exceed it and the request is rejected outright. It says nothing about quality. What matters in practice is the **effective context length**: the input size beyond which the model's accuracy on your task degrades to the point where you would not ship it. These two numbers are frequently an order of magnitude apart, and the industry vocabulary for the gap has settled on phrases like *context rot* — quality decaying as the window fills, rather than failing cleanly at the edge. ## Why the gap exists Three forces stack up. **Attention dilution.** Attention is a soft, normalized competition: every token's attention mass is spread across all the tokens it can see. As the sequence grows, the share any individual token can command shrinks, and a single decisive fact competes with far more irrelevant material. Nothing forbids the model from finding it; the signal-to-noise ratio simply gets worse. **Training-distribution scarcity.** Even when a model has been trained or mid-trained to a very long window, the number of training examples that require genuinely resolving a dependency across hundreds of thousands of tokens is tiny compared to short-range examples. The model has far less practice at long-range work than at short-range work, so its skill drops off with distance even where its architecture permits the span. **More text, more competition.** A longer prompt is not just longer, it is richer in near-misses: older versions, similar phrasings, adjacent-but-wrong facts. Degradation is therefore partly a property of *what* is in the window, not only *how much*. ## How to measure it The standard instrument is a length-by-depth grid, historically called needle-in-a-haystack. You take filler text, insert a known fact ("the needle") at a controlled fractional depth, ask a question only that fact answers, and score the answer. Sweeping two axes — total input length and insertion depth — gives a heat map instead of a single pass/fail. A realistic run looks like this. A clinical-trial protocol assistant is being built over a 300-page protocol plus its amendments, on a model advertising 200K. You build inputs at 8K, 32K, 64K, 128K and 190K tokens, insert a real exclusion criterion at 0%, 25%, 50%, 75% and 100% depth, and run each cell many times. The result is typically clean, near-perfect recall out to something like 40K at every depth, then a ragged band: some depths still fine at 128K, the middle depths dropping sharply, and high variance run to run. That break point is your effective length **for that task**. Change the task — reasoning over the fact rather than quoting it — and the break point moves down again. Two design rules make the number trustworthy. Repeat each cell: long-context failures are stochastic, and a single sample tells you nothing about a cell that succeeds 60% of the time. And use filler that resembles your real corpus, because degradation depends on how confusable the surrounding text is, not just on token count. ## What the number is for Once you have it, effective length becomes a **budget**, not a trivia fact. It tells you the maximum you may put in a prompt before you must retrieve, summarize, or split the work — and it tells you which of your failures are prompt-length failures rather than model-capability failures. It also reframes vendor announcements: a jump from 200K to 1M advertised is not a promise that your 700K-token prompt will work, and published suites in 2026 consistently show large drops at the top of the advertised range. The honest reading of a bigger window is "a bigger place to put things", not "a bigger place to put things and expect the same accuracy". ## Common mistakes Assuming the advertised number is a quality guarantee is the biggest one. Second is measuring once, with one task shape: retrieval, multi-hop reasoning and whole-document aggregation each have their own, decreasing, effective length. Third is treating the cliff as sharp — it is a slope with a widening variance band, so a prompt that works today at 150K can fail tomorrow on slightly different input. Fourth is assuming effective length transfers across models or across versions of the same model; it is cheap to re-measure and expensive to be wrong.
- How would you tell a long-context degradation apart from a plain retrieval bug in a RAG pipeline?Isolate the variable. Feed the model a prompt that provably contains the answer — verify the string is in the input — and re-ask. If it answers correctly at short length but fails at long length with the identical evidence, that is long-context degradation. If it fails at both lengths, or the evidence was never in the prompt, the retrieval step is the problem. Logging which chunks were actually included makes this a five-minute check rather than a guess.
- Does raising the effective length require a different model, or can you get there with prompt design?Both, at different scales. Prompt design buys real gains: trimming irrelevant material, deduplicating near-identical passages, placing critical evidence at the edges, and structuring the input so sections are addressable. Those typically move the break point meaningfully but not by an order of magnitude. Getting a genuinely longer usable window is a model property — you change models or you change architecture, which is a procurement decision, not a prompting one.
- Why report a distribution rather than a single accuracy number for each cell in the grid?Long-context failures are stochastic. A cell that passes once may pass 60% of the time, and a single sample cannot distinguish 100% from 60%. Repeating each length-by-depth cell many times and reporting the success rate, plus its spread, tells you where behaviour is merely noisy versus genuinely broken — and noisy is what you actually ship against, since production sees one sample per user request.
saying these in an interview costs you the question
- Treating the advertised window as a quality guarantee
- Claiming accuracy is flat until the hard limit, then fails
- Measuring effective length once and reusing it across models
- Assuming a bigger window removes the need for retrieval
- Testing with one prompt length and calling it validated