skip to content

Does position bias get worse as you fill more of a million-token context window?

level: seniorimportance: should knowfreq 31%

answer

  1. advertised capacity is not usable capacity
  2. the penalty scales with how full it is
  3. the curve's shape changes, not just its depth
  4. the tail keeps its influence, the head loses some
  5. degradation arrives as a cliff, not a slope

basics

~20 s

Yes. Position effects are mild when a prompt uses a small fraction of the window and grow sharply as utilization rises. As of mid-2026, reported behaviour is that past roughly half the window the U-shape flattens into a distance-based bias favouring the end, so early material loses its protection.

solid answer

~50 s

Advertised window size and usable window size are different numbers, and position bias is one reason. At low utilization — a few thousand tokens in a million-token window — the accuracy difference between the front, the middle and the back is usually small enough to ignore. As you fill the window, the penalty on interior content grows, and analyses reported through 2026 describe a transition: beyond roughly half the window the symmetric U flattens toward a **distance-based bias**, where influence falls off with distance from the generation point and the front-of-prompt advantage weakens. The operational reading is counter-intuitive — the fuller the window, the more the *end* is the only reliably-attended region, so a leading instruction becomes less trustworthy exactly when the prompt is longest. Degradation also tends to be cliff-shaped rather than linear and varies by model, so treat any specific threshold as something to measure rather than inherit.

go deeper

for a junior

Know that a model's advertised context window is a capacity, not a promise that everything inside it is used equally, and that longer prompts make placement matter more.

for a middle

Explain that the position penalty grows with how full the window is and that the curve's shape shifts toward favouring the end, so a leading instruction is least reliable exactly when the prompt is longest.

for a senior

Show operational instinct: log prompt size against quality, sample multiple lengths because degradation is cliff-shaped, and set a working ceiling below the nameplate with an explicit reduction path when a request exceeds it.

for a principal

Own the recall-versus-utilization tradeoff as an explicit policy rather than a default, with per-workload measurement, a re-measurement gate on every model change, and an architecture that references bulk material rather than preloading it.

## Two numbers that are not the same A model advertising a very large context window is stating a capacity, not a guarantee of uniform quality across that capacity. Long-context evaluations consistently report that useful performance holds over some fraction of the advertised span and then degrades, often sharply. Position bias is one of the mechanisms behind that gap, alongside distractor accumulation and the general difficulty of reasoning over more material. ## How the position curve changes as the window fills At low utilization, the ordering penalty is real but small. A prompt of a few thousand tokens has a short 'middle', every token is close to the generation point in relative terms, and moving content around produces modest differences. This is why teams who developed a prompt at 4k tokens often believe ordering does not matter — at that size, it nearly does not. As the prompt grows, two things change. The absolute penalty on interior content increases, because there is far more interior. And the shape of the curve itself shifts. Reports through 2026 describe the symmetric U giving way, past roughly half of window utilization, to a bias dominated by distance from the end: the model still attends strongly to the tail, but the primacy advantage at the head erodes. That matters practically because the front of the prompt is exactly where teams habitually put system instructions, role definitions and output contracts. Degradation is typically **cliff-shaped**, not a gentle slope. A prompt design can look healthy across a range of sizes and then fall off within a relatively narrow band. That behaviour defeats spot-checking at one length, because the length you happened to test may sit on either side of the cliff. ## Why bigger windows did not solve ordering The intuitive expectation was that once everything fits, placement stops mattering. The opposite happened at the level of practice: bigger windows made it easy to build prompts long enough for position effects to become dominant, and they encouraged the habit of preloading everything that might be relevant. A model with a million-token window running at 700k tokens is operating deep in the regime where interior content is weakest — an area no 8k-window prompt could ever reach. This is why the current default in context engineering is to treat window space as a budget to spend sparingly rather than a capacity to fill: load lightweight references and pull detail in on demand, keep working context small, and hand off bulk exploration to isolated processes that return short summaries. Those practices are motivated by attention quality, not only by cost. ## What to do about it **Do not rely on a leading instruction in a very long prompt.** If the prompt regularly runs long, the output contract belongs at or near the end, whatever it also says at the top. **Treat utilization as an operational metric.** Log the prompt token count per request alongside quality signals. A quality regression that correlates with prompt length is a different problem from one that correlates with a content type or a user segment, and only the log tells you which you have. **Set a working ceiling below the advertised one.** Many teams cap effective usage well under the nameplate figure and trigger reduction — selecting less material, summarizing, or splitting the work — when a request would exceed it. The specific ceiling should come from your own measurements, because it varies by model and by task difficulty; a simple extraction survives far more filling than multi-hop reasoning does. **Re-measure after a model change.** Position behaviour and effective range are model properties. A drop-in model upgrade can move the cliff in either direction, and prompts tuned around the old threshold are exactly the ones that break quietly. ## Reading the tradeoff honestly There is a real tension here. Filling the window is the cheapest way to make sure the needed information is present; keeping it small is the most reliable way to make sure the present information is used. Recall and utilization pull against each other, and the balance is workload-specific. The mistake is to resolve it by default — either by preloading everything on the assumption that capacity equals usability, or by starving the prompt on the assumption that all long context is bad. Measure where your task sits, then hold that line as a budget with an explicit reduction path when a request would exceed it.

  • Why is a leading system instruction less safe in a 700k-token prompt than in a 7k-token one?
    Because the primacy advantage that protects it weakens as utilization rises, while the tail keeps its influence. At small sizes the whole prompt is effectively near the generation point, so a leading instruction is fine. At high utilization the head is far away and increasingly outcompeted by everything between it and the answer, so the contract needs a presence near the end.
  • Why does cliff-shaped degradation make spot-checking a prompt at one length dangerous?
    Because a single length tells you which side of the cliff you were on that day, not where the cliff is. A prompt validated at 100k tokens can behave completely differently at 200k, and real traffic varies in length. Sampling several lengths across the range you actually serve, and logging prompt size against quality, is what makes the boundary visible.
  • Doesn't a bigger window at least remove the need to select what goes in?
    No — it removes the hard truncation error and replaces it with a soft quality failure, which is worse operationally because nothing raises an alarm. Selection is still what keeps the important material in a strongly-attended region. Current practice treats window space as a budget spent on high-signal tokens, with bulk material referenced and loaded on demand rather than preloaded.

saying these in an interview costs you the question

  • Assumes a bigger window makes ordering irrelevant
  • Treats advertised window size as usable capacity
  • Expects smooth linear decay instead of a cliff
  • Reuses one model's threshold for a different model
  • Validates a long prompt at a single token length

context