In a 20-document prompt, why is accuracy lowest when the answer sits in the middle?
answer
- accuracy versus position is not flat
- the shape of the curve is a letter
- best at the edges, worst between them
- more documents can lower accuracy
- evidence present but positionally disadvantaged
basics
~20 sThis is the lost-in-the-middle effect: retrieval accuracy traces a U-shaped curve against position. Models use information best when it appears first or last in the prompt and measurably worst when it is buried in the middle, even though every document is inside the window.
solid answer
~50 sThe finding comes from multi-document question answering experiments (Liu et al., 2023) where the same set of documents is fed in different orders and only the position of the gold document changes. Accuracy is highest when the gold document is first, drops as it moves toward the centre, and partially recovers when it is last — a **U-shaped curve**. In the original setup the mid-position result could fall below the model's accuracy with no documents at all, which is the striking part: adding correct evidence in the wrong position made things worse than omitting it. Two forces explain the shape. Causal attention gives the earliest tokens outsized influence, and distance-decaying positional encodings favour the tokens nearest the generation point. The practical consequence is that stuffing more retrieved material into the prompt has a ceiling: past some count, extra documents mostly enlarge the weak middle.
go deeper
Recall the name and the shape: lost in the middle, a U-shaped accuracy curve where the first and last positions beat the middle. Say that being inside the window is not the same as being used.
Explain the controlled setup — same documents, only the gold position varies — and give both mechanisms, attention sinks at the start and distance decay at the end. Be ready to say why more documents can hurt.
Show the diagnostic instinct: separate 'evidence was not retrieved' from 'evidence was retrieved but ignored' by moving it and re-running, and describe a front-versus-back evaluation that quantifies your own exposure.
Own the design implication across the product: a policy of fewer, higher-signal items rather than maximal recall, with the accuracy-versus-count curve measured for your workload rather than assumed, and cost and latency scored alongside it.
## The experiment The cleanest demonstration is a controlled multi-document question-answering setup. The model is given a question and a fixed number of documents — say 10, 20 or 30. Exactly one of them, the *gold* document, contains the answer; the rest are relevant-looking distractors. The experimenters then vary one thing only: which slot the gold document occupies. Position 1, position 5, position 10, position 20. Because the token count, the distractors and the question are held constant, any change in accuracy is attributable to position alone. What comes out is a **U**: highest accuracy at the first position, a decline through the middle, and a partial recovery at the last position. The same shape appeared in a synthetic key-value retrieval task, which rules out the explanation that it is something about natural-language documents specifically. The headline result from Liu et al. (2023) is that for some models the mid-position accuracy fell below the closed-book baseline — the model's accuracy with no documents supplied at all. Correct evidence, placed badly, was worse than no evidence. ## Why the curve has that shape **Primacy — attention sinks.** In a causal decoder every token can attend to all tokens before it, so the earliest positions are visible to the entire sequence. Empirically, attention mass concentrates on those opening tokens; they behave as a sink. Content sitting there is structurally hard for the model to overlook. **Recency — distance decay.** Widely used positional schemes make attention weaker as the distance between two tokens grows. Generation happens right after the final prompt token, so the tail is the nearest and cheapest region to condition on. **Training distribution.** Instruction-tuned corpora put the salient thing at an edge far more often than in the interior. The model learned a prior about where important content lives. These are model properties, not dataset accidents. Analyses of causal attention suggest a position bias is present in decoder-only architectures before any training signal shapes it, with training then modulating how strong it is. ## What it does *not* say It does not say middle content is unreadable. Ask a pointed question — 'what does paragraph 14 say?' — and the model will usually retrieve it. The failure is one of *spontaneous use*: when the model must simultaneously decide what is relevant and reason with it, the middle loses the competition. That is why the effect bites hardest on tasks that combine search with synthesis, and least on tasks where the user has already pointed at the passage. It also is not a claim that all models degrade equally. The magnitude is model-specific and has narrowed on some frontier models, but the shape keeps reappearing in newer long-context evaluations, and no released model has been shown to be flat across positions on hard multi-document tasks. ## Consequences for prompt design **More context is not monotonically better.** Every document you add past the useful few mostly grows the weak middle. There is a count beyond which extra retrieved material lowers answer quality while raising cost and latency — the curve for accuracy-versus-document-count typically peaks and then declines. **Fewer, better items beat more items.** Cutting a candidate set from thirty passages to five raises the share of the prompt that sits in a strong position, independent of any change in retrieval quality. **Order is a lever you already own.** If you know which material is most likely to matter, putting it at an edge rather than in the interior is free — no extra tokens, no extra call. **Instructions are the most fragile passengers.** A constraint stated in the middle of a long prompt is the classic silent failure: no error, no refusal, just quietly unfollowed rules. ## Diagnosing it in your own system The telltale signature is an answer quality that depends on *where* the right evidence landed rather than *whether* it was retrieved. If your retrieval logs show the correct passage was in the prompt but the model answered from a different one, and if reordering that passage to the front or back fixes the answer, you are looking at position bias rather than a retrieval failure. Those two failures have completely different fixes, and conflating them wastes weeks: improving retrieval will not help a system whose evidence is already present but positionally disadvantaged. A cheap standing check is to run a fixed evaluation set twice with the supporting material at the front and then at the back. A large gap between those two runs tells you how exposed your prompt shape is; a small gap tells you position is not currently your bottleneck.
- If the correct passage is in the prompt but the model answers from a different one, how do you tell position bias from a retrieval bug?Check whether the answer changes when you move that passage to the front or the back without changing the set. If moving it fixes the answer, retrieval did its job and position is the bottleneck. If the answer stays wrong at every position, the problem is elsewhere — ambiguous phrasing, a distractor that reads as more authoritative, or a genuine reasoning failure. The fixes are unrelated, so distinguishing them first matters.
- Why can adding more supporting documents make answer quality worse rather than better?Beyond a small number, extra documents mostly enlarge the low-influence middle region and add distractors competing for attention. Accuracy against document count typically rises to a peak and then declines, while cost and latency rise monotonically. That is why selecting fewer, higher-signal items usually beats maximizing recall into the prompt.
- Does the U-shaped curve go away on a model with a very large context window?No. A larger window changes how much you can fit, not how attention is distributed across it. The shape has kept reappearing in long-context evaluations of newer models; what varies is amplitude, which is model-specific. Assuming a big window makes ordering irrelevant is the most common wrong inference from the capacity numbers.
- Was the effect shown only on natural-language documents?No — it reproduced on a synthetic key-value retrieval task, where the model must return the value for a given key from a list. That rules out explanations based on document semantics or distractor plausibility and points at the architecture and training distribution as the source of the bias.
saying these in an interview costs you the question
- Says the effect disappeared with long-context models
- Treats every in-window token as equally usable
- Confuses position bias with a retrieval miss
- Assumes accuracy rises monotonically with document count
- Claims mid-prompt content cannot be retrieved at all