How does lost-in-the-middle position bias affect where you place key facts in a prompt?
answer
- position matters, not just presence
- the curve is not flat
- primacy and recency both help
- the middle region sags
- hold length fixed, sweep depth
basics
~20 sModels attend most reliably to the beginning and end of a long prompt and least reliably to its middle, producing a U-shaped recall curve. Put instructions and the most critical evidence at the edges, not buried mid-context.
solid answer
~50 sLong-context recall is not uniform across position. When you hold the prompt length fixed and move a single required fact through it, accuracy is typically highest near the start and near the end and sags in the middle — a U-shaped curve, usually called lost-in-the-middle. Primacy and recency both help; the middle gets neither. A concrete case: a clinical-trial protocol assistant reliably finds an exclusion criterion when it sits at the top or the bottom of a 300-page protocol, and misses the same criterion when it sits at roughly 55% depth. The practical response is to treat position as a controllable variable — put the task instruction and the highest-scoring retrieved evidence at the edges, restate the question after a long body of text, and cut filler so that less material lands in the sag. Ordering is a cheap mitigation, not a fix; verify it with a depth-swept eval rather than assuming it worked.
go deeper
Recall that a fact buried in the middle of a very long prompt is more likely to be missed than the same fact at the start or the end, and that repeating the question at the end of a long prompt helps.
Explain the U-shaped recall curve and why primacy and recency both favour the edges. Describe the experiment that shows it: fix the prompt length, move one fact through it, measure accuracy at each depth.
Demonstrate that you design around it in production — ordering retrieved evidence, restating instructions before generation, trimming filler — and that you verify each mitigation with a depth sweep instead of assuming it worked.
Own the boundary where layout stops being the answer. Aggregation and whole-document reasoning cannot be fixed by ordering, so decide when the architecture must shift to decomposition and combination rather than one large prompt.
## The shape of the failure If you take a long prompt, hold its total length constant, and slide one required fact from the very beginning to the very end while re-asking the same question, you do not get a flat accuracy line. You get a curve that is high at both ends and lower in between. This is the finding usually referred to as **lost in the middle**, and it is one of the most robust empirical facts about long-context behaviour: presence in the window does not imply usability, and *where* something sits matters as much as *whether* it is there. ## Why the ends win Two effects overlap. **Primacy**: the opening of the prompt is where instructions, system framing and role material conventionally live, so models are heavily trained to weight it — and every subsequent token attends back over it. **Recency**: the tokens nearest the point of generation are the freshest and most strongly attended, which is why an instruction repeated at the very end often overrides one given ten thousand tokens earlier. The middle inherits neither advantage: it is far from the framing at the top and far from the generation point at the bottom, so it competes for attention on the merits alone against a very large field. The exact shape is model-specific and moves with training. Some models are notably more end-heavy than start-heavy; some flatten the curve considerably after long-context mid-training. What has been stable is the direction: the middle is the weakest region, and the sag deepens as the prompt gets longer. ## A concrete case A clinical-trial protocol assistant is answering eligibility questions over a 300-page protocol plus its amendments. The team runs a depth sweep: the same exclusion criterion — a renal-function threshold — is planted at 0%, 25%, 50%, 75% and 100% of the way through, and the same question is asked each time. The criterion is found essentially every time at the top and at the bottom. At around 55% depth, the model answers as though the criterion does not exist and declares the patient eligible. Nothing about the fact changed; only its coordinates did. That single experiment is the most persuasive artifact you can bring to a design review, because it converts a vague worry about "long prompts" into a measured, reproducible bias. ## What to do about it **Order deliberately.** If you are assembling retrieved passages, do not hand them over in raw score order top-to-bottom. Place the highest-scoring evidence at the edges of the block and let the weaker material occupy the middle, where it will be underweighted anyway. Some teams interleave — best first, second-best last — precisely to exploit both ends. **Restate the task at the end.** After a long body of text, repeat the question or the output requirement immediately before generation. This costs a handful of tokens and reliably recovers instruction-following that the body of the prompt had washed out. **Cut, don't only reorder.** The sag is a function of length. Reducing the amount of low-value material shortens the region where the bias bites and often outperforms any reordering scheme, because it moves the useful content closer to an edge as a side effect. **Make position visible.** Section headers, document identifiers and explicit chunk boundaries give the model addressable structure, so an answer can be grounded in "section 7.2" rather than in a wall of text. This does not eliminate the bias but it makes mid-prompt material easier to locate when the model is prompted to look for it. **Verify, do not assume.** Reordering feels like it should work, and it usually helps, but the size of the help is model-specific. The check is the same instrument that revealed the problem: hold length constant, sweep depth, measure. If your reordering strategy is not measurably better on that sweep, it is superstition. ## The limits of the mitigation Ordering tricks address *retrieval of a single fact*. They do much less for tasks that must use material from everywhere at once — counting occurrences across a whole document, reconciling a dozen sections, or summarizing faithfully. In those tasks there is no edge to promote the important content to, because all of it is important. That is the point at which the answer stops being prompt layout and becomes decomposition: process the document in pieces and combine the results, rather than hoping a single pass attends evenly to everything. It is also worth being honest that the effect is not universal or fixed. Newer models flatten the curve, and some report near-flat recall on simple single-fact retrieval while still sagging badly on harder mid-prompt reasoning. Claiming a specific U shape for a specific model without having measured it is exactly the kind of stale prior that long-context work punishes.
- If reordering retrieved chunks helps, why not just retrieve fewer of them?Usually you should. Cutting the number of chunks shortens the sag region and removes competing near-miss material at the same time, which typically beats any ordering scheme. The reason not to cut aggressively is recall: fewer chunks means a higher chance the answer was never retrieved at all. The practical move is to tune k against measured end-to-end accuracy, then order what survives — treating ordering as a cheap refinement on top of a well-chosen k, not a substitute for it.
- Does putting the instruction at both the top and the bottom of the prompt cause conflicts?Only if the two copies differ. Duplicating an identical instruction is a standard and safe technique: it gets primacy and recency at once, at the cost of a few tokens. The failure case is drift, where the copies were edited at different times and now disagree — the model will typically follow the later one, which makes the bug hard to spot. Generate both copies from a single source string so they cannot diverge.
- Which long-context tasks do ordering tricks fail to help?Anything that must draw on the whole document at once: counting or aggregating across all sections, reconciling many mutually-referencing passages, or producing a faithful full-document summary. There is no edge to promote the key evidence to when every part is key. Those tasks need decomposition — chunked passes with a combining step — rather than clever layout of one enormous prompt.
saying these in an interview costs you the question
- Assuming any fact inside the window is equally usable
- Believing recall is flat until the hard token limit
- Handing retrieved chunks over in raw score order without thought
- Claiming a bigger window removes position bias
- Treating reordering as a complete fix for aggregation tasks