After reranking retrieved chunks, in what order should you place them in the RAG prompt?
answer
- score order is not read order
- the middle of a long block
- strongest evidence at both edges
- rank 1 first, rank 2 last
- measure the position curve yourself
basics
~20 sScore order is not read order. Long prompts under-use material buried in the middle, so a common choice is to put the top-ranked chunks at the very start and end of the retrieved block and let weaker chunks fill the middle.
solid answer
~50 sReranking decides *which* chunks go in; ordering decides *where* they sit, and those are separate decisions. The default is most-relevant-first — simple, and fine when only a handful of short chunks are injected. Once you are packing many chunks into a long block, the middle of that block is the weakest read position, so the strongest evidence ends up exactly where it is most likely to be skimmed past. The usual fix is bookending: rank 1 first, rank 2 last, rank 3 second, rank 4 second-to-last, and so on, so both edges of the block carry strong material. Neither ordering is universally right — the effect size depends on the model and on how long the block is — so treat it as a measured choice: hold retrieval fixed, vary only the order, and score answer accuracy on a held-out set.
go deeper
Know that where a chunk sits in the prompt affects whether the model uses it, and that pasting chunks in rerank-score order is a choice rather than the only option.
Be ready to explain both orderings and the mechanism behind them: most-relevant-first for short blocks, bookending for long ones, because the middle of a long block is the least reliably read region.
Show you would measure rather than assume — freeze retrieval, vary only the permutation, sweep the gold chunk through every position, and log the order with every answer so you can tell a retrieval miss from a burial.
Own the framing that ordering is a cheap, testable lever that competes with expensive ones. Argue for spending on it only when the block is long enough for it to pay, and keep the permutation in one place so the whole system's prompt shape stays explainable.
## Score order is not read order A RAG pipeline typically ends with a reranker that assigns each candidate chunk a relevance score, and it is tempting to treat the sorted list as finished output: concatenate top to bottom, wrap it in a prompt, send it. But the reranker answered one question — *which* chunks are worth the budget — and the prompt asks a second one — *where in the block each chunk sits*. Those are independent decisions, and conflating them is one of the most common design gaps in a RAG system that otherwise retrieves well but answers poorly. The reason position matters at all is that a language model does not read a long block of text with uniform care. Material at the very beginning of a long block and material at the very end are used more reliably than material in the middle; the middle is the weak slot. Modern long-context models flatten this effect compared with earlier generations, but as of mid-2026 it has not disappeared, and it grows with the length of the block. If you paste twenty chunks in descending score order, ranks 3 through 12 — often the bulk of your real evidence — land squarely in the weak region while a low-value tail chunk gets the strong final position. ## The two orderings people actually ship **Most-relevant-first.** Straight descending score. It is the right default when the injected block is short — three or four chunks of a few hundred tokens each — because there is no meaningful "middle" to fall into. It also produces the most predictable prompts, which matters when you are debugging. **Bookending (edge-loading).** Reorder so the ranking runs inward from both ends: rank 1 at the head, rank 2 at the tail, rank 3 second from the head, rank 4 second from the tail, and so on. The strongest evidence occupies both privileged positions, and the material you would least miss occupies the middle. This is the standard move once the block gets long, and it costs nothing at runtime — it is a list permutation applied after reranking. A worked shape: a travel-policy assistant reranks twenty chunks and has roughly a 6,000-token slot. Packing them in score order puts the top per-diem rule at position 1 and the second-best rule — often the one that actually answers the question, since the top hit is frequently a near-duplicate — around position 2 to 4, which is fine, but the corroborating exception clauses at ranks 5 to 10 are now mid-block. Bookending pulls half of that set to the tail, right before the question, where it is read most reliably. ## Within a document, reading order beats score order When several retrieved chunks are consecutive sections of the same document, score order shreds a narrative: section 4 above section 1 above section 7. Restore document order within each source and apply the relevance ordering *across* sources. Sequential material — a procedure, a numbered policy, a timeline — is much easier for the model to use when its internal order is intact, and any position benefit you gained by shuffling it is usually smaller than the coherence you lost. ## Deciding it on your corpus rather than by folklore Ordering is cheap to evaluate, so do not take the strategy on faith. Build a held-out set of questions where you know which chunk contains the answer. Hold retrieval and reranking completely fixed so the same chunks are injected every run, and vary only the permutation. Score two things: end-to-end answer correctness, and whether the model actually used the gold chunk (checkable by inserting the gold chunk at each position in turn and reading off the accuracy curve for *your* model at *your* block length). If the curve is flat, ship most-relevant-first and spend the effort elsewhere; if it dips, bookend. ## When it stops mattering Three short chunks: the dip needs distance, and there is none. A model whose answer is dominated by a single decisive passage: whichever position it lands in, it wins. A block where every chunk is near-equally relevant: permuting equals is a no-op. Recognizing these cases is as much a part of the skill as knowing the bookending trick — arguing hard for a reordering scheme on a prompt that injects 800 tokens total signals cargo-culting. ## Practical notes Keep the ordering deterministic and log it alongside the answer. When a RAG answer is wrong, the first diagnostic question is whether the right chunk was retrieved at all, and the immediate second one is where it sat; without the recorded order you cannot separate "retrieval missed it" from "retrieval found it and the prompt buried it." Also keep ordering logic in one place: a permutation applied at prompt-assembly time, not smeared across the retriever, the reranker and the template.
- How would you verify on your own corpus that bookending actually helps?Build a held-out set where you know which chunk contains the answer. Freeze retrieval and reranking so the same chunks are injected every run, and vary only the permutation. Score end-to-end correctness, and separately sweep the gold chunk through every position to read off the accuracy curve for your model at your block length. If the curve is flat, keep the simpler most-relevant-first ordering.
- What order do you use when the retrieved chunks are consecutive sections of one document?Restore document order within that document and apply relevance ordering across documents. Sequential material — a numbered procedure, a policy with exceptions, a timeline — becomes much harder to use when its internal order is shuffled, and that coherence loss usually outweighs any position gain from interleaving it by score.
- Does ordering matter when only three short chunks are retrieved?Barely. The weak-middle effect needs distance, and a block of three short chunks has no meaningful middle. Ship most-relevant-first, keep it deterministic, and spend the effort on retrieval quality instead. Reordering schemes start earning their keep when the injected block runs to thousands of tokens across many chunks.
saying these in an interview costs you the question
- Claims order cannot matter because attention sees the whole context
- Sorts by rerank score and considers the ordering question settled
- Applies bookending to three short chunks and calls it an optimization
- Confuses reranking (which chunks) with ordering (where they sit)
- Shuffles consecutive sections of one document by score, destroying its order