In an LLM prompt containing a long document, where should the instruction go?
answer
- position is a design choice, not cosmetic
- two privileged slots in a prompt
- the model generates right after the last token
- instruction after the document by default
- state at the top, restate at the bottom
basics
~20 sPut the instruction after the document, or at both ends. Language models attend most reliably to the beginning and the end of a prompt, so an instruction stated once before a long body is the easiest part to miss.
solid answer
~40 sPlacement is a real design decision, not cosmetics. Decoder-only models show **primacy and recency bias**: content at the very start and the very end of the prompt influences the output far more reliably than content sitting in the middle. With a long document, the safest default is to state the task **after** the document — the instruction then lands in the recency zone, immediately before generation. If the instruction is long or has a strict output contract, state it briefly up front (so the model reads the document knowing what it is looking for) and restate the contract at the end. Few-shot examples usually go after the task description and before the final input, so the last thing the model sees is the actual work item, not another example.
go deeper
Know that models attend most reliably to the start and the end of a prompt, and say plainly that with a long document the instruction should come after it, or at both ends.
Be ready to explain the mechanics: causal attention favours the earliest tokens, positional decay favours the nearest ones, and instruction-tuning data puts tasks at the edges. Explain where examples and output contracts belong and why.
Expect to be asked how you would confirm a placement change helped. Demonstrate an A/B with a programmatic compliance metric, and show judgement about which element deserves the recency slot given its length and decisiveness.
Own the ordering as a contract across a whole prompt fleet. Argue for a house template that fixes where framing, data, examples and output contracts sit, so placement is not re-decided per feature and regressions are attributable.
## The claim Where you put text inside a prompt changes whether the model actually uses it. Two positions are privileged: the very beginning (**primacy**) and the very end (**recency**). Material in the middle of a long prompt is measurably less likely to influence the answer, even though it is fully inside the context window and the model can quote it if asked directly. This is not superstition or prompt folklore. It shows up as a reproducible accuracy curve in controlled experiments where the only variable changed is the position of the needed information, and it has plausible mechanical causes in how decoder-only transformers attend. ## Why the two ends win Three effects stack up. **Attention sinks.** In causal (left-to-right) attention, the earliest tokens are visible to every later token and tend to accumulate disproportionate attention mass. The opening of a prompt is structurally hard to ignore. **Recency.** The next token is generated immediately after the last token of the prompt. Positional encodings in common use decay attention with distance, so the tail of the prompt is the closest, cheapest thing to condition on. **Training distribution.** Instruction-tuning data overwhelmingly places the task statement at the start or the end of an example, rarely buried inside a wall of text. The model learned where instructions usually live. ## The practical ordering A workable default for a single-turn prompt over a long body of text: 1. **Role and short task framing** (one or two lines) — primacy slot; tells the model how to read what follows. 2. **The document(s)**, clearly delimited so the model knows where the data starts and stops. 3. **Few-shot examples**, if any. 4. **The full instruction and output contract** — recency slot, the last thing before generation. 5. **The specific input to act on**, if it is separate from the document. The pattern of stating the task briefly at the top and restating the output contract at the bottom is sometimes called an instruction sandwich. It costs a duplicate of the instruction in tokens on every call — usually tens to a couple hundred tokens — which is cheap relative to the document it is protecting. ## Where examples go Few-shot examples are content, and they obey the same rule. Placing twenty examples *before* the task description means the model reads them without knowing what they are demonstrating; placing them *after* the task description gives them a purpose. Either way, do not let an example be the final thing in the prompt — the model may continue the pattern of producing another example rather than answering. End on the real input, or on a short cue like a restated output format. There is a second-order effect worth knowing: with a large example set, the ones nearest the end exert more influence than the ones in the middle. If your examples are not all equally representative, the most typical ones belong closest to the task. ## What this does not mean It does not mean middle content is invisible. Ask a direct question about a fact in the middle and a strong model will usually retrieve it. The failure mode is subtler: the model does not *spontaneously* apply a mid-prompt constraint while doing something else. Output-format instructions, tone constraints, forbidden-action rules and 'always cite the source' style requirements are exactly the kind of thing that quietly stops being honoured when buried. It also does not mean placement replaces content quality. If the instruction is ambiguous, moving it to the end makes it reliably ambiguous. ## Verifying rather than believing Placement effects are model-specific and prompt-specific. The honest way to settle an argument about ordering is a small A/B: same inputs, same model, same temperature, two orderings, and a measurable compliance signal — did the output match the required schema, did it use the required template, did it cite. Twenty to fifty cases is usually enough to see a real placement effect, because these effects tend to be large when they exist rather than marginal. ## Rules of thumb - Long data plus short instruction: instruction last. - Long instruction plus short data: instruction first, data last. - Strict output contract with a long prompt either way: state it at the top, restate it at the bottom. - Never let the prompt end on an example or on filler; end on what you want the model to act on.
- Where would you place twenty few-shot examples relative to the task description, and why?After the task description, so the model reads them already knowing what they demonstrate, and never as the final block. Ending on an example invites the model to continue the pattern and emit another example instead of answering. Put the real input last. If the examples vary in quality, the most representative ones belong closest to the end, since later examples exert more influence than middle ones.
- When would you put the instruction first instead of last?When the instruction is long relative to the data, or when reading the data correctly depends on knowing the task — for example, extracting only three named fields from a table. Then the instruction leads and the short input trails. The general rule is that whichever element is shorter and more decisive should occupy the recency slot immediately before generation.
- How would you prove that a reordering actually helped rather than just felt better?Run a small A/B on fixed inputs: same model, same temperature, two prompt orderings, and a programmatically checkable compliance signal such as schema validity, template match or citation presence. Twenty to fifty cases usually suffices, because placement effects are large when real. Reporting a single cherry-picked output is not evidence.
saying these in an interview costs you the question
- Claims position does not matter if content is in the window
- Buries the output format in the middle of a long prompt
- Ends the prompt with a few-shot example rather than the input
- Treats ordering advice as vendor-specific folklore
- Assumes a bigger window removes the need to order anything