A 40-page RFP review ignores the output template stated at the top — how do you fix it?
answer
- distinguish three causes before fixing
- one fix costs tokens, one costs nothing
- the last thing before generation wins
- shrink what sits between the two ends
- score compliance, do not eyeball it
basics
~20 sTwo levers exist: restate the template at the end of the prompt, or reorder so the instruction follows the document instead of preceding it. Restating costs a duplicate instruction per call; reordering is free but changes the prompt shape. Shrinking the document reduces the exposure behind both.
solid answer
~50 sAn instruction that sits at token 3 of a 40-page prompt is competing with everything that came after it by the time generation starts, and formatting contracts are the first thing to be silently dropped. **Reordering** is the free fix: move the template below the document so it occupies the recency slot immediately before generation. **Restating** is the cheap fix when the instruction must also frame how the document is read — keep a short version up top and repeat the exact output contract at the bottom, paying a duplicate of maybe 100 to 300 tokens on every call. The third lever is **shrinking the middle**: if only twelve of the forty pages are relevant, selecting them removes the low-influence bulk that caused the problem. Whichever you pick, validate with a compliance metric — schema or template match rate over a fixed set — not by eyeballing one output.
go deeper
Know the first move: put the output template at the end of the prompt, right before the model answers, rather than only at the top of a long document.
Explain both fixes and what each costs — reordering is free but changes framing, restating duplicates tokens on every call — and say why the restatement must be verbatim rather than paraphrased.
Demonstrate diagnosis before repair: separate position failure from ambiguity and from an in-document format conflict using a short-input test, then validate the chosen fix with a programmatic compliance rate over a fixed set.
Own the durable answer: make format contracts machine-checkable so violations become caught retries rather than shipped defects, and set a house policy on prompt shape so this failure is designed out rather than rediscovered per feature.
## Reading the symptom correctly 'The model ignores the template' can mean several different things, and the fixes diverge. Establish which one you have before touching the prompt. - **Position failure.** The instruction is unambiguous and the model follows it on a short input, but drifts on the long one. Compliance degrades as the document grows. - **Ambiguity failure.** The template is under-specified, and the model is filling gaps plausibly. Compliance is bad even on a two-page input. - **Conflict failure.** The document itself contains a competing format — the RFP shows its own response layout — and the model follows the nearer, more concrete example instead of your instruction. Only the first is an ordering problem. The cheap discriminator is to run the same instruction against a truncated document: if compliance is fine at two pages and collapses at forty, position is implicated. If it fails at both, reordering will not save you. ## Lever one: reorder Move the output contract so it is the last thing before generation. This costs nothing in tokens and is usually the largest single improvement, because the recency slot is the most influential position in the prompt. The cost is not zero in other senses. The prompt's shape changes, which matters if anything downstream depends on a stable prompt structure, and it can make the prompt harder for a human to read, since the reader now meets forty pages before learning what they are for. It also removes the framing benefit: a model that reads the document already knowing it must extract nine specific sections reads it differently from one that discovers the task afterwards. ## Lever two: restate Keep a compact framing instruction at the top — 'you are reviewing an RFP; you will produce a compliance response using the client template' — and put the **full, literal output contract** at the bottom. This is the instruction sandwich. It buys both the framing benefit and the recency benefit. The cost is a duplicated block on every call. For a template of a few hundred tokens against a 40-page document of perhaps 25,000 tokens, that is roughly one percent overhead — trivially worth it if it converts a 60 percent compliance rate into a 95 percent one. It stops being trivial when the instruction is itself enormous, or when the restatement is repeated many times through the body, which is a common overcorrection: interleaving the instruction every few pages fragments the document, multiplies cost, and rarely beats one good restatement at the end. One detail matters: restate the contract *verbatim*, not paraphrased. Two slightly different versions of the same rule inside one window is a contradiction the model has to resolve, and it may resolve it by splitting the difference. ## Lever three: shrink what sits between The underlying problem is the mass of low-influence material separating the instruction from the output. Reducing it helps regardless of ordering — selecting the sections of the RFP that actually bear on the response, or processing the document in sections and assembling the results, both shrink the distance the instruction has to survive. This is more work than reordering and changes the system's shape, so it is the right move when the position sensitivity is severe rather than marginal. ## Making the constraint checkable A formatting contract that a program can verify is worth far more than one only a human can judge. If the template maps to a schema, ask for structured output and validate it; a violation then becomes a caught error and a retry rather than a silent defect discovered by a client. Position bias does not disappear, but its blast radius shrinks from 'wrong document shipped' to 'one retry'. A related trick is asking the model to emit the contract's key constraints before the body of the answer. That forces the constraint into the model's own recent context, where it exerts maximum influence over what follows. It costs a few output tokens and is worth trying when reordering alone is not enough. ## Proving the fix Build a fixed set of representative documents — twenty to fifty is normally enough — and score template compliance programmatically: required sections present, order correct, headings matching, no extra prose. Compare the original ordering against each candidate fix on the same set at the same temperature. Report the compliance rate, not an anecdote. Placement effects tend to be large enough that a set of this size resolves them clearly, and having the number lets you decide whether the token cost of restating is justified rather than arguing about it. ## What not to do Do not respond to a position failure by escalating the language — ALL CAPS, 'IMPORTANT', threats of consequences. It occasionally helps a little, it is unmeasurable, and it leaves the actual structural problem in place. Do not scatter the instruction throughout the document. And do not conclude the model is incapable of following the template before you have tried it on a short input.
- How would you tell this apart from an instruction that is simply too vague?Run the identical instruction against a two-page excerpt. If compliance is high on the short input and collapses on the full document, the instruction is clear and position is the problem. If it fails on both, the contract is under-specified and no amount of reordering fixes it. This single test costs one call and prevents the most common wasted fix.
- What is the token cost of the sandwich, and when does it stop being worth paying?You pay a duplicate of the output contract on every call — often 100 to 300 tokens against a document of tens of thousands, so roughly one percent. It stops being worth it when the instruction block is itself very large, when the surrounding document is short enough that recency already covers it, or when a checkable structured-output contract achieves the same compliance without duplication.
- Would repeating the instruction every few pages be better than once at the end?Almost never. Interleaving fragments the document, multiplies token cost linearly with document length, and introduces many near-copies of the same rule that the model must reconcile. One verbatim restatement in the recency slot captures nearly all the benefit. If one restatement is not enough, the real answer is usually to shrink the document rather than to repeat harder.
- The RFP itself contains a response layout that differs from the client's template. How does that change your diagnosis?That is a conflict failure, not a position failure. The model is following a concrete in-document example over your abstract instruction. Reordering may help at the margin, but the durable fixes are to state explicitly that the document's own layout must be ignored, to strip or clearly fence that section, and to make the required template concrete enough to outweigh the example.
saying these in an interview costs you the question
- Escalates to ALL CAPS instead of moving the instruction
- Repeats the instruction on every page of the document
- Blames model capability without testing a short input
- Paraphrases the restatement instead of repeating it verbatim
- Declares the fix successful from one good-looking output