How do you resolve cache-friendly prompt ordering against putting instructions last?
answer
- separate the bulk from the directive
- large stable blocks stay in the head
- short restatement after the variable content
- duplication is cheaper than fragmentation
- measure quality before conceding reuse
basics
~20 sSplit the instruction from its bulk. Keep the large stable blocks — persona, tools, corpus, examples — in the cached head, and repeat only a short restatement of the task after the variable content. Duplicating a few dozen tokens beats fragmenting the prefix.
solid answer
~50 sThe tension is real: caching rewards putting everything stable at the front, while many teams' evals show the model follows a task instruction better when it sits adjacent to the input. Resolve it by separating the *bulk* from the *directive*. The heavyweight stable material stays where it is and stays cached; a compact restatement of the task and output format is appended after the variable content, where it is re-sent every request anyway and costs almost nothing. That is not a compromise so much as recognising the two blocks were never the same thing. When a genuine restructure is on the table — moving a large corpus below per-request content, for instance — decide it on measured answer quality first: caching is an optimization applied to a correct prompt, never a reason to ship a worse one. At organizational scale, write the layer order down and make "nothing new is inserted above the boundary" a review rule.
go deeper
Know that stable content goes first for caching, and that if an instruction needs to be near the user's input you can repeat a short version at the end rather than moving everything.
Be ready to distinguish the bulk of a system prompt — persona, examples, corpus — from the short task directive, and explain why only the directive needs to move to the tail.
Show that you would settle the question with an eval rather than a hunch, size both the quality delta and the reuse delta before conceding either, and refuse to ship a layout that answers worse just because it caches better.
Own the layout as a cross-team contract: documented tiers with change rates, additions appended within a tier, a review question about content above the boundary, and a named owner. Also own the revisit trigger — a model change or a shift in traffic mix makes the boundary worth re-deciding.
## The tension, stated honestly Two pressures pull the prompt in opposite directions. Prefix caching wants the maximum number of tokens to be identical from position zero, which argues for putting everything stable at the top and everything variable at the bottom. Meanwhile many teams find, in their own evals, that the model adheres better to a task directive when the directive sits close to the input it applies to — which argues for pushing instructions down, past the variable content. Both pressures are legitimate. The trap is treating this as a single either/or decision about "where the instructions go", because it hides the fact that "instructions" is not one thing. ## Bulk and directive are different blocks Unpack what usually gets bundled as the system prompt: a persona and tone specification, safety and policy rules, tool usage guidance, few-shot examples, a domain glossary, and — somewhere in there — three sentences saying what to do with this request and in what format to answer. The first group is bulk. It is large, it is stable, it is background knowledge, and nothing about it needs to be adjacent to the user's input. The last group is a directive. It is small, and it is the part whose position plausibly affects adherence. Once you see them as separate, the resolution writes itself: bulk stays at the top inside the cached region, directive is repeated at the tail after the variable content. The tail restatement is a few dozen tokens — recomputed on every request regardless, because everything after the boundary is — so it costs essentially nothing in reuse while giving you the placement your evals asked for. Duplication in the prompt feels wasteful to engineers trained on DRY, but the arithmetic here favours a small duplication over a fragmented prefix by a wide margin. ## When the tail restatement is not enough Sometimes the thing that needs to be near the input genuinely is large — a detailed rubric, a long output schema, a set of examples selected per request. Three options, in order of preference: 1. **Shrink it.** A rubric that must be adjacent is often a rubric that could be a checklist. Compress it for the tail position and keep the full version above. 2. **Split it.** The stable part of the schema lives in the head; only the request-specific selections go in the tail. 3. **Accept the loss.** If per-request examples materially improve quality and must sit before the user's question, they live outside the cached prefix by definition, and the corpus above them still caches. That is a smaller loss than people assume — the boundary just sits higher than you wanted. What you should not do is move the large stable corpus *below* the variable content to keep it near the instructions. That converts your best cacheable asset into per-request work for a placement effect you can usually get with a restatement. ## Deciding when they genuinely conflict The tie-break is quality first. Caching is an optimization over a prompt that already works; a layout that measurably degrades answers is not made acceptable by reuse. But "measurably" is doing real work in that sentence — run the two layouts against a task eval set before conceding anything, because the assumed quality difference often fails to show up, and then the cache-friendly layout wins for free. When a difference is real, size both sides before choosing. How large is the quality delta on the eval set, on which slices? How large is the reuse delta — is the block in question a thousand tokens or thirty thousand, and how many requests share it? A one-point quality gain that costs reuse on a 30k-token corpus shared by every user is a different trade from the same gain costing reuse on a 400-token header used by one workflow. ## The organizational half of the problem At any scale beyond one prompt, the real failure is not a bad decision but an unmanaged one. A shared prompt template accumulates contributions from many feature teams, and each one has a locally sensible reason to add its block near the top: it feels important, so it goes first. Six months later the header contains a date, a feature-flag dump and a per-user greeting, and the prefix is worthless. What works is treating the layout as a contract rather than a style preference. Write the layer order down with the change-rate of each tier next to it. Make additions default to appending within a tier rather than inserting above one. Put "does this add variable content above the cached boundary?" on the prompt review checklist. Give the template an owner who can say no. None of this is technically interesting, and it is the difference between a design that holds for a year and one that quietly decays within a quarter. ## Anti-patterns Letting the cache dictate a prompt that answers worse. Refusing to duplicate a small instruction on principle. Restructuring on a hunch about placement effects without an eval to back it. And treating the layout as fixed forever — when the model, the corpus size or the traffic mix changes materially, the boundary is worth revisiting.
- Is duplicating the instruction not just wasted tokens on every request?It is duplicated, but the copy sits after the cache boundary where the tokens are processed fresh anyway, and it is tiny. The alternative — moving a large stable block below the variable content so the instruction can lead it — turns thousands of reusable tokens into per-request work. A few dozen duplicated tokens against thousands of preserved ones is not a close call.
- How would you decide whether a proposed reorder is worth its cost in reuse?Run both layouts against a task eval set first, because the assumed quality difference frequently does not appear. If it does, weigh its size and which slices it affects against how many tokens and how many sharing requests the reorder costs. Reuse on a 30k-token corpus shared by every user is a very different stake from reuse on a small per-workflow header.
- What stops a shared prompt template from decaying as teams add to it?Ownership and a written layer contract. Document the tiers with their change rates, default to appending within a tier rather than inserting above one, and add "does this put variable content above the cached boundary?" to prompt review. Without that, every team's locally reasonable decision to put its block first collectively destroys the prefix.
saying these in an interview costs you the question
- Moves a large static corpus below the user input for recency
- Refuses to duplicate a short instruction on DRY grounds
- Ships a worse-answering layout because it caches better
- Reorders on a placement hunch with no eval evidence
- Treats the layout as one team's private style choice