How do you standardize section order across a library of thirty production prompts?
answer
- a template, not thirty documents
- slots you fill, not prose you write
- one variable changes, evals attribute
- invariant first, volatile last
- version it, lint it, allow exceptions
basics
~20 sFix one skeleton — role, rules, tools, data, task — and make every prompt fill its slots rather than invent an order. The gain is reviewability, clean diffs, eval attribution and a stable shared prefix; the cost is ceremony, so keep a documented escape hatch.
solid answer
~50 sTreat the skeleton as an interface. Pick one order — typically role, then standing rules, then available tools, then delimited data, then the specific task — and express it as a template with named required and optional slots, so a prompt is authored by filling slots, not by writing a document. Four things pay for this. Review becomes possible, because a reviewer knows where to look for the constraint they care about. Diffs become meaningful, since a change to the rules block does not reflow the whole string. Evals become attributable: when quality moves, you know which slot changed. And an identical leading section across prompts is a stable prefix, which matters for reuse of unchanged context. The costs are real — tokens spent on structure, and the temptation to force an unusual task into the shape. So version the template, lint for slot presence and order in CI, and allow documented deviations rather than pretending none are needed.
code
markdown · 18 lines<role>
You review incident postmortems for a logistics platform.
</role>
<rules>
- Never state a root cause the report does not support.
- Flag missing timelines rather than inferring them.
</rules>
<data>
<postmortem id="inc-2291">
...report text...
</postmortem>
</data>
<task>
List the action items that have no named owner.
</task>go deeper
Know that production prompts are usually built from a fixed set of named sections rather than written as free prose, and that data blocks are delimited and named within that shape.
Explain the concrete order — role, rules, tools, data, task — and why the specific instruction sits after the data. Be able to say what a template plus one assembly function buys over hand-written strings.
Show the engineering benefits: localised diffs, single-variable eval changes, one place where untrusted spans are wrapped, and a CI lint that catches drift. Be candid that the quality effects of ordering are tendencies to be measured, not laws.
Own the skeleton as a versioned interface with a deprecation path, an eval gate on template changes, and a documented escape hatch. Be ready to argue when the ceremony is not worth it and when a high-traffic prompt should be tuned outside the standard.
## What a skeleton is A skeleton is a fixed sequence of named sections that every prompt in a library fills. A typical one: **role** (who the model is acting as), **rules** (standing constraints that hold regardless of the request), **tools** (what is callable, if anything), **data** (delimited, named blocks of retrieved or user-supplied content), **task** (the specific thing to do this time), and optionally a **format** slot for how the answer should be shaped. The content of each slot varies per prompt. The order does not. ## Why the order stability is the point **Reviewability.** Thirty free-form prompts are thirty documents a reviewer must read end to end. Thirty slotted prompts are thirty forms: a reviewer checking that no prompt permits a refund above a threshold reads one slot in each. This is the single largest practical benefit and the one that shows up in incident response. **Diffs that mean something.** When prompts are prose, an edit reflows the string and the diff is unreadable. When they are slots, a change is localised, and a reviewer can tell a rules change from a task rewording without reading the surrounding text. **Eval attribution.** With a fixed shape you can hold every slot constant and vary one — swap the rules block, keep everything else — and attribute the score movement. Without it, every prompt edit is a multi-variable change and the eval tells you only that something moved. **Prefix stability.** Sections that are identical across requests sit at the front, unchanged, request after request. That is exactly the property that lets an unchanged leading segment be reused rather than reprocessed. Ordering volatile content — the user's data, the task — after the invariant content is what makes the invariant part actually invariant. Note this is a consequence of the ordering discipline, not its justification; the reviewability argument holds even where no reuse mechanism exists. **Onboarding and authoring speed.** A new engineer writing prompt thirty-one copies the template and fills it. Nobody relitigates where the rules go. ## Where the order genuinely matters to output quality Two effects are worth stating rather than assuming. First, putting the specific task **after** the delimited data means the model reads the instruction with the material already in context, which tends to help on long inputs — and it also puts the instruction at the end, where models generally attend well. Second, standing rules ahead of the data mean the constraint is established before untrusted content appears, which is the ordering most teams settle on for that reason. Both are tendencies, not laws. If your evals say otherwise for your workload, follow the evals — but change the template, not one prompt. ## The costs, honestly Structure costs tokens on every call: headings, tags, and slot labels for sections that may be two lines long. It also creates a pull toward uniformity that does not always fit — a one-shot classification prompt with no tools and no data does not need five sections, and forcing it into the shape adds noise the model has to read past. And a rigid template can ossify: teams keep a slot long after its content became meaningless because removing it means touching thirty files. ## Governance that makes it stick - **A versioned template** with slots marked required or optional, and a changelog. Bumping the template version is a deliberate act with an eval run attached. - **A lint in CI** that asserts each prompt's slots are present, correctly named and in order. This is cheap and catches drift immediately. - **One assembly function.** Prompts should be built by code that walks the slots, not by string concatenation at each call site. That is also where delimiter handling for untrusted spans belongs. - **A golden eval set per prompt**, run when the template changes, so a library-wide edit cannot silently regress one consumer. - **A documented escape hatch.** Let a prompt deviate, but require the deviation to be recorded with a reason. Undocumented deviation is drift; documented deviation is a design decision and often the seed of the next template version. ## What to say in an interview The strong answer frames the skeleton as an interface with owners, versioning and tests — the same discipline you would apply to a shared library — and is candid that the benefits are mostly engineering benefits (review, diff, attribution, reuse) rather than a claim that the model performs better because sections are in a particular order.
- What goes wrong if each team orders sections however they like?Review degrades first: nobody can find a given constraint without reading whole prompts, so constraint drift goes unnoticed until an incident. Then eval attribution breaks, because every change is multi-variable. Then any reuse of an unchanged leading segment disappears, since no two prompts and often no two revisions share a prefix. The cost is organisational long before it is a quality cost.
- Which slot would you put last, and why?The specific task. Placing it after the delimited data means the model reads the instruction with the material already in context, and end-of-prompt instructions are generally well attended to. It also keeps the volatile part at the tail, leaving the role, rules and tool sections identical across requests — which is what makes the leading segment genuinely stable.
- How do you retire a slot that thirty prompts still carry?Same as deprecating a field in a shared library: mark it optional in the template, add a lint warning, migrate consumers in batches with the golden eval set run per batch, then make its presence an error and delete it. The cost of doing this badly is why you resist adding speculative slots in the first place.
- When is a shared skeleton the wrong call?When the library is small and heterogeneous — five prompts doing unrelated one-shot tasks gain little and pay full ceremony cost. It is also wrong when a single prompt dominates traffic and cost, since that one is worth hand-tuning outside the template. The skeleton earns its keep at the point where nobody can hold the whole library in their head.
saying these in an interview costs you the question
- Claims a fixed order improves model quality by itself
- Lets each prompt invent its own section order
- Builds prompts by string concatenation at each call site
- Adds slots speculatively and never retires them
- Changes the template with no eval run behind it