skip to content

Before an agent starts a multi-turn run on your webhook sender, what must the request state beyond the goal?

level: middleimportance: should knowfreq 50%

answer

  1. A goal is not a brief
  2. Fence, done, and what comes back
  3. Artefacts, not sentiments
  4. Stated up front measures every turn

basics

~20 s

Three things the goal itself does not carry: the fence, naming artefacts the run must not change; what done means, in terms you could check against the diff; and the decisions it must bring back rather than settle alone.

solid answer

~50 s

A goal says what to build. It says nothing about where to stop, what may be disturbed on the way, or which forks are mine to decide — so the opening request carries three more blocks. The fence names artefacts rather than sentiments: the stored shape of a delivery record, the sender's public contract, the module another team owns. "Be careful" fences nothing, because a diff cannot be checked against it. Done is stated as something observable: a delivery rejected as malformed is never retried, an exhausted event is recorded where an operator sees it. And the decisions to bring back are named by category — anything that changes what the receiver observes, anything that adds a dependency — since I cannot enumerate the forks of a run I have not watched. Stated up front, all three measure the whole run.

code

text · 19 lines
text
GOAL
  The webhook sender retries deliveries that failed for a transient reason.

MUST NOT CHANGE
  - the stored shape of a delivery record (other services read it)
  - the sender's public contract: no renamed events, no new required fields
  - anything inside the receiver team's module

DONE MEANS (each line checkable against the diff)
  - a delivery the receiver refused as malformed is never retried
  - a delivery that got no answer is retried, with growing gaps, to a fixed bound
  - a delivery that exhausts the bound is recorded where an operator will find it
  - a repeat carries the identity of the original attempt
  - nothing in flight is lost when the process restarts

BRING BACK TO ME INSTEAD OF DECIDING
  - anything that changes what the receiver observes
  - any new dependency
  - any case where you believe one of the three fences above has to move

go deeper

for a junior

Name the three blocks and give one concrete example of each for the work in front of you. Saying what the run must not touch and what finished looks like is more useful than any amount of description of the goal itself.

for a middle

Explain why each block exists: a fence that names artefacts can be checked against a diff, an observable done removes the run's choice of stopping point, and a named category turns a silent fork into a question you get asked.

for a senior

Show the judgement in what you fence and what you leave free. Over-fencing produces a run that stops constantly and a request nobody reads; under-fencing shows up as a stored shape quietly extended in a change nobody read closely.

for a principal

Own the question of which non-negotiables are worth stating on every run at all, and what it means that a request is an instruction rather than a guarantee. Where a boundary genuinely must hold, something other than wording has to hold it.

## A goal is not a brief *Add a retry policy to the webhook sender* tells a run what to produce. For one bounded edit that can be enough: you see the result immediately and the blast radius is a file. A run that will take many turns differs in one specific way — **by the time you read any of it, later turns have been built on earlier decisions you never saw made.** So the opening request carries three blocks the goal itself does not. ## 1. The fence, named as artefacts "Be careful" and "don't break anything" fence nothing. They name no object, so no diff can be checked against them and no reader can tell whether they were honoured. A usable fence names artefacts: - **the stored shape of a delivery record** — other services read it, so a new required field is not yours to add here; - **the sender's public contract** — no renamed events, no new required fields in what the receiver is shown; - **a module another team owns** — a change there is their review, not a line in your diff. Write it so that a breach is **visible in the diff**, because that is how you will actually check it. Then pair the fence with a way to ask: a boundary the run must either obey or silently cross is worse than one that says *if you believe this has to change, stop and tell me*. A stated fence is an instruction the run is asked to honour rather than a mechanism that prevents the change — what actually enforces a boundary is a separate subject — and the escape hatch is what keeps an honest run from quietly choosing between your rule and the task. ## 2. "Done", stated as something observable Left unstated, done is whatever the run decides looks finished, and on a long run that is a real decision with no owner. Stated well, it is a short list you could hold beside the diff: 1. a delivery the receiver refused as malformed is **not** retried; 2. a delivery that got no answer is retried, with growing gaps, to a fixed bound; 3. a delivery that exhausts the bound is recorded where an operator will find it; 4. a repeat carries the identity of the original attempt; 5. nothing in flight is lost when the process restarts. Each line is checkable by reading, which is the property that matters. "Add retries, and make it robust" is not a weaker version of this list; it is a different kind of sentence, one that cannot be failed. ## 3. The decisions you want brought back rather than made Across many turns a run meets forks the request never mentioned: a dependency that would make the work easier, a field added to a stored record, a change to what the receiver observes. Each is a decision, and **a fork you did not name is a decision made silently and discovered in the diff.** Name them by **category**, not by instance: - anything that changes what the receiver observes; - anything that adds a dependency; - anything that would require changing the stored shape you fenced. You cannot enumerate the forks of a run you have not watched yet. Categories survive contact with the fork you failed to foresee, which is precisely the one that would have cost you. ## What each block buys | what the request carries | what it settles | what you find out from the diff instead | |---|---|---| | the goal alone | what to build | everything else, after it is built | | + the fence | which artefacts stay as they are | a shared record shape quietly extended | | + observable done | when the work is finished | a run that carried on into work you never asked for | | + decisions to bring back | which forks are yours | a dependency added, or a contract changed, without a word | ## Why up front rather than part-way through A constraint in the opening request measures **every turn**, including the ones you never read. The same constraint added at turn twelve measures only what comes after it, and hands you a second job: deciding what earlier work to revisit in the light of it. That asymmetry — not tidiness — is why the three blocks go in at the start. It is also why they are worth writing even when you expect to watch the run: watching tells you what is happening now, while the blocks are what the parts you skim are graded against.

  • How do you write a fence you can actually check afterwards?
    Name artefacts, not attitudes: this stored shape, this contract, that module. Then ask whether a breach would be visible in the diff — if you could not point at the line that broke it, the fence is a sentiment and the run will be graded on nothing.
  • Why name the decisions to bring back by category rather than listing them?
    Because you cannot enumerate the forks of a run you have not watched. A list covers the decisions you already foresaw, which are the ones you would have caught anyway; a category such as "anything that changes what the receiver observes" catches the fork you did not imagine.
  • What if stating done this precisely takes longer than the change itself?
    Then the change is small enough that you will read all of it, and the case for the blocks is weaker. They earn their cost on work spanning many turns, where you will skim most of the result and need something the skimmed parts were measured against.

saying these in an interview costs you the question

  • Say be careful and thorough; the model reads tone
  • Done is obvious from the goal and needs no statement
  • You can list up front every decision the run will face
  • Constraints are just as good added once you see the run drifting
  • A fence written into the request guarantees the run will not cross it