Before an agent starts iterating on a change, what makes a check usable as the loop's steering signal?
answer
- Decide it before the first turn
- Runs without you, same answer twice
- Cheap enough to run every turn
- Names what failed, not that something did
- Already green gives no direction
basics
~20 sA usable verdict runs without you, returns the same answer on the same input, is cheap enough to run every turn, and names what failed rather than that something did — and it has to be failing now, for the right reason.
solid answer
~50 sA loop is only as good as the verdict it turns on, so decide before the first turn whether you have one. A usable signal runs unattended, is deterministic on the same input, is cheap enough that running it every turn is not a decision, and is **specific**: a named case with an expected and an actual value localises a fault, while a red suite only proves that one exists. It also has to be failing right now, and failing for the thing you are chasing — a check that is already green gives the next turn no direction. If no such check exists you are not in a signal-driven loop; you are the verdict, and every turn now costs a human read, which argues for fewer turns and closer supervision. For some subjects — wording, a naming choice — there is rarely a mechanical verdict worth steering on.
go deeper
Know that a loop needs something that says pass or fail without you, and that a check which is already green gives the next change no direction.
Explain the properties that make a verdict usable — unattended, deterministic, cheap per turn, and specific enough to localise — and why a flaky check is worse than none.
Show the judgement call when nothing mechanical exists: build the cheapest check, supervise a short loop, or decline to loop, and say what each costs.
Own which classes of work your team is willing to run unattended at all, and treat an unreliable or unreadable verdict as a defect that costs turns on every loop that uses it.
## The question to settle before turn one A loop steers on a verdict. Whether one exists, and what it is worth, is decided before the agent starts — not discovered three turns in. The test is not "do we have tests?" but "is there something that will tell us, every turn, without us, whether we are closer?" ## Five properties of a signal that can steer 1. **Automatic.** It runs without a person looking at anything. The moment a human has to inspect output to decide pass or fail, each turn costs a human read, and the loop's economics change completely. 2. **Deterministic.** The same input gives the same verdict. A check that passes and fails on unchanged code is worse than no check: the loop will attribute the change in verdict to the last edit, and the agent will "fix" noise, producing real edits that address nothing. 3. **Cheap enough to run every turn.** If running it is a decision, it will be skipped, and the turns between runs are guesses. 4. **Specific.** A named case with an expected and an actual value says *where*; a red suite says only *that*. In a billing calculator where one case out of forty catches a rounding fault, "one failed" and "per-line rounding expected `1204.80`, got `1204.60`" are the same verdict carrying wildly different amounts of steering. 5. **Failing now, for the right reason.** A green check gives the next turn no direction — it can only guard what already works — and a check failing for an unrelated reason steers you into the unrelated thing. Confirm that the failure you are about to loop on is the failure you care about before handing it over. | signal | automatic | deterministic | specific | usable as a loop's verdict | |---|---|---|---|---| | a named failing case with expected and actual | yes | usually | high | yes | | a type or compile error | yes | yes | high | yes | | a lint rule identifier on a line | yes | yes | high | yes, for what the rule covers | | a whole suite reporting only red or green | yes | usually | low | weakly — it cannot localise | | a timing-dependent or order-dependent case | yes | no | varies | no, repair it first | | a person looking at the result | no | no | varies | not a loop signal; you are the oracle | ## The case with no signal Plenty of real work has no mechanical verdict: making a message clearer, choosing names, judging whether a layout reads well. Three honest responses, and the choice between them is the actual skill: - **Build the cheapest check that expresses the rule, first.** For the billing fault, one case that fails on the disputed total is worth more than any amount of describing the fault. What that check should contain and how it gets written is a separate subject; here the point is only that the loop cannot start without it. - **Run a short supervised loop and accept that you are the verdict.** Legitimate, but price it honestly: every turn now costs your attention, so the sensible shape is few turns and small units rather than a long unattended run. - **Do not run a loop at all.** If nothing mechanical can judge the result and you cannot judge it quickly either, iterating just re-rolls the answer and the last version is not reliably the best one. ## Why specificity outranks coverage here For steering, a signal that names one failing case beats a broader one that reports only a colour. Breadth tells you how much is wrong; specificity tells the next turn where to go. This is why an error naming an identifier and a line is often a better loop signal than a suite result, even though the suite checks far more: the identifier is searchable and the colour is not. That does not mean narrow checks are better checks. It means the property a loop needs from a check — **localisation** — is not the property that makes a check valuable overall. ## Signals people mistake for verdicts - **"It ran without errors."** Absence of a complaint is not a verdict on behaviour; nothing asserted anything. - **"The agent says it is fixed."** That is a report from the thing under test. - **A flaky case.** It produces verdicts, just not ones that carry information — and a loop consumes verdicts indiscriminately. - **A check the change itself introduced this turn.** A new check written alongside the fix tends to encode the same belief the fix does, so it can pass while the original fault stands. ## What this does not claim - **Not that a subject without a mechanical check is off limits** to an agent. It is off limits to an *unattended loop*. - **Not that tests are the only usable signal.** Type errors, compile failures, lint rule identifiers and a reproducible runtime failure all steer, sometimes better. - **Not that one green check means done.** The verdict is only as broad as what it asserts, and what it asserts was chosen by somebody.
- The only failing check is flaky. Is that still a usable loop signal?No — repair it before looping. A verdict that changes on unchanged code will be attributed to the last edit, so the agent makes real changes chasing noise, and a single verdict cannot tell a fix from a coin flip. Pinning the source of variation is the first turn, not a detour from it.
- How specific does a signal have to be before it is worth looping on?Enough to name a place. A case name with expected and actual values, a rule identifier on a line, or a frame in your own code all qualify. A bare red verdict tells the next turn only that something is wrong, which it already assumed.
- Can you loop on a check the agent writes itself at the start of the run?With care, and knowing what it does not prove. A check written in the same run as the fix tends to encode the same belief as the fix, so it can pass while the original fault stands. Before it becomes the loop's verdict it has to have failed, for the reason you are chasing. How such a check gets written well is a separate subject.
saying these in an interview costs you the question
- Any failing test will do as a signal; a red build is a red build
- You can always just ask the agent whether it fixed it
- A flaky check is fine as a signal if it usually fails
- If the code compiles and runs, the loop has a verdict
- Work with nothing mechanical to check is simply unsuitable for an agent