skip to content

A check failed after an agent's change — why hand back the run's own output rather than your summary of it?

level: juniorimportance: must knowfreq 70%

answer

  1. Two things could go back
  2. Evidence, or your conclusion about it
  3. Expected, actual, and the failing case
  4. The size of the gap is the diagnosis
  5. Verbatim means selected, not everything

basics

~20 s

The run's output carries the failing case, the expected and actual values, and the line that produced them. A summary replaces those facts with your diagnosis, and if the diagnosis is wrong the next turn inherits it.

solid answer

~50 s

Two different things can go back into the session: what the run printed, or what you concluded from it. The printed text is evidence — the case that failed, the value expected against the value produced, the frames that reached it, any rule identifier the tool attached. Your restatement is a hypothesis, and a loop is running precisely because nothing is yet understood well enough to hypothesise safely. The raw text also carries identifiers the agent can search the codebase for, which a paraphrase usually strips out. Two qualifications matter. Verbatim means *selected*, not *everything*: one failing case with its assertion and its own frames steers well, while a wall of repeated failures and environment noise mostly displaces the context that was helping. And when you do hold a real hypothesis, send it beside the output and labelled as a guess, rather than in place of it.

code

text · 16 lines
text
WHAT THE RUN PRINTED  (hand this back)

  FAIL  invoice/perLineRounding_halfUp
        expected  1204.80
        actual    1204.60
        at  billing.Money.round             line 61
        at  billing.InvoiceTotals.monthly   line 118
        39 other cases passed

WHAT A SUMMARY SAYS INSTEAD  (this is a guess, not evidence)

  "the rounding is still wrong in the invoice total"

The gap is 0.20 across 40 lines - half a hundredth per line.
That points at the per-line step, not at the total. The summary
throws that away and sends the agent to the wrong one of the two.

go deeper

for a junior

Know to paste what the run actually printed — failing case, expected against actual, the lines that produced it — and to keep your own explanation separate from it.

for a middle

Explain the mechanism: output is evidence, a restatement is a hypothesis, and a wrong hypothesis stated as fact steers the following turn and costs it.

for a senior

Show how you select from a noisy run: one actionable failure per turn, frames inside your own code, rule identifiers over prose messages, and what you do when the run printed nothing.

for a principal

Own the judgement that a loop's input quality is a team habit, not a personal one, and that an unreadable failure output is a defect worth fixing before it costs turns repeatedly.

## The two messages you could send A check goes red after an agent's change. There are two quite different things you could put into the next turn, and they are not two styles of the same message. - **What the run printed.** The name of the case that failed, the value it expected against the value it produced, the frames that led to the line, and any rule identifier the tool attached to the complaint. - **What you concluded from it.** "The rounding is still wrong in the invoice total." The first is **evidence**. The second is a **hypothesis about the evidence** — and a loop is running because something is not yet understood, which is exactly the situation in which your hypothesis is most likely to be the wrong part. ## What a summary drops Take a billing calculator with forty test cases, of which exactly one catches a rounding fault: an invoice of forty identical lines, priced so that rounding each line and rounding only the final total give different answers. The run prints an expected total of `1204.80` against an actual of `1204.60`, on a case named for per-line rounding. | carried by the raw output | survives "the rounding is wrong" | |---|---| | the name of the case that failed | usually not | | expected against actual, exactly | rarely — it becomes "off by a bit" | | the size and sign of the gap | no, and here the size *is* the diagnosis | | the frame that produced the value | no | | which of the other cases passed | no | | identifiers the agent can search for | no | The gap is the part people underrate. Twenty hundredths spread over forty lines is **half a hundredth per line**, which points at rounding applied once at the end instead of per line; a flat "the rounding is wrong" leaves the agent to choose between those two readings, and it will choose. Handing over the number hands over the diagnosis for free. ## Why a wrong summary costs more than no summary A prose restatement arrives as a statement of fact about the code, not as a guess, and the turn that follows builds on it. If you say the fault is in the total and it is in the per-line step, the agent will now look at, and change, the total — and the case will still fail, on a number that has moved. You have spent a turn and made the next signal harder to read. This is the ordinary way a loop starts wasting turns, and it is not a failure of the tool. It is a wrong input, supplied confidently. ## Verbatim, but selected "Paste the output" is not "paste everything." A run can print the same failure many times, plus banners, timings, and dependency noise. All of that competes for the same limited space as the code the agent needs to hold while it works. A good hand-back is usually: 1. **One failing case**, named, with its expected and actual values. 2. **The frames inside code you own**, trimmed of framework interior. 3. **The rule identifier**, where the complaint came from a linter or a type check, since the identifier is searchable and the prose message often is not. 4. **A pointer, for yourself**, to where the full log sits. If several checks are red at once, one actionable failure per turn reads better than five, because a turn carrying five failures produces a change addressing whichever one was easiest — and you are then left inferring which of the five the change was even about. ## When your hypothesis is worth sending Often it is. You may have seen this failure before, or you may know the module. The rule is about **placement, not suppression**: put the output first, then your reading of it, marked as a reading — "the gap looks like half a hundredth per line, which would mean the rounding happens once at the end." The agent can then disagree with you against the evidence, which it cannot do when the evidence never arrived. ## What this does not claim - **Not that prose is useless.** Constraints, intent and what you already ruled out are prose, and they belong in the turn. - **Not that raw output is always enough.** Some failures print almost nothing useful — a process that stops, a check that times out — and then what the run *did not* print is the fact you carry over. - **Not that more text is better.** The loop's scarce resource is attention over a limited context, and repeated noise spends it. - **Not that you are always the one relaying it.** Where the agent runs the check itself, the same distinction applies one level in: what steers the next attempt is the output, not the run's own summary of the output. The durable form of the rule: **hand back what was observed, and keep what you concluded clearly separate from it.**

  • The failing run prints hundreds of lines. What do you actually paste?
    The failing case with its expected and actual values, the frames inside code you own, and any rule identifier. Drop repeats, banners and framework interior. Aim for one actionable failure per turn: a turn carrying five failures produces a change addressing whichever was easiest, and afterwards you are left inferring which it was even about.
  • Is your own reading of the failure ever worth sending?
    Often, but beside the output and labelled as a reading rather than a fact. The cost only appears when it replaces the evidence, because the agent will build on a stated diagnosis instead of contradicting it — and it has nothing left to contradict it with.
  • What do you hand back when the run printed nothing useful at all?
    The absence is the signal: what you ran, where it stopped, how long it took, and what was last printed. A check that hangs or dies without a message is a weaker signal than a failed assertion, and the honest next move is usually to make the failure louder before asking for another fix.

saying these in an interview costs you the question

  • My description of the failure is better input than the raw output
  • Give the agent your diagnosis instead of the failure text; it saves a turn
  • Paste the whole log every turn — more context is always better
  • Saying it still fails is enough; the agent knows what it wrote
  • Exact expected and actual numbers are detail the model does not need