skip to content

Beyond the change itself, what does a long autonomous agent run cost you that a small assisted edit does not?

level: seniorimportance: should knowfreq 44%

answer

  1. The diff is the cheap part
  2. Count the human decisions, not the lines
  3. How far could a mistake travel?
  4. Could you explain it in six months?
  5. Would a second run agree?

basics

~20 s

A long autonomous run concentrates a large change into one review decision, widens what a mistake can touch before anyone notices, leaves less of the session reconstructable, and produces a result a re-run will not reproduce.

solid answer

~40 s

The diff is the cheap part. **Review surface**: the same number of changed lines arriving as one proposal carries far more decisions per human glance than the same lines arriving a few at a time, and attention is a fixed budget spread over whatever shows up. **Blast radius**: how far a mistake travels before anyone notices is set by what the run could reach, not by how wrong it was. **Reconstructability**: if the run left only a final change, you have the what and not the why, which matters when someone asks in six months why a rule was rewritten. **Reproducibility**: re-issuing the same request generally will not give the same change, so a run is not a trial you can repeat. None of that shows up in a green build.

go deeper

for a junior

Know that a bigger automated change is not automatically a bigger win. Somebody still has to decide it is right, and that decision does not get easier as the change grows.

for a middle

Be able to separate the size of a change from the number of decisions it forces, and to say why a run that left no trail is harder to defend later than one that did.

for a senior

Show that you price a run before starting it: how far its reach goes, what you will be able to reconstruct, and whether the team can absorb the review it will produce.

for a principal

Own the standard for what a run must leave behind when its change is likely to be questioned later, and be honest that the requirement costs throughput rather than pretending it is free.

## The diff is the cheapest thing the run produced A long autonomous run ends with a change you can read, a build you can execute and, often, a green result. All of that is visible, and little of it is where the cost sits. The costs that distinguish this shape from a small assisted edit leave no trace in the diff: how many decisions the change forces into one moment, how far a mistake could have travelled while nobody was watching, how much of the session survives the session, and whether the result can be obtained again. Each is a consequence of the shape you chose, not of how well the change happened to turn out. ## Review surface: decisions per glance, not lines changed **Review surface** is the amount of judgement a change demands per human decision point. A completer offers one small unit and takes an answer immediately. A long run offers everything it did, at once, at the end. Two changes of identical size can differ enormously here: forty lines applying one pattern you have already approved is close to a single decision, while forty lines spread across a dozen files, each carrying its own judgement call, is a dozen decisions that arrive together, with one reviewer's attention spread across all of them. The same total size can therefore mean very different things: - **one pattern, many places, already approved** — close to a single decision, whatever the line count; - **many small judgements across many files** — one decision per file, and they all arrive together; - **one subtle change buried in a wide diff** — the expensive case, because the decision that matters is the one least likely to be noticed. This is why *the agent wrote it in ten minutes* is an incomplete sentence. The work did not disappear. It moved to a place where it is harder to do, because it is now concentrated, less familiar, and delivered after every decision has already been taken. What you actually look for while reviewing is a discipline of its own; the point here is only that the shape sets how much of it arrives at once. ## Blast radius: set by reach, not by severity **Blast radius** is what a mistake can touch before anyone notices, and it is determined by what the run could reach rather than by how wrong the mistake was: - **Proposal-only reach**: the worst case is wrong text in front of a person who must act before anything happens. - **Project-wide edit reach**: the worst case is a wrong change resting in a file nobody opened. - **Execution reach**: the worst case includes state that no diff displays at all. Severity is independent of this. A trivially wrong edit with wide reach can be a bigger incident than a badly wrong one that never left a proposal. How you constrain that reach, and how you recover once something has gone wide, are separate subjects with their own homes; what belongs here is that choosing the shape is already choosing the radius. ## Reconstructability: the what without the why Months later, somebody asks why a rule reads the way it does. What answers them is whatever the run recorded **at the time**: what it was asked, what it read, what it ran, what came back. A shape that leaves only a final change leaves the *what* and not the *why*, and a diff is a poor witness because it shows the decision taken and never the alternative rejected. Asking the model afterwards does not recover it either — a later account is a fresh generation, plausible and unevidenced, not a log. ## Reproducibility: a run is not an experiment Re-issuing the same request generally does not produce the same change. The sampling differs, the assembled context differs, and the state of the working copy has moved on. Two consequences follow: 1. A difference between two runs is the expected behaviour rather than a signal, so it tells you very little about whether either one was right. 2. Anything you want to keep from a run has to be captured while it happens, because there is no re-derivation later. This is the property that most surprises people arriving from deterministic tooling, where *run it again and look* is a legitimate move. ## The four costs against the two shapes | | review surface | blast radius | reconstructable | reproducible | |---|---|---|---|---| | **small assisted edit** | one small unit at a time | the edit you were watching | trivially, because you were there | not needed at this size | | **long autonomous run** | the whole change, at the end | whatever the run could reach | only what the run recorded | no | ## Answering this in an interview Price the run before you start it rather than after it. Say what reach it will have, what you expect to be able to reconstruct, and whether you can afford to review what it will produce — and be willing to say that a shape is a bad deal on a given task even though it would reach a working change faster. A candidate who can only describe the upside of autonomy is describing a demo. A candidate who can price it has run one.

  • Two runs change the same number of lines. Why might one be far more expensive to review?
    Because review cost tracks decisions, not lines. Forty lines applying a pattern you have already approved is roughly one decision; forty lines across a dozen files, each with its own judgement call, is a dozen — and they arrive together, so attention is thinnest exactly where the change is least familiar.
  • Why can a second run of the same request not be treated as a repeat of the first?
    Because the sampling, the assembled context and the state of the working copy all differ, so a different change is a normal outcome rather than a signal. Runs are not repeatable trials: anything worth keeping — the commands, the outputs, the intermediate state — has to be recorded while it happens.
  • Does a bigger blast radius mean the change was more likely to be wrong?
    No, and conflating the two leads people astray. Radius is about what could be touched; likelihood is about how well specified and checked the task was. A well-checked wide change and a badly specified narrow one can have exactly the opposite risk profile to the one their sizes suggest.

saying these in an interview costs you the question

  • Treats a green build as evidence the run was cheap
  • Measures review cost in lines changed rather than decisions required
  • Assumes re-running the same request will reproduce the change
  • Thinks a mistake's reach is set by how serious the mistake was
  • Believes the final diff is a complete record of the session