Why does building a whole feature in one agent run fail differently from building it as five?
answer
- the difference is structural, not skill
- when does the first check happen
- later work rests on an early choice
- one moment of truth, at the end
basics
~20 sA single run has no interior — nothing inside it meets a check you already trusted. An early wrong decision surfaces only after later work is built on it; five units catch it while it is still one unit wide.
solid answer
~50 sThe difference is not that a model handles small jobs well and large ones badly — it is where the checks sit. Migrating a ticketing system's reporting module off a date library that is being removed, as one request, gives you **one point where the work meets a check you already trusted, and it is at the end**. If the run settled the week-boundary rule differently from the old library on an early turn, every report it touched afterwards is consistent with that choice, and the suite it wrote asserts it. Cut the same work into five units and the same decision meets a real check while it is still one report wide. What you buy is not accuracy per turn. It is earlier evidence, a smaller thing to throw away, and a failure you can attribute to one instruction.
go deeper
Remember the shape: one run gives you a single moment of truth and it arrives at the end. Be able to say what that costs when an early decision was wrong.
Explain the mechanism rather than asserting it. Later turns produce work consistent with an early choice, and tests written afterwards assert that choice, so the end-of-run signals agree with a decision nobody reviewed.
Show that you treat it as a trade. Name what splitting costs — setup, review and plan upkeep per unit — and name the conditions under which you would still hand the whole job over in one go.
The interesting question is what a team standardises. A rule of thumb about unit size is cheap to state and wrong half the time; a shared habit of naming the check before the unit travels better and does not need policing.
## Two plans for the same work A ticketing system's reporting module reads ticket history and produces three things: a weekly ticket-volume report, a month-to-date resolution-time report, and a downloadable export. It takes its dates from a library whose maintainers are removing it, so the module has to move onto something else. **Plan one** is a single request describing the whole migration, handed to an agent that runs until it reports it is finished. **Plan two** is five units — the module's date work behind one internal helper, then that helper onto the replacement a few operations at a time, then the old dependency deleted — each unit checked and landed before the next begins. The work is identical. What differs is where the checks sit, and that decides most of the rest. ## The explanation that sounds right and is not the useful one The reflex answer is that models are worse at bigger jobs. Hold that loosely. What a run gets right is bounded by what it could see and what it was checked against, and asking for less does not by itself improve either — plenty of wide, mechanical changes come back fine, and plenty of narrow, subtle ones come back wrong. The difference that holds up is structural: **one run has no interior.** It compiles things and runs the suite as it goes, but those checks sit downstream of its own decisions, and the tests among them assert the behaviour it just built. The first check that *can* disagree is the one you apply at the end, and by then every decision the run made is in the same diff. ## What "built on top of it" means here The old library rounded week boundaries one way — say a week starts on Monday, and the last few days of a year are reported as their own short week. A replacement can round differently. Nobody wrote that rule down, so the run settles it early, in passing, while getting something to compile. Everything after that is consistent with the choice: - the weekly-volume report groups by the boundary the run picked; - the month-to-date comparison is computed over those groups; - the export inherits the same grouping; - the tests the run wrote assert the behaviour it built. At the end the module builds and the suite is green. Both facts are true, and neither is evidence about the rule, because the tests came after the choice and encode it. You are not looking at a bug in one file. You are looking at a premise, spread across a diff, with a passing suite on top of it. ## The comparison, laid out | | one run over the whole module | five units, each ending checked | |---|---|---| | first check that can disagree | at the end | at the end of unit one | | what a failed check tells you | something in here is wrong | this unit is wrong | | what an early wrong decision costs | the feature | one unit | | what you discard to recover | the run's whole output | the last unit | | what it costs when all goes well | one setup, one review | five setups, five reviews | That last row is the honest one: this is a trade, not a rule. Splitting is not free, and a plan cut so fine that each unit costs more to set up and read than to perform is its own mistake. ## What you are actually buying 1. **Earlier evidence.** The first unit meets a real check while the decision under it is still one report wide. 2. **A bounded loss.** What you abandon when a unit is wrong is a unit. How expensive an abandonment is also depends on what the run could reach outside your working tree, which is a separate subject. 3. **An attributable failure.** With one unit in flight you can say which instruction produced the result. With a whole feature in flight you can say that the feature is wrong. None of that makes any individual unit more likely to be correct, and saying so plainly separates a solid answer from an overreaching one. Splitting changes **when you find out** and **how much you lose**, not the quality of any single turn. ## When one run is the right call When the change is confined to a place you can name, when a check you trust already exists and runs in seconds, and when a wrong version would fail that check rather than sail past it — hand over the whole thing. The five-unit discipline earns its overhead on work whose interior contains decisions, whose verification has to be built first, or where a late discovery is expensive. Asking which of those is true is the scoping decision; answering it with a habit is not. ## What a good answer sounds like Name the structural difference first — *one run has one moment of truth and it is at the end* — then say what that costs in the concrete case, then say when you would still do it in one go. An answer that claims models are simply bad at large tasks has skipped the mechanism. An answer that never says when one run is fine has turned a trade-off into a rule.
- When is handing over the whole job in one run the right call?When the change sits in a place you can name, a check you already trust runs in seconds, and a wrong version would fail that check rather than pass it. Splitting costs a setup and a review per unit, so on small bounded work the overhead buys nothing. The discipline earns its keep when verification has to be built first, or when finding out late is expensive.
- Does cutting the work into units make each unit's output more accurate?Not by itself. Each unit is a smaller job with a check attached, and a check does not improve the turn that produced the code — it tells you sooner whether that turn was any good. The purchase is earlier evidence and a smaller loss. If a unit is wrong, you still have to notice it, and the check is what lets you.
- The one-run migration ended with a green suite. Why is that weak evidence?Because the run wrote most of that suite after making the decisions the suite now asserts. A test written from the behaviour just built agrees with it by construction. Green tells you the module is self-consistent and still compiles; it says nothing about whether the week-boundary rule matches what the reports meant before the migration.
saying these in an interview costs you the question
- The only problem with a big request is the model forgetting its end
- Splitting the work is just slower for the same result
- A green build at the end proves the one-run migration was right
- Any mistake in a big run is one edit away from fixed
- Always cut work as small as possible; smaller is safer