skip to content

What tells you a unit of work is too big for one agent run before you start it?

level: middleimportance: should knowfreq 56%

answer

  1. describe the finish line first
  2. what state is the repository left in
  3. the check has to pre-date the unit
  4. done and wandered must look different

basics

~10 s

Three tells, all visible before you type: the unit has no end state where the project still builds, no check that existed before the unit did, and no area you can name.

solid answer

~40 s

I try to describe the unit's finish line, and the tells show up while I do it. First, **does it end somewhere runnable?** "Take the reporting module off the old date library" ends with the module not building until its call sites move, so it is the middle of a unit rather than a unit. Second, **what would I read to know it landed, that exists today?** If the answer is the tests the run will write, the unit carries its own verification and nothing independent can disagree with it. Third, **can I name where it lands?** Without an area, done and wandered look the same. A unit that fails any of the three gets re-cut — usually into one group of date operations at a time, behind a helper that already exists.

code

text · 15 lines
text
CANDIDATE UNIT A
  "Take the reporting module off the old date library"
    repository when done .... module does not build; call sites still expect the old API
    check that exists now ... none (nothing runs until the call sites move)
    area ................... the whole module
    verdict ................ not a unit - this is the middle of one

CANDIDATE UNIT B
  "Route every date operation in the module through one internal helper.
   The helper still calls the old library. No behaviour changes.
   The existing reporting suite must pass untouched."
    repository when done .... builds; suite green; old library called from one place
    check that exists now ... the existing reporting suite, unmodified
    area ................... the module's date call sites
    verdict ................ a unit - and the next four become checkable behind it

go deeper

for a junior

Get into the habit of describing the finish line before handing work over. If you cannot say what the repository looks like when the unit is done, the unit is not ready to start.

for a middle

Explain why a check has to pre-date the unit. Tests written from the behaviour just built agree with it, so they cannot tell you the unit did what was wanted — only that it is consistent with itself.

for a senior

Show the re-cut, not just the diagnosis. Taking a migration and putting a behaviour-preserving helper first, so the existing suite becomes the check for every later unit, is the move an interviewer is listening for.

for a principal

Worth owning: which checks a team is willing to treat as unit boundaries at all. A suite nobody trusts makes a poor boundary, so unit size and test-suite credibility are the same investment seen from two directions.

## The decision you make before you type anything By the time you hand work to an agent you have already chosen where it stops. That choice is made quickly, usually by restating a line from a ticket, and a surprising share of what goes wrong in agent sessions was decided there rather than during the run. The useful skill is noticing in advance that the unit you are about to hand over cannot be finished. Take the running case: a ticketing system's reporting module has to move off a date library that is being removed. The obvious first unit is *take the reporting module off the old date library*. Three tells say it is too big, and all three are visible before anything runs: 1. **No runnable end state** — the repository cannot be left where this unit leaves it. 2. **No check that pre-dates it** — the only thing that would say it landed is written by the run itself. 3. **No nameable area** — you cannot tell finished from wandered. ## Tell one: it has no runnable end state In a codebase, a unit ends in a state the project can be left in. The obvious cut here — replace the library, then fix everything that called it — has a middle where the old dependency is gone and the call sites still expect it. That middle is not a state you can hand back, check, or leave overnight. It is the inside of a unit, and calling it a unit does not give it an end. The test is one sentence: **describe the repository at the moment the unit is done.** If the honest description contains "and then", the boundary is in the wrong place. ## Tell two: its check does not exist yet Ask what you would read or run to decide the unit landed, and require the answer to name something that exists today. Things that qualify, because they pre-date the unit and can therefore disagree with it: - a suite that passes today and is expected to pass unchanged; - an output you already hold, such as last week's report; - a number someone outside the work can confirm. If instead the answer is *the tests the run will write*, the unit contains its own verification. Those tests are written from the behaviour just built, so they agree with it, and a green result tells you the unit is self-consistent rather than right. That is not an argument against generated tests — it is an argument against counting them as the unit's check, which is a scoping decision rather than a testing one. ## Tell three: you cannot name where it lands If you cannot say roughly which area the change belongs in, you have no way to tell finished from wandered, and no way to size the review in advance. ## The three tells against the running migration | candidate unit | runnable end? | check that already exists? | area you can name? | verdict | |---|---|---|---|---| | take the module off the old date library | no — call sites break in the middle | no | the whole module | re-cut | | put all date work behind one internal helper, behaviour unchanged | yes — suite must pass untouched | yes — the existing suite | the module's date call sites | a unit | | move the helper's week-boundary operations onto the replacement | yes | yes — last week's report output | the helper | a unit | | make the reports right | no finish line at all | no | unbounded | not a unit | Two rows of that table carry the lesson. The second row's whole value is that **nothing observable changes**, which is exactly what makes the existing suite a real check: any difference it reports is a regression the unit caused. The last row is what a ticket line looks like when it is copied across unexamined. ## The opposite mistake is real too None of this argues for the smallest possible unit. Every boundary costs a setup, a review and a place for the plan to go stale, and a unit that changes nothing anyone can observe cannot be checked any more than an oversized one can. Cutting a three-line change into three units buys nothing and spends attention you will want later. The tells above are about the top end, and they are silent about the bottom. ## Two things this is not about - **It is not a file count.** Reach and size correlate loosely at best; what matters is whether one check covers what changed. - **It is not how long the run takes.** A long run over a bounded unit with a real check at the end is fine. A short run whose result nobody can judge is not. ## What a good answer sounds like Give the tells as things you can see in advance, and show one being applied: *I said the unit out loud, noticed that its finish line was "and then fix the callers", and re-cut it so the first unit puts the date work behind a helper and the suite stays green.* A weak answer sizes units by feel or by file count and then discovers the problem during the run, which is the expensive place to discover it.

  • Why does a unit whose only check is the tests the run writes count as unchecked?
    Because those tests are derived from the behaviour just built, so they agree with it by construction. They are useful afterwards, as a guard against the next change, but at the unit boundary you want something that pre-dates the work: a suite that already passed, an output from last week, a number someone can confirm.
  • Is a unit that touches forty files automatically too big?
    No. Reach and size are only loosely related. A mechanical rename across forty files with one check covering all of them is a clean unit, because a wrong version fails the check. A two-file change whose correct result nobody can state is the harder one, however small the diff looks in review.

saying these in an interview costs you the question

  • Size a unit by how many files it will touch
  • Any unit is fine if the agent writes tests for it
  • The smaller the unit, the safer the run — always
  • There is no way to tell a unit is too big until the run fails
  • A unit is finished when the agent says it is finished