When a feature is split across several agent runs, what stops the fourth reinventing what the first built?
answer
- boundaries are places to disagree
- the next run reads the code
- a decision needs somewhere to live
- close the route around the helper
basics
~20 sOnly what is legible in the code carries across a boundary. A later run reads the repository, not the earlier run's reasoning — so a decision survives as a helper with no route around it, or as a failing test.
solid answer
~40 sSplitting the work does not make the units agree — it gives you more boundaries across which they can disagree. Each run starts from what it can see, and what it can see is the repository. So a decision from unit one survives only if it left something behind: the date helper exists **and there is no working route around it**, and the year-end rounding rule is a named test rather than a sentence someone repeats. The failure looks tidy — a second helper appears beside the first, or the same rule is implemented twice with different edge behaviour — and it passes review one unit at a time, because each diff is individually reasonable. None of this guarantees agreement. It makes disagreement fail something instead of merging quietly.
go deeper
Remember that a later run reads the repository, not the earlier run's conversation. If a decision is not visible in the code, treat it as not carried across the boundary.
Explain why the failure survives review: each unit's diff is reasonable on its own, each piece passes its own tests, and nothing is textually duplicated, so only a look across the whole set finds it.
Show the mechanism rather than the intention. A helper with no working route around it and a test named for the awkward case both fail when contradicted; a repasted paragraph fails nothing.
Worth deciding deliberately: which decisions a team requires to be expressed as code or tests rather than as convention. That choice sets how much coherence survives when work is cut into many short runs by many people.
## The problem splitting creates Cutting a feature into units buys earlier evidence and a smaller loss when a unit goes wrong. It also creates something that did not exist before: **boundaries across which the work can disagree with itself.** Each run starts from what it can see. What it can see is the repository — not the reasoning of the run that went before it, which lives in a transcript nobody re-reads, and not the plan in your head. So the question for every decision made in an early unit is narrow and practical: *what did that decision leave behind that a later run will trip over?* ## What the failure looks like Across the reporting migration, unit one puts every date call site behind an internal helper. Three units later, a run adding a period column to the export writes its own small date routine beside it, because that was the shortest path to a passing test. Both are reasonable code. Together they are two date policies in one module, and the year-end rounding differs between them. It is a tidy-looking failure, which is why it survives: - each unit's diff is individually sensible, so unit-by-unit review passes it; - the tests pass, because each piece was tested against itself; - nothing is duplicated textually, so a mechanical check has nothing to flag; - it is only visible across the set, and the set is never looked at as one thing. ## What actually carries a decision across a boundary | what you want to survive | what really carries it | what does not | |---|---|---| | all date work goes through one helper | the helper exists and the old route is gone | a line in the plan | | week boundaries round this way at year end | a test named for the year-end case | a sentence in the previous session | | the export's times are the customer's, not the server's | the helper's parameters make it explicit | a convention people know | | why the rule was chosen | a note recorded with the change | the code, which shows only the what | The right-hand column is the point. Intentions expressed only as text outside the repository do not cross the boundary, because the next run does not read them and a colleague picking the work up on Thursday does not either. ## Three things that actually hold a set of units together 1. **Make the shared thing the first unit's product, and close the alternative.** A helper that later units *may* use is a suggestion. A helper they must use, because the direct route no longer compiles or the old dependency is gone from the module, is a constraint. The second costs a little more to build and is the one that survives contact with four more runs. 2. **Encode decisions as tests, not as prose you re-paste.** A test named for the last week of a year fails when a later unit contradicts it, in the ordinary run of the build. A paragraph carried from session to session degrades every time it is restated, and it never fails anything. 3. **Review the set once, not only the units.** Splitting moves a review cost; it does not remove one. One pass over the finished feature asking *is this one design or four?* catches what unit-by-unit review is structurally unable to see. ## What this does not claim It does not claim these measures guarantee agreement — a later run can still add a second path, and people do it too, for the same reason: it was the shortest way to green. What the measures change is the cost and visibility of the disagreement. A contradicted test fails loudly at the moment it is contradicted; a contradicted intention merges. Nor does it claim splitting is the cause of the disagreement. Two developers, two weeks apart, produce the same duplication for the same reasons. Agent units simply run the sequence faster and more often, so a habit that leaked slowly now leaks at speed. ## The cost side Closing the alternative route is real work: sometimes it means deleting a path that still has callers, or taking the dependency out of a module before you would otherwise have bothered. It is worth it when several units queue behind the seam, and it is over-engineering when the feature is two units long and one person is doing both this afternoon. ## What a good answer sounds like Say plainly that nothing about splitting makes the units agree, then name the carrier: **the code, not the conversation.** Give one concrete example — the helper with no way around it, the test named for the year-end case — and finish with the review pass over the set, because that is the part people leave out. A weak answer promises to "keep context between sessions", which describes the problem rather than a mechanism that fails when it is violated.
- Why is repasting the previous session's decisions into the next one a weak fix?Because it is restated by hand each time, degrades as it is summarised, and fails nothing when it is ignored. It is also invisible to the colleague who picks the work up. A test that fails when the rule is contradicted, or a route that no longer exists, does the same job without needing anyone to remember.
- How do you catch a second implementation that duplicates the first without copying it?Read the feature once as a whole, after the units are done, asking whether it is one design or several. Unit-by-unit review cannot see it, because each diff is individually reasonable and nothing is textually duplicated. It is the one review pass that splitting adds rather than removes.
saying these in an interview costs you the question
- Splitting the work makes the units consistent by itself
- Repasting the last session's decisions keeps later units consistent
- If each unit's diff is reasonable, the feature is coherent
- Duplicate logic always shows up as duplicated text
- A convention the team knows does not need encoding