skip to content

Task Scoping

Cutting a feature into units an agent can finish in one coherent pass, and sequencing them so the context window never becomes the limiting factor. Interviewers ask because the failure of large one-shot prompts is the clearest lesson people bring back from real agent use.

part ofAI-assisted developmentoverview, primer and where to startread it →
on this pageshow

questions

5

Why does building a whole feature in one agent run fail differently from building it as five?

level: juniorimportance: must knowfreq 66%

answer

  1. the difference is structural, not skill
  2. when does the first check happen
  3. later work rests on an early choice
  4. one moment of truth, at the end

basics

~20 s

A single run has no interior — nothing inside it meets a check you already trusted. An early wrong decision surfaces only after later work is built on it; five units catch it while it is still one unit wide.

solid answer

~50 s

The difference is not that a model handles small jobs well and large ones badly — it is where the checks sit. Migrating a ticketing system's reporting module off a date library that is being removed, as one request, gives you **one point where the work meets a check you already trusted, and it is at the end**. If the run settled the week-boundary rule differently from the old library on an early turn, every report it touched afterwards is consistent with that choice, and the suite it wrote asserts it. Cut the same work into five units and the same decision meets a real check while it is still one report wide. What you buy is not accuracy per turn. It is earlier evidence, a smaller thing to throw away, and a failure you can attribute to one instruction.

go deeper

for a junior

Remember the shape: one run gives you a single moment of truth and it arrives at the end. Be able to say what that costs when an early decision was wrong.

for a middle

Explain the mechanism rather than asserting it. Later turns produce work consistent with an early choice, and tests written afterwards assert that choice, so the end-of-run signals agree with a decision nobody reviewed.

for a senior

Show that you treat it as a trade. Name what splitting costs — setup, review and plan upkeep per unit — and name the conditions under which you would still hand the whole job over in one go.

for a principal

The interesting question is what a team standardises. A rule of thumb about unit size is cheap to state and wrong half the time; a shared habit of naming the check before the unit travels better and does not need policing.

## Two plans for the same work A ticketing system's reporting module reads ticket history and produces three things: a weekly ticket-volume report, a month-to-date resolution-time report, and a downloadable export. It takes its dates from a library whose maintainers are removing it, so the module has to move onto something else. **Plan one** is a single request describing the whole migration, handed to an agent that runs until it reports it is finished. **Plan two** is five units — the module's date work behind one internal helper, then that helper onto the replacement a few operations at a time, then the old dependency deleted — each unit checked and landed before the next begins. The work is identical. What differs is where the checks sit, and that decides most of the rest. ## The explanation that sounds right and is not the useful one The reflex answer is that models are worse at bigger jobs. Hold that loosely. What a run gets right is bounded by what it could see and what it was checked against, and asking for less does not by itself improve either — plenty of wide, mechanical changes come back fine, and plenty of narrow, subtle ones come back wrong. The difference that holds up is structural: **one run has no interior.** It compiles things and runs the suite as it goes, but those checks sit downstream of its own decisions, and the tests among them assert the behaviour it just built. The first check that *can* disagree is the one you apply at the end, and by then every decision the run made is in the same diff. ## What "built on top of it" means here The old library rounded week boundaries one way — say a week starts on Monday, and the last few days of a year are reported as their own short week. A replacement can round differently. Nobody wrote that rule down, so the run settles it early, in passing, while getting something to compile. Everything after that is consistent with the choice: - the weekly-volume report groups by the boundary the run picked; - the month-to-date comparison is computed over those groups; - the export inherits the same grouping; - the tests the run wrote assert the behaviour it built. At the end the module builds and the suite is green. Both facts are true, and neither is evidence about the rule, because the tests came after the choice and encode it. You are not looking at a bug in one file. You are looking at a premise, spread across a diff, with a passing suite on top of it. ## The comparison, laid out | | one run over the whole module | five units, each ending checked | |---|---|---| | first check that can disagree | at the end | at the end of unit one | | what a failed check tells you | something in here is wrong | this unit is wrong | | what an early wrong decision costs | the feature | one unit | | what you discard to recover | the run's whole output | the last unit | | what it costs when all goes well | one setup, one review | five setups, five reviews | That last row is the honest one: this is a trade, not a rule. Splitting is not free, and a plan cut so fine that each unit costs more to set up and read than to perform is its own mistake. ## What you are actually buying 1. **Earlier evidence.** The first unit meets a real check while the decision under it is still one report wide. 2. **A bounded loss.** What you abandon when a unit is wrong is a unit. How expensive an abandonment is also depends on what the run could reach outside your working tree, which is a separate subject. 3. **An attributable failure.** With one unit in flight you can say which instruction produced the result. With a whole feature in flight you can say that the feature is wrong. None of that makes any individual unit more likely to be correct, and saying so plainly separates a solid answer from an overreaching one. Splitting changes **when you find out** and **how much you lose**, not the quality of any single turn. ## When one run is the right call When the change is confined to a place you can name, when a check you trust already exists and runs in seconds, and when a wrong version would fail that check rather than sail past it — hand over the whole thing. The five-unit discipline earns its overhead on work whose interior contains decisions, whose verification has to be built first, or where a late discovery is expensive. Asking which of those is true is the scoping decision; answering it with a habit is not. ## What a good answer sounds like Name the structural difference first — *one run has one moment of truth and it is at the end* — then say what that costs in the concrete case, then say when you would still do it in one go. An answer that claims models are simply bad at large tasks has skipped the mechanism. An answer that never says when one run is fine has turned a trade-off into a rule.

  • When is handing over the whole job in one run the right call?
    When the change sits in a place you can name, a check you already trust runs in seconds, and a wrong version would fail that check rather than pass it. Splitting costs a setup and a review per unit, so on small bounded work the overhead buys nothing. The discipline earns its keep when verification has to be built first, or when finding out late is expensive.
  • Does cutting the work into units make each unit's output more accurate?
    Not by itself. Each unit is a smaller job with a check attached, and a check does not improve the turn that produced the code — it tells you sooner whether that turn was any good. The purchase is earlier evidence and a smaller loss. If a unit is wrong, you still have to notice it, and the check is what lets you.
  • The one-run migration ended with a green suite. Why is that weak evidence?
    Because the run wrote most of that suite after making the decisions the suite now asserts. A test written from the behaviour just built agrees with it by construction. Green tells you the module is self-consistent and still compiles; it says nothing about whether the week-boundary rule matches what the reports meant before the migration.

saying these in an interview costs you the question

  • The only problem with a big request is the model forgetting its end
  • Splitting the work is just slower for the same result
  • A green build at the end proves the one-run migration was right
  • Any mistake in a big run is one edit away from fixed
  • Always cut work as small as possible; smaller is safer
open as a page

Your five agent units can only be tested once the fourth lands — how do you re-cut the plan?

level: middleimportance: should knowfreq 48%

basics

~10 s

Re-order so evidence arrives first: open with a behaviour-preserving unit the existing suite already checks, then the smallest unit that settles the riskiest assumption. Merge any unit that cannot fail on its own.

open as a page

What tells you a unit of work is too big for one agent run before you start it?

level: middleimportance: should knowfreq 56%

basics

~10 s

Three tells, all visible before you type: the unit has no end state where the project still builds, no check that existed before the unit did, and no area you can name.

open as a page

In a long agent coding session, what fills the context window, and why does the run degrade rather than stop?

level: middleimportance: should knowfreq 52%

basics

~20 s

A long session is filled mostly by its own by-products: files read entire, full command output, failed attempts, its own restatements. The room is finite, so older material goes and the work carries on with less.

open as a page

When a feature is split across several agent runs, what stops the fourth reinventing what the first built?

level: seniorimportance: nice to knowfreq 38%

basics

~20 s

Only what is legible in the code carries across a boundary. A later run reads the repository, not the earlier run's reasoning — so a decision survives as a helper with no route around it, or as a failing test.

open as a page