skip to content

How do you pick subtask granularity when an agent decomposes a goal?

level: middleimportance: must knowfreq 65%

answer

  1. chosen per goal, not fixed
  2. two opposing costs meet here
  3. ask how you would check it
  4. unbounded scope, unmeasurable verb
  5. overhead per step versus work per step

basics

~20 s

Size each subtask so its completion can be checked objectively and it still fits one focused stretch of work. Too coarse and nobody can tell whether it succeeded; too fine and per-step overhead costs more than the work itself.

solid answer

~50 s

Granularity is chosen, not given, and the deciding question is "how would I verify this step finished correctly?". A subtask like *review all contracts* has no answer to that — it is a theme, not a task. Rewritten as *review the 14 supplier contracts signed after 2023 and list the ones with auto-renewal clauses*, the inputs are bounded, the output has a shape, and success is checkable without re-reading everything. The opposite failure is real too: splitting the same work into 14 one-contract steps multiplies turns, tokens and latency, and gives the agent 14 chances to lose the thread for no extra safety. My working heuristic is that a subtask should have named inputs, a named output artifact, and a check that a script or a reviewer could apply in seconds. If I cannot write that check, the subtask is still too coarse.

code

markdown · 7 lines
markdown
# Plan (rewritten for verifiability)

- [ ] BAD:  Review all contracts
- [ ] GOOD: Review the 14 supplier contracts signed after 2023;
            write one row per contract to findings/auto-renewal.csv
            with columns: contract_id, has_auto_renewal, notice_days
- [ ] TOO FINE: Read page 3 of contract 7

go deeper

for a junior

Be ready to say that a subtask needs a clear scope and a clear output, and to rewrite a vague step like 'review all contracts' into a bounded, checkable one.

for a middle

Explain both failure directions with mechanics: unverifiable coarse steps produce false success claims, and micro-steps multiply per-turn overhead and context restatement without adding safety.

for a senior

Show the judgment: derive granularity from the verification you actually have and the tools' natural unit, and name the run-time symptoms that tell you the size was wrong.

for a principal

Own the tradeoff at the fleet level — granularity sets cost per task, auditability and how much of the work can be checked mechanically, so it is a standard your teams should agree on rather than each agent inventing.

## What granularity means here When an agent turns a goal into a plan, it emits a list of subtasks. Granularity is how much work each of those subtasks represents — the difference between *do the due diligence*, *review all contracts*, *review the 14 supplier contracts signed after 2023*, and *read page 3 of contract 7*. All four describe the same underlying work at different resolutions. Choosing the resolution is one of the few decisions in decomposition that reliably decides whether the run succeeds. ## The verification floor The lower bound on coarseness is verifiability. A subtask exists so that the agent — and whoever audits the run — can say "this is done" and move on. If completion cannot be checked, the agent has no signal, and the usual outcome is a confident claim of success over work that was never finished. *Review all contracts* fails this test twice over: the set is unbounded (which contracts? in which repository?) and the verb is unmeasurable (what does "reviewed" produce?). A useful drill is to write the check before the task. If the check is "a list of contract IDs with a boolean auto-renewal flag, one row per contract in the named set", the subtask writes itself. If no check can be written in a sentence, the subtask is still a theme and needs to be split until each piece has one. Verification also determines whether a subtask can be safely handed off. A step whose output is a small, named artifact can be produced by anyone — a later stage, a fresh session, or a separate subagent — and the caller can judge the result without redoing it. A step whose output is "understanding" cannot be handed anywhere. ## The overhead ceiling The upper bound on fineness is cost. Every subtask carries fixed overhead: at minimum a model turn that re-reads the plan and the relevant context, and often a tool round trip, a status write and a token bill for restating what the previous step already established. Splitting a coherent piece of work into a dozen micro-steps multiplies that overhead without buying accuracy — and each additional hop is another place where the agent can misread state or drift off the goal. So the practical band is: fine enough that failure is localized and checkable, coarse enough that a step is worth the turn it costs. A good test is whether the step's *body* is substantially larger than the *bookkeeping* around it. If restating the context takes more tokens than doing the work, merge it upward. ## Heuristics that hold up - **Bound the input set explicitly.** "The 14 supplier contracts signed after 2023" beats "the relevant contracts" — the agent cannot silently redefine the scope mid-run. - **Name the output artifact.** A file, a table, a list, a patch. "Analyze X" produces nothing; "write findings for X to `findings/x.md`" does. - **One verb, one resource class.** A subtask that both fetches documents and judges them is two steps whose failures are indistinguishable from the outside. - **Match the tool's natural unit.** If a tool returns a page at a time, a per-page subtask is fine; if it returns a corpus, a per-corpus subtask is fine. Fighting the tool's unit is what creates either 200 tiny steps or one unbounded one. - **Size to the verification you actually have.** With a programmatic checker (tests, a schema, a diff), steps can be larger, because failure is caught mechanically. With only human review, keep steps small enough that a person will actually read the output. ## Signals you chose wrong Too coarse looks like: a step that runs for many turns without producing an artifact; a step marked done with a summary rather than an output; two runs of the same step producing incomparable results. Too fine looks like: adjacent steps whose outputs are trivially concatenated; a plan longer than the work; most tokens spent restating context rather than doing anything. ## Interview framing What separates a strong answer is treating granularity as a *decision with two opposing costs* rather than a style preference, and naming the verification test as the thing that resolves it. Weak answers assert that smaller is safer, which is the intuition from human project management and does not transfer: for an agent, each extra step is another opportunity to lose state, not another checkpoint of safety. Strong answers also note that granularity is workload-dependent — the right size shifts with how reliable the tools are, how expensive a step is, and whether an automated verifier exists.

  • How does granularity shift when the agent has a programmatic verifier for a step?
    A mechanical check — tests, a schema validation, a diff that must apply — lets you make steps larger, because failure is caught reliably regardless of how much work the step did. Without one, you are relying on the model's own claim of success, so steps should be small enough that a human or a later step can spot-check the output cheaply. The verifier, not the prose, sets the safe size.
  • What granularity do you use when only a human can judge the subtask's output?
    Small enough that the human will actually read it. A reviewer will genuinely check a bounded artifact — fourteen rows, one patch, one page of findings — and will rubber-stamp a fifty-page dump. So when the check is human, size the subtask to the reviewer's attention budget, and keep the output shape identical across steps so review becomes routine rather than a fresh reading task each time.
  • When is it right to merge two subtasks that each look verifiable on their own?
    When the second adds no independent check and cannot fail separately — for example, fetching a document and extracting one field from it, where a failed fetch makes the extraction moot and the extraction has no failure mode of its own. Merging removes a turn, a context restatement and a bookkeeping write while losing nothing observable. Keep them split only if you would ever want to retry or delegate one without the other.

saying these in an interview costs you the question

  • Claims finer subtasks are always safer than coarser ones
  • Leaves subtasks like 'analyze the data' with no completion criterion
  • Treats granularity as fixed rather than chosen per goal and toolset
  • Ignores that every subtask costs a turn, tokens and latency
  • Sizes subtasks by wording style instead of by verifiability

context