skip to content

How do you tell truly parallel subtasks from false parallelism in an agent plan?

level: seniorimportance: should knowfreq 40%

answer

  1. not obviously ordered is not independent
  2. three sets, checked disjoint
  3. last writer wins, silently
  4. the read set includes decisions
  5. independent but still contending

basics

~20 s

Two subtasks are only independent if neither reads what the other writes and they touch no shared mutable resource. Absence of an obvious ordering is not evidence of independence — check data flow, shared write targets and shared bottlenecks before branching.

solid answer

~50 s

The test is not "do these read like sequential steps?" but "does either one observe the other's effects?". Three things create hidden edges. **Shared write targets**: two subtasks that both edit the same file or record are ordered even though neither consumes the other's output — run them together and the last writer wins silently. **Hidden data flow**: one step's result narrows another's scope, like a survey step that determines which services actually need patching. **Shared bottlenecks**: two branches that both saturate one rate-limited API are logically independent but gain nothing from concurrency, and may lose throughput to retries. So for each proposed parallel pair I ask what each writes, what each reads, and what scarce resource each consumes. If all three are disjoint, the branch is real; otherwise it is false parallelism, and the failure mode is a corrupted artifact rather than an error you can see.

code

yaml · 9 lines
yaml
# False parallelism: disjoint reads, overlapping writes
tasks:
  - id: patch_billing
    reads: [services/billing]
    writes: [services/billing, shared/lockfile]
  - id: patch_search
    reads: [services/search]
    writes: [services/search, shared/lockfile]
# shared/lockfile appears in both write sets -> must not run concurrently

go deeper

for a junior

Know that two steps are only parallel if neither needs the other's result, and that steps writing to the same file are ordered even when neither reads the other.

for a middle

Explain the check concretely: compare what each subtask writes, what it reads including scope-defining decisions, and what scarce resource it consumes, and say what goes wrong when write sets overlap.

for a senior

Demonstrate that the dangerous failure is silent — last-writer-wins on a shared artifact with both branches reporting success — and describe mitigations: declared write sets enforced by the executor, or isolated state with an explicit merge.

for a principal

Own the policy angle: decide when concurrency is permitted at all given the blast radius of a silent overwrite, and whether the platform requires declared write sets before any branch may run.

## Why this question exists Decomposition produces a set of subtasks and an ordering. The tempting move is to treat every step not obviously downstream of another as parallelizable, because that is where the wall-clock win is. False parallelism is what happens when that judgement is wrong: work is launched concurrently that was actually ordered, and — this is the important part — the failure usually does not look like a failure. It looks like a plausible result that is missing half its inputs, or an artifact where one branch's edits overwrote the other's. ## The three sources of hidden dependency **1. Shared write targets.** Two subtasks that both modify the same artifact are ordered, even when neither reads the other's output. An agent upgrading a shared dependency across twelve services will happily plan twelve independent-looking subtasks; if three of them edit the same lockfile or the same shared config, running them together means the last write wins and two services are silently left unpatched. Nothing errors. The plan reports twelve successes. **2. Hidden data flow.** A step's *scope* can depend on another step's *result* even when its inputs look self-contained. A survey step that determines which services are actually affected must precede the per-service work; skipping the edge means patching services that did not need it and missing ones that did. This edge is easy to miss because the dependency is on a decision, not on a file. **3. Shared bottleneck resources.** Two branches can be perfectly independent in state terms and still gain nothing from concurrency because they contend for one scarce thing — a rate-limited API, a single database connection, a serialized review queue, a human approver. This is the benign case: nothing corrupts, but the parallelism is an illusion and often a net loss once retries and backoff enter the picture. ## A usable test For each candidate pair, write down three sets and check they are disjoint: - **Writes**: every artifact, file, record or external system state the subtask mutates. - **Reads**: everything it consumes, *including* decisions made by earlier steps that define its scope. - **Consumes**: the scarce resources it draws on. Disjoint writes and no read-of-the-other's-write means the branch is safe. Disjoint consumes means the branch is worth taking. Both conditions matter and they fail for different reasons — the first produces wrong results, the second produces no speedup. The part candidates most often miss is that a subtask's read set includes *implicit* reads. A step that says "summarize the affected contracts" reads the definition of "affected", which some earlier step produced. If that definition is not yet final, the step is not ready to run regardless of what the plan says. ## Making dependencies visible rather than inferred The reason this is hard for an LLM-authored plan is that the model orders steps by narrative plausibility, not by state analysis. It has no direct view of what a tool actually mutates. The practical mitigation is to make the plan state it: each subtask names what it will write, and the executor refuses to run two subtasks whose write sets intersect. That converts a reasoning problem into a mechanical check — and mechanical checks are the only ones that hold up across many runs. Where the write set genuinely cannot be predicted ahead of time — an agent editing whichever files turn out to be relevant — the conservative move is to keep that stage sequential, or to give each branch its own isolated copy of the mutable state and reconcile afterwards. Isolation converts a shared write into two private ones plus an explicit merge step, which at least surfaces the conflict instead of losing it. ## Signals you got it wrong Watch for: two branches reporting success while the combined artifact contains only one branch's changes; results that vary run to run in ways the plan does not explain; wall-clock that barely improves after parallelizing, which points at a shared bottleneck; and retry storms as branches contend for the same rate-limited dependency. The first is the dangerous one, because it is silent — which is why the write-set check belongs before execution rather than in post-hoc review. ## What a strong answer sounds like Name all three sources of hidden dependency, distinguish the *correctness* failure (shared writes, hidden data flow) from the *throughput* non-win (shared bottleneck), and say plainly that absence of an obvious ordering is not evidence of independence. The best answers add the mitigation: declare write sets so the check is mechanical, and isolate state when it cannot be declared.

  • How do you handle a stage where a subtask's write set cannot be known in advance?
    Either keep that stage sequential, or give each branch an isolated copy of the mutable state and add an explicit merge step afterwards. Isolation converts an invisible last-writer-wins overwrite into a visible conflict at merge time, which you can then resolve or escalate. Guessing the write set and hoping is the option that produces silent corruption, so it is the one to rule out.
  • What early signal suggests branches you declared independent actually were not?
    Two branches both reporting success while the combined artifact only contains one branch's changes, and results that differ between otherwise identical runs. A weaker but useful signal is wall-clock that barely improves after parallelizing, which usually means a shared bottleneck rather than a correctness bug. Compare the produced artifacts against the union of what each branch claimed to write.
  • Is parallelism ever worth it when two branches contend for the same rate-limited dependency?
    Rarely for throughput — the bottleneck caps you either way, and concurrent branches add retry and backoff churn on top. It can still be worth it if the branches spend substantial time on work that does not touch the limited resource, so the contention is only a fraction of each branch. Measure the share of time in the bottleneck before assuming a win.

saying these in an interview costs you the question

  • Assumes any step not obviously downstream can run concurrently
  • Checks only data flow and ignores shared write targets
  • Treats concurrent write conflicts as something retries will fix
  • Expects speedup from branches sharing one rate-limited resource
  • Trusts the model's step ordering as an analysis of state

context