skip to content

How do you decide how much work belongs in a single orchestrated task?

level: seniorimportance: should knowfreq 48%

answer

  1. ask what you want to retry
  2. and what you want to see separately
  3. different failure causes, different tasks
  4. never bury a side effect mid-task
  5. minutes of work, one destination

basics

~20 s

Size a task around what you want to retry, observe, parallelise and rerun. Split where failure causes differ or where a step is not safely repeatable; keep together work that shares state or must commit as one visible effect.

solid answer

~50 s

A task is simultaneously four units: the unit of **retry**, the unit of **observability**, the unit of **parallelism**, and the unit of **rerun blast radius**. Sizing follows from those. Split at boundaries where the failure modes differ — a flaky API pull, a schema validation and a warehouse load fail for unrelated reasons and want different retry policies — and split any step that is not safely repeatable away from the compute around it, so a retry of the transform does not re-send an email. Keep together work that shares expensive local state, or that must land as one all-or-nothing effect, because the orchestrator will not roll anything back for you. Too coarse and you rerun forty minutes to fix the last two; too fine and you pay per-task scheduling overhead, fragment your logs, and load the scheduler with thousands of nodes doing seconds of work each. Roughly: minutes of work per task, one destination, retryable on its own.

code

text · 6 lines
text
# Before: one node, one retry policy, one log, 40-minute reruns
daily_sales  (extract + validate + load + swap + email)

# After: boundaries follow failure causes and side effects
extract_orders(retry=3) -> validate_orders(retry=0) -> load_staging(retry=2)
    -> publish_swap(retry=1) -> notify(retry=0)

go deeper

for a junior

Remember that a retry re-runs the entire task, so anything inside happens again — which is why long tasks that mix a data load with sending a notification are a bad idea.

for a middle

Explain the four roles a task plays — retry, observability, parallelism, rerun scope — and give concrete split points: different failure causes, expensive steps you do not want repeated, and non-repeatable side effects.

for a senior

Demonstrate you have paid for bad boundaries: the mid-task email re-sent by retries, the forty-minute rerun to fix two, the partial write left behind. Talk about staging plus atomic publish, and about not rebuilding an engine's parallelism as graph nodes.

for a principal

Own it as a standard: task boundaries follow failure domains and commit boundaries, side effects are isolated, publication is atomic, and orchestrator fan-out is bounded — so retry policy and alert routing become consequences of the shape rather than per-pipeline improvisation.

## A task is four units at once Granularity questions get answered badly when people reason about "logical steps". Reason instead about the four roles a task actually plays: 1. **Unit of retry.** Retrying re-executes the whole task. Everything inside it happens again. 2. **Unit of observability.** Duration, state, logs and lineage are recorded per task. Anything inside is invisible to the platform. 3. **Unit of parallelism.** Tasks are what the orchestrator can run concurrently and place on workers. 4. **Unit of rerun blast radius.** Reruns operate on a task and its downstream closure. A task is the smallest thing you can redo. Every sizing rule below falls out of one of those. ## Split here **Different failure causes.** An API extract fails on rate limits and timeouts; a validation fails on bad data; a warehouse load fails on locks and permissions. Different causes deserve different retry counts and often different owners — and separate tasks tell you which one broke without reading a log. **A non-repeatable side effect.** Sending an email, calling a payment API, posting to a webhook, incrementing an external counter. Isolate these into their own task so a retry of the neighbouring compute cannot re-fire them, and so you can give them a retry policy of zero if that is what safety requires. The classic incident is a forty-minute task that emails a report at minute thirty-eight and fails at minute thirty-nine; three retries later, four copies have gone out. **Expensive work you do not want to repeat.** If a thirty-minute extract is followed by a two-minute load that fails intermittently, one task means every load failure costs thirty-two minutes. Splitting makes the retry cost two. **Meaningfully different resource shapes.** A memory-hungry transform and a trivial metadata update should not share sizing, a worker class, or a concurrency limit. ## Keep together here **Shared expensive local state.** If step two needs the in-memory or on-local-disk output of step one, splitting them forces you to materialise the intermediate to shared storage. Sometimes that is worth it; often the intermediate is worthless on its own and the write is pure cost. **One atomic effect.** The orchestrator gives you no transaction across tasks: if the second half fails, nothing undoes the first half. Work that must be all-or-nothing to consumers should either sit in one task that commits once, or be structured as write-to-staging tasks plus a final publish task that swaps atomically. **Trivial work.** A task that runs for two hundred milliseconds but takes seconds to schedule, queue and record is nearly all overhead. ## The two failure modes **Too coarse** looks like: one task called `run_pipeline`, all failures reported identically, a rerun redoing everything, no idea where time goes, and partial writes left behind by a mid-task failure that nothing cleans up. **Too fine** looks like: thousands of nodes, scheduler and metadata pressure, per-task startup dominating actual work, a graph nobody can read on screen, and logs fragmented across hundreds of task views so tracing one record means opening dozens of them. ## Do not rebuild an engine's parallelism in the graph A common mistake is generating one task per file for ten thousand files. The orchestrator becomes the bottleneck doing work a data engine does natively: hand it the prefix and let it parallelise internally. Use orchestrator-level fan-out when the units are genuinely independent, coarse (minutes each), individually retryable, and few enough to read — dozens to low hundreds, not tens of thousands. ## Atomicity, concretely A well-shaped task leaves exactly one durable effect that is visible either fully or not at all. Techniques: - write to a staging path or table, then swap, rename or repoint in the final step; - overwrite a whole partition rather than appending rows, so a rerun replaces rather than duplicates; - key writes so that a repeat is a no-op or an upsert. The test to apply to any task you design: *if this is killed halfway and retried, is the outcome the same as if it had run once cleanly?* If the answer is no, either fix the task or move the offending part out of it. ## A worked reshape A single forty-minute `daily_sales` task becomes: `extract_orders` (retry 3, network-flaky), `validate_orders` (retry 0 — bad data will not fix itself), `load_staging` (retry 2), `publish_swap` (retry 1, atomic), `notify` (retry 0, isolated side effect). Now a bad-data day fails in two minutes at validation with an obvious cause, lock contention on load retries cheaply, and no retry can double-send the notification. ## Granularity is also a communication choice Task names are what an on-call engineer reads at 3 a.m. and what appears in lineage and dashboards. A graph whose node names describe meaningful stages — extract, validate, load, publish — diagnoses itself; a graph of `step_1` through `step_12`, or one giant node, does not. If a boundary makes the failure message obviously actionable, that is evidence it is the right boundary. ## Interview framing Lead with the four units, give the split-here and keep-together rules, name both failure modes, and finish with the retry test. Mentioning the isolated side effect and the atomic publish is what marks the answer as coming from operating pipelines rather than reading about them.

  • You must process 10,000 files. One task per file, or one task for the whole prefix?
    One task for the prefix in almost all cases — hand the whole set to an engine built for partitioned parallelism instead of making the scheduler manage ten thousand nodes with per-task overhead, unreadable graphs and fragmented logs. Orchestrator-level fan-out earns its keep when units are coarse, genuinely independent, individually retryable, and countable in dozens or low hundreds.
  • How do you test whether a task is the right size?
    Ask what happens if it is killed halfway and retried: the outcome must match a single clean run. Then ask what a failure costs in redone work, and whether the task name alone tells you which step broke. If a retry duplicates a side effect, or one flaky minute forces redoing thirty, the boundaries are wrong.
  • Should a task that transforms data also send the completion email?
    No. The email is a non-repeatable external effect; the transform is retryable compute. Bundled together, any retry of the transform re-sends the mail, and you cannot give the two different retry policies. Split the notification into its own downstream task with retries set to what is actually safe.
  • Two adjacent tasks keep failing together because the second depends on scratch files the first wrote locally. Merge them?
    Either merge them or make the handoff durable. Tasks may run on different hosts, so local scratch is not a supported channel. If the intermediate has no value on its own and is expensive to materialise, one task is the honest answer; if it is reusable or the steps fail for different reasons, write it to shared storage and keep them apart.

saying these in an interview costs you the question

  • Sizes tasks by logical tidiness rather than retry and rerun cost
  • Leaves an email or external API call in the middle of a long task
  • Generates one task per file for thousands of small files
  • Assumes a failed later task undoes an earlier task's writes
  • Says finer is always better without mentioning scheduling overhead

context