skip to content

In a workflow orchestrator, when one upstream branch of a fan-in task fails, what decides whether it runs?

level: middleimportance: must knowfreq 68%

answer

  1. the arrow says after, not after-what
  2. each task has a rule over parent outcomes
  3. succeeded, failed, skipped are different endings
  4. the default is the strictest one
  5. cleanup and publishing want opposite rules

basics

~20 s

The downstream task's trigger policy — the rule saying which combination of upstream terminal states lets it start. The default everywhere is all upstreams succeeded, so one failed branch blocks the join and the run ends incomplete rather than publishing partial data.

solid answer

~50 s

Every task has a policy describing which upstream outcomes make it eligible to run. The default in every mainstream orchestrator is *all upstreams succeeded*, so a fan-in task with one failed parent never starts; it is marked blocked or upstream-failed rather than failed itself, and the run finishes in a non-success state. Other policies exist and are chosen deliberately: run once all upstreams have **finished** regardless of outcome (teardown, releasing a lock, a summary notification), run only if **at least one failed** (an alert path), run if at least one succeeded (best-effort union), or run when none failed but skips are acceptable (joins after a conditional branch). The dangerous choice is giving a data-producing join the all-finished policy: it will happily publish output built from two of three regions and report green. Cleanup and alerting get the permissive policies; anything that writes a consumer-visible dataset keeps the strict one.

code

text · 6 lines
text
extract_us    (failed)    ─┐
extract_eu    (succeeded) ─┼─► merge ─► publish
extract_apac  (succeeded) ─┘

# policy = all succeeded  -> merge is blocked, run ends incomplete
# policy = all finished   -> merge runs on 2 of 3 regions, silently short

go deeper

for a junior

Know that a task waits for its parents and that, by default, all of them must have succeeded — so one failed branch means the join simply does not run and the pipeline ends in a non-success state.

for a middle

Explain the policy menu in terms of upstream states — all succeeded, all finished, none failed, at least one failed — and say which kind of task each suits, especially teardown versus publishing.

for a senior

Show judgment about partial output: when best-effort is legitimate, how to encode a completeness floor as an assertion, how to keep alerting pointed at the root failure, and what the rerun of a fixed branch has to overwrite.

for a principal

Own the policy as a platform default: publishing joins are strict, partial results must be explicitly marked and still surface as a failed run, and permissive policies are reviewed rather than adopted to make a red pipeline turn green.

## Fan-out, fan-in, and the join question Fan-out is one task with several children; fan-in is one task with several parents. Fan-in is where mental models break, because an edge only says "after", and after says nothing about *after what outcome*. That gap is filled by the downstream task's **trigger policy** (orchestrators name it differently, but the concept is universal): a rule over the terminal states of its upstreams that decides whether this task becomes eligible to run. ## The terminal states involved A task run ends in one of a few states — typically **succeeded**, **failed**, **skipped** (deliberately not applicable this run), and a derived **upstream-failed / blocked** state meaning "this task never ran because its dependencies did not satisfy its policy". That last state matters: a join with one failed parent is not itself failed. Nothing in it broke. Treating it as a failure produces noisy alerting that points at the wrong task. ## The policy menu Broadly, orchestrators offer some subset of: - **All succeeded** — the default. Every upstream must have finished successfully. - **All finished** — every upstream reached *any* terminal state; run regardless of outcome. - **None failed** — succeeded or skipped are both acceptable, but a failure blocks. This is the join-after-a-branch policy. - **At least one failed** — the alerting or compensation path. - **At least one succeeded** — best-effort merge over whatever completed. The names differ by tool; the semantics above are what you are actually being asked about. ## Choosing by what the task does The rule of thumb is: **match the policy to the blast radius of running with incomplete inputs.** - A task that writes a dataset consumers read — a merge, a load, a mart build — keeps the strict *all succeeded* policy. Publishing two of three regions as though it were the whole world is worse than publishing nothing, because downstream consumers cannot tell the difference and the number ends up in a report. - A task that releases a resource — deleting a scratch prefix, tearing down a cluster, dropping a lock, closing a ticket — needs *all finished*, or a failure upstream leaks the resource on exactly the days you can least afford it. - A notification summarising the run also wants *all finished*, otherwise the notification you most need is the one that never fires. - An escalation or compensation step wants *at least one failed*. - A join sitting after a conditional branch wants *none failed*, because one side of the branch is skipped by design on every run. ## The classic defect Someone sets a merge task to all-finished because "the run kept getting stuck". Now a failed regional extract no longer blocks anything: the merge runs over the partitions that exist, publishes a smaller table, and — depending on how run state is derived — the pipeline may even show green because the last task succeeded. Revenue is under-reported for a day and nobody is paged. If partial output is genuinely acceptable, make it **explicit** rather than silent: - record which inputs were present in a run-level metadata row or a column on the output; - publish to a location marked incomplete, or set a flag consumers check; - fire a distinct warning notification, not the normal success one; - and still fail the run's overall state so freshness and completeness monitoring sees it. Silence is the failure mode, not partiality. ## Best-effort fan-in done properly Some pipelines really are best-effort — scraping fifty sources where three are always down. Then *at least one succeeded* is right, but pair it with a floor: the merge itself asserts that, say, at least forty-five of fifty inputs landed and fails if not. That converts "how many is too few?" from an invisible property of the graph into an assertion in code that you can tune and test. ## What happens on the rerun When the failed branch is fixed, you rerun that task and everything downstream of it. This is where a clean graph pays off: because edges define the downstream closure, "clear this task and its descendants" is a well-defined operation. It is also why the join should not have quietly produced output — if it did, the rerun has to overwrite rather than fill a gap, and any consumer that already read the partial version has been served bad data. ## Fan-in also shapes timing A join with many parents starts only when the slowest parent lands, so its completion is governed by the worst branch, not the average. A retry deep in one branch pushes the whole join out. When one slow source keeps delaying every consumer, intermediate joins per group — so fast sources publish without waiting for the laggard — are usually a better answer than relaxing the policy. ## Interview framing State the default, name the alternative policies by their semantics, and — most importantly — say *which kinds of task get which*. Candidates who list policies without connecting them to cleanup versus publishing have memorised a table; candidates who say "the merge stays strict, the teardown goes all-finished, and partial output must be explicit" have run pipelines.

  • Your fan-in task publishes a table. Someone sets it to run once all upstreams finish. What is the risk?
    It publishes silently partial data. If one regional extract failed, the merge builds the table from the remaining regions and consumers cannot tell — the numbers are simply low, and if run state is taken from the final task the pipeline may even look green. Data-producing joins keep the strict policy; if partial is acceptable, mark it explicitly and still fail the run.
  • Which state does the downstream task get when its policy is not satisfied — failed?
    No. It is typically recorded as upstream-failed, blocked, or skipped: it never ran, and nothing inside it broke. That distinction matters for alerting, because paging on every such task means one bad extract produces a dozen pages pointing at innocent tasks. Alert on the root failure and on the run's overall state instead.
  • How do you make a teardown task reliably release resources when a branch fails?
    Give it a policy that fires once all upstreams reach any terminal state, so it runs after failures too, and make the teardown itself idempotent and independent of the upstream's output — it should not need a value that a failed task never produced. Keep it separate from any task that publishes data, so its permissive policy cannot leak into the publish path.
  • Is a fan-in with many parents a scheduling problem in itself?
    It can be. A join with hundreds of parents starts only when the slowest one lands, so its timing is governed by the worst-case branch, and any retry deep in one branch pushes the whole join out. Watch the critical path rather than the average, and consider intermediate joins per group so one slow source does not hold every consumer.

saying these in an interview costs you the question

  • Thinks a failed parent automatically fails the downstream task
  • Sets a data-publishing join to run regardless of upstream outcome
  • Cannot name a case where all-finished is the correct policy
  • Assumes edges alone determine behaviour on failure
  • Publishes partial output without recording that it is partial

context