skip to content

In an orchestrated pipeline with a conditional branch, why does a skipped task also skip everything downstream?

level: middleimportance: should knowfreq 50%

answer

  1. not every ending is success or failure
  2. the default rule wants all parents green
  3. a skipped parent fails that test too
  4. trouble starts where the branches rejoin
  5. let the join accept skips, not failures

basics

~20 s

Skipped is a terminal state distinct from success, and the default rule for starting a task is that all upstreams succeeded. A skipped parent does not satisfy it, so the child skips too, and the skip cascades down the whole branch.

solid answer

~50 s

Conditional branching works by marking the untaken path as **skipped** — a terminal state meaning "not applicable this run", separate from success and from failure. Because the default eligibility rule is *all upstreams succeeded*, a task whose parent skipped is not eligible either, so it skips as well, and the effect propagates to the end of that branch. That is what you want inside a branch, and it is a trap at the point where the branches rejoin: a join task under the default rule sees one skipped parent on every single run and therefore never executes. The fix is to give the join a policy that accepts skipped upstreams as long as none failed. Two other things matter in production: a skip is not an error and should not page anyone, and the branch decision should be a function of the run's own inputs so a rerun of the same window takes the same path.

code

text · 6 lines
text
┌─► full_rebuild    (skipped)  ─┐
decide_load_mode ─┤                              ├─► publish
                  └─► incremental_load (success) ─┘

# publish, policy = all upstreams succeeded -> never runs, on any run
# publish, policy = none failed             -> runs on whichever path was taken

go deeper

for a junior

Recall that skipped is its own ending, separate from success and failure, and that a task whose parent was skipped normally skips as well — which is why an untaken branch goes grey all the way down.

for a middle

Explain the cascade in terms of the default all-succeeded eligibility rule, and know the join-after-a-branch trap plus its fix: a policy that tolerates skipped parents but still blocks on a failure.

for a senior

Show that a skip is invisible to failure alerting, so a mis-evaluated condition silently leaves an output stale — and that the real safety net is a freshness or completeness check on the data, plus deterministic, recorded branch decisions.

for a principal

Weigh branching against keeping the graph shape constant: uniform graphs make runs comparable and alerting simple, so branch only where the paths differ in real cost, and require every conditional output to be covered by an output-level check.

## Three terminal states, not two People think in terms of pass and fail. Orchestrators need a third outcome: **skipped**, meaning the task was deliberately not run because it does not apply to this run. It is not a failure — nothing broke — and it is not a success, because the task produced nothing. Conditional branching is built on that state. A decision task evaluates a condition and the paths not chosen are marked skipped rather than left pending, which is what allows the run to reach a terminal state at all: an untaken branch that stayed pending forever would mean the run never finishes. ## Why the skip cascades Every task's eligibility is decided by a policy over its upstreams' terminal states, and the default policy is *all upstreams succeeded*. A skipped parent fails that test just as surely as a failed one does — for a different reason, but with the same effect. So the child skips, its children skip, and the entire subtree beneath the untaken branch is marked skipped without anyone writing a rule for it. This is the correct behaviour. If you take the full-refresh path, you do not want the three incremental-load tasks running "just in case", and you certainly do not want them sitting pending and blocking completion. ## The join-after-a-branch trap Now put a task after both branches — a publish step, a notification, a metrics refresh — under the default policy. On a full-refresh run, the incremental branch is skipped; on an incremental run, the full branch is skipped. Either way one of the join's parents is skipped, so the join is never eligible and never runs. The pipeline reports a state that looks unremarkable, and the publish silently never happens. The fix is to give the join a policy that treats **skipped as acceptable while failure is not**: run if none of the upstreams failed. Then the join fires on either path but is still blocked by a genuine error. Choosing the fully permissive all-finished policy instead would make the join run even when the taken branch failed, which reintroduces the silent-partial-output problem. ## Skip is not failure, and must not be treated as one Alerting rules that fire on "task did not succeed" will page on every skipped task, which on a branching pipeline is most of the graph. Skips must be filtered out of failure alerting. But there is a second-order risk, and this is the senior point: **a skip can hide a missing output.** If the branch that publishes today's mart is skipped because a condition was mis-evaluated, no task failed, no alert fired, and the mart is simply stale. Task-level state cannot catch this. The defence is a freshness or completeness check on the *output* — "the mart must have a partition for this window" — which fails loudly whether the cause was a failure, a skip, or a pipeline nobody triggered. ## Determinism of the decision A branch decision should be a pure function of the run's own inputs: the data interval, the run parameters, a value read from a table as of that interval. When the condition is computed from `now()`, from a mutable config file, or from a flag a human toggles, the same run re-executed later takes a different path — and reproducing an incident becomes guesswork. Two practices help: - Record the decision as run metadata or as a column on the output, so you can see later which path a given window took. - Prefer decisions that are recoverable: if the branch chose "skip today's load because the source was empty", write a marker recording that emptiness rather than leaving an ambiguous gap. Otherwise a missing partition means either "no data" or "never ran" and you cannot tell which. ## Alternatives to branching Branching is not always the best tool. Two alternatives keep the graph shape constant, which makes runs comparable: - **Make the task itself a no-op** when it does not apply — it runs, finds nothing to do, and succeeds. The graph is uniform and the join needs no special policy. Best when the check is cheap. - **Push the condition into the data**, filtering rows rather than choosing tasks. Often the incremental-versus-full distinction is a predicate, not a topology. Use a real branch when the two paths do materially different work with different costs — spinning up a large cluster for a full rebuild versus a small incremental job — and when running the wrong one has real cost. ## Reruns and skipped subtrees When you rerun a window, remember that clearing a skipped task does not necessarily make it run: it becomes eligible only if the branch decision, re-evaluated, now selects it. If you actually want the other path, you change the inputs or the parameter that drives the decision, not the state of the individual task. This is another reason the decision must be an explicit function of run parameters rather than hidden inside opaque logic. ## Interview framing Say: skipped is a third terminal state; the default eligibility rule is all-succeeded; therefore skip cascades; therefore the join needs a policy accepting skips but not failures. Then add the operational half — skips must not page, but the outputs they were supposed to produce still need a freshness check — and you have answered at senior depth.

  • A publish task sits after two mutually exclusive branches and never runs. What is the fix?
    Change its policy from all-upstreams-succeeded to one that accepts skipped parents while still blocking on failure. On either path exactly one branch is skipped by design, so the strict rule can never be satisfied. Avoid the fully permissive all-finished rule here, because that would let the publish run even when the branch that was actually taken failed.
  • How do you notice that a branch skipped a task that should have produced today's data?
    Not from task states — nothing failed. You need a check on the output itself: a freshness or completeness assertion that the target has a partition or row count for this window, run independently of the pipeline that fills it. That single check catches skips, silent partial publishes, and runs that were never triggered at all.
  • What makes a branch decision safe to rerun?
    It must be a deterministic function of the run's own inputs — its data interval, its parameters, or data read as of that interval — never wall-clock time or a mutable flag. Otherwise re-executing the same window takes a different path than it did originally, which makes incidents unreproducible. Record the chosen path as run metadata so you can see later what happened.
  • When would you avoid branching and just let the task run and do nothing?
    When the check is cheap and the work is naturally empty — the task queries its input, finds no rows, and succeeds. That keeps the graph shape identical on every run, so run histories are comparable and no join needs a special policy. Reserve real branching for paths whose costs differ materially, such as a full rebuild on a large cluster versus a small incremental job.

saying these in an interview costs you the question

  • Treats a skipped task as a failure and alerts on it
  • Cannot explain why a join after a branch never runs
  • Fixes the join by making it run regardless of upstream outcome
  • Bases the branch condition on wall-clock time or a mutable flag
  • Assumes green task states prove the expected output exists

context