skip to content

In an orchestrated pipeline, what does a dependency edge between two tasks actually guarantee?

level: middleimportance: should knowfreq 64%

answer

  1. it constrains when, not what
  2. two tasks, two processes, maybe two machines
  3. the arrow carries no payload
  4. storage is the real interface
  5. and nothing rolls back if it fails

basics

~20 s

Only ordering: the downstream task will not start until the upstream one reaches a terminal state the downstream accepts. The edge moves no data, shares no memory, and spans no transaction — data travels through storage both tasks agree on.

solid answer

~50 s

A dependency edge is a scheduling constraint and nothing more. It says the downstream task is not eligible to start until the upstream reached a terminal state that satisfies the downstream's trigger policy — by default, success. What it does **not** provide is a data channel. The two tasks may run in different processes, on different machines, minutes apart, so any value the upstream held in memory is gone. Data moves through a location both sides agree on: an object-storage prefix, a table, a partition, a queue — usually derived from the run's parameters or its data interval so the contract is deterministic. Some orchestrators do offer a small key-value channel between tasks, but it is backed by the metadata store and is sized for identifiers, paths and counts, not datasets. The edge also gives you no transaction: if the downstream fails, nothing rolls back the upstream's writes.

code

python · 9 lines
python
# BAD: the edge does not carry this value
rows = []

def extract():
    global rows
    rows = fetch_from_api()      # only exists in this worker process

def load():
    write_warehouse(rows)        # empty when this task runs elsewhere

go deeper

for a junior

Remember the one-line rule: the arrow controls order, not data. If a later step needs data from an earlier one, the earlier one writes it somewhere durable and the later one reads that location.

for a middle

Explain why in-memory handoff fails — separate processes and hosts — and describe how a deterministic path or table derived from the run's parameters serves as the interface. Know that the small metadata channel is for identifiers only.

for a senior

Bring up the cases that bite in production: a task that reports success before its external job committed, overlapping runs breaking the ordering assumption, and building all-or-nothing publication yourself with staging plus an atomic swap.

for a principal

Frame inter-task handoff as an interface contract the platform should standardise — naming conventions, manifests, publish-by-swap — so that ownership, reruns and impact analysis do not depend on each team inventing its own convention.

## The precise guarantee When you declare `A → B`, you have told the scheduler one thing: **do not start B until A has reached a terminal state that B's trigger policy accepts.** That is the whole contract. Everything else people assume about edges is folklore. Three consequences follow immediately, and each one is a real production bug when it is forgotten. ## 1. The edge is not a data channel Tasks are separate executions. Depending on the deployment they run in different processes, different containers, or on different hosts, and they may be minutes or hours apart. Nothing in an ordinary orchestrator copies A's heap into B's. A pattern like this fails as soon as the pipeline runs anywhere but one developer's laptop: ```python rows = [] # module-level state def extract(): global rows rows = fetch_from_api() # lives only in this worker process def load(): write_warehouse(rows) # empty list when this runs elsewhere ``` The cruel part is that it *works* on a single-process local runner, so it passes review and fails in production. The correct pattern makes the handoff a **location**, not a variable: ```python def extract(interval_start): write_parquet(f"s3://raw/events/dt={interval_start}/", fetch_from_api(interval_start)) def load(interval_start): copy_into_warehouse(f"s3://raw/events/dt={interval_start}/") ``` The path is derived deterministically from the run's own parameters, so both sides compute the same answer without talking to each other, and a rerun of the same logical window addresses the same data. ## 2. The small-metadata channel is not a data channel either Most orchestrators offer a small key-value channel so a task can hand a value to a downstream task — a file path, a row count, a job identifier, a chosen branch. It is backed by the orchestrator's metadata database and is meant for **bytes, not megabytes**. Pushing a dataframe or a blob of records through it bloats the metadata store, slows the scheduler, and eventually hits a size limit. Pass the pointer; leave the payload in storage. ## 3. The edge is not a transaction If B fails, nothing undoes what A wrote. Orchestrators have no distributed rollback. If the run must look all-or-nothing to consumers, you build that yourself: write to a staging location, and make the last task an atomic publish — a partition swap, a view repoint, a rename. Then a failure anywhere before the swap leaves consumers on the previous good version. ## Ordering is per run, not across runs `A → B` orders A and B *within one instance of the graph*. It says nothing about the previous run's B versus this run's A. If the graph can have two runs in flight — because a run overran, or a backfill is executing — then A from a later window may run while B from an earlier one is still going. If your tasks cannot tolerate that, you need an explicit concurrency limit or a cross-run dependency; the edge alone will not save you. ## The subtle failure: success is not the same as visible The edge guarantees A reported a terminal state. It does not guarantee A's effects are **observable** to B. Two common ways this bites: - **Asynchronous submission.** If A fires off a job on an external engine and returns as soon as the job is accepted, A succeeds while the data is still being written. B then reads a half-written dataset. The fix is that a task should not report success until the work it represents has committed — poll to completion, or make the wait its own task. - **Eventually consistent listings.** Some object stores can serve a stale listing right after a write. Prefer reading an explicit manifest or a known path over listing a prefix and hoping. ## What good practice looks like - Treat the storage location as the interface between tasks, and write it down: this task produces `<table>/<partition>`, that task consumes it. - Derive locations from run parameters, never from wall-clock `now()` read inside the task — that alone makes reruns non-reproducible. - Keep the metadata channel for small identifiers. - Put the atomic publish in one task, at the end. - Assume nothing about memory, filesystems, or temporary directories being shared between tasks; if two steps genuinely need the same scratch disk, that is an argument for making them one task. ## Why the misconception is so common Code-level pipelines in a single process really do pass values along, and many tutorials show a function returning a value that the next step consumes. Orchestrators borrow that syntax while changing the execution model underneath: the call you write is a *declaration* of a node, not an invocation whose return value flows onward. Reading declarations as calls is the single most common beginner error in this area. ## The interview answer Say "ordering, not data transfer", then name the three things people wrongly assume: shared memory, a data pipe, and a transaction. Follow with how data really moves — an agreed location derived from the run's parameters — and you have covered what the question is testing.

  • Two adjacent tasks need to share a 2 GB intermediate file. How should they exchange it?
    Through shared durable storage — an object-store prefix or table derived from the run's parameters — with the upstream writing and the downstream reading the agreed location. Never through the orchestrator's metadata channel, and never through local disk unless you can guarantee both run on the same host, which is exactly the assumption that breaks first. If the intermediate is worthless on its own, consider making both steps one task.
  • An upstream task submits a job to an external engine and returns immediately. What breaks downstream?
    The edge fires as soon as submission succeeds, so the downstream task starts reading output that is still being written and sees partial or missing data. A task must not report success until the work it represents has committed: poll the external job to a terminal state inside the task, or add an explicit wait task between submit and consume.
  • Does an edge stop a later run's upstream task from overlapping an earlier run's downstream task?
    No. Edges order tasks inside a single run instance. If runs can overlap — a slow run, a backfill, a manual trigger — later windows may execute alongside earlier ones. If that is unsafe, cap the graph's concurrent runs or add an explicit dependency on the previous run; the edge itself gives you nothing across instances.

saying these in an interview costs you the question

  • Believes the edge passes the upstream's return value automatically
  • Shares data through module-level variables or local disk
  • Pushes whole datasets through the small metadata channel
  • Assumes a failed downstream rolls back upstream writes
  • Treats task success as proof the data is fully written

context