skip to content

Orchestration Concepts

The tool-agnostic ideas every scheduler implements: dependency graphs, data intervals, idempotent reruns, retries and SLAs, and lineage. Interviewers lean on these when they want to know whether you understand pipelines or only one vendor's UI, and the answers transfer directly between Airflow, Prefect, Dagster and managed services.

on this pageshow

explore

questions

page 1 of 2

Why do workflow orchestrators model pipelines as directed acyclic graphs and reject cycles?

level: juniorimportance: must knowfreq 72%

answer

  1. a graph you can never walk in circles
  2. the scheduler needs somewhere to start
  3. and needs to know when it is done
  4. cycles have no valid ordering
  5. loops live in time, not in edges

basics

~20 s

An acyclic graph always has a valid execution order, so the scheduler can find work that is ready and can tell when the run is finished. A cycle leaves tasks waiting on each other forever, with no defensible starting point and no termination.

solid answer

~50 s

A pipeline is a directed graph whose nodes are units of work and whose edges mean *must finish before*. Requiring it to be acyclic guarantees a topological order exists: the scheduler can repeatedly compute the set of tasks whose upstreams are all in an accepted terminal state, launch those, and stop when nothing is left — which is also how it knows the run completed. A cycle destroys both properties: if A depends on B and B depends on A, neither is ever ready, and there is no principled answer to which goes first. Pipelines still need repetition, but they get it **in time**, not in graph structure — task retries, a polling task that waits with a timeout, the next scheduled run, or a dependency on the previous run of the same task. Cycle detection normally runs when the pipeline is loaded, so the error surfaces before anything is scheduled.

code

text · 4 lines
text
# Legal: fan-out, fan-in and a diamond are all acyclic
extract_us   ─┐
extract_eu   ─┼─► merge ─► validate ─┬─► publish
extract_apac ─┘                      └─► notify

go deeper

for a junior

Be ready to define nodes and edges, say that an edge means must-finish-before, and state plainly that a cycle would leave tasks waiting on each other forever so nothing could start or finish.

for a middle

Explain the scheduling loop the acyclic property enables — repeatedly launch whatever has all upstreams satisfied — and show where repetition actually lives: retries, a polling task with a timeout, and the next scheduled run.

for a senior

Show you know cycle detection happens at parse time and stops at the boundary of one graph, so two pipelines waiting on each other's datasets deadlock silently. Talk about how you review cross-pipeline dependencies.

for a principal

Own the argument that graph shape is a platform contract: the acyclic model is what makes impact analysis, rerun scope, and completion targets computable at all, and that is why you refuse designs that need feedback edges between pipelines.

## The model A workflow orchestrator represents a pipeline as a graph. Each **node** is a task — a unit of work the orchestrator can start, watch, and record a terminal state for. Each **directed edge** from A to B means "B may not start until A has finished acceptably". *Acyclic* means there is no path that leaves a node and comes back to it, including a node pointing at itself. That single constraint is what makes the graph schedulable. ## What the acyclic property buys the scheduler A directed graph is acyclic if and only if a **topological order** exists — an ordering of all nodes in which every edge points forward. The scheduler does not need to compute that order explicitly; it runs the equivalent loop: 1. Find every task whose upstream tasks have all reached a state the task accepts. 2. Launch them (possibly many at once). 3. When a task finishes, repeat. 4. When nothing is runnable and nothing is running, the run is over. Acyclicity guarantees step 1 finds something on the first pass (some node has no upstreams) and that step 4 eventually happens. It also gives you the things people build on top: a finish time to compare against a freshness target, a critical path to reason about duration, a well-defined "everything downstream of X" set for reruns and impact analysis, and a progress measure that means something. With a cycle, none of that holds. Members of the cycle can never satisfy the readiness test, so they never start, and because they never start, the run never completes. The question "which of A and B goes first?" has no answer inside the graph — answering it requires state the graph does not model, such as an iteration counter or a convergence condition. ## What a cycle would actually mean A cycle is an attempt to express "do this again under some condition". But the condition lives outside the dependency structure: it is data, or a counter, or a timeout. Dependency edges have no place to put it. That is why orchestrators do not offer a "loop edge" with a bound — they push repetition somewhere it can be reasoned about and limited. ## Where repetition actually lives The usual objection is "but my process loops". Orchestrators express iteration along the time axis instead of the graph axis: - **Retries.** A failed task is re-executed as the same node, a bounded number of times. That is repetition without an edge. - **Waiting tasks.** A task that polls for a file, a partition, or an upstream signal loops internally (or reschedules itself) until the condition holds or a timeout fires. The loop is inside one node. - **The next scheduled run.** Pipelines are re-instantiated per interval; the graph runs again from the top on fresh inputs. This is the main loop of most data platforms. - **Cross-run dependencies.** A task can be made to wait for its own previous run, or a downstream graph can wait for an upstream graph's output. That is an edge between *instances* in time, not a cycle in the graph definition. - **Genuinely data-dependent iteration** — train until the loss converges, retry a reconciliation until deltas are empty — belongs inside a single task, or in an engine designed for it, with its own bound. ## When cycles are caught Cycle detection is cheap — a depth-first search or Kahn's algorithm over the declared edges — and normally runs when the pipeline definition is parsed or registered, not when it runs. The pipeline is rejected with an error naming the offending nodes, so the failure is a development-time error rather than a stuck production run. One important gap: this check sees only the edges of one graph. If graph A waits for a dataset produced by graph B while graph B waits for a dataset produced by graph A, most orchestrators will not detect it. Both simply wait, and you find out from timeouts or a missed freshness target. Cross-pipeline dependencies deserve an explicit review for exactly this reason. ## Shapes that are allowed and often confused Acyclic does not mean tree-shaped. A node may have many parents and many children: - **Fan-out**: one extract feeding three independent transforms. - **Fan-in**: three regional extracts feeding one merge. - **Diamond**: A feeding B and C, both feeding D. Perfectly legal; D simply waits for both. The common mistake is adding "one edge back" so a cleanup or notification step re-runs after a later stage. The fix is a new node downstream with a trigger policy that fires whatever happened upstream — not an edge that closes a loop. ## What to say in an interview Name the two guarantees — a valid order exists, and the run terminates — then show you know where loops went: retries, waiting tasks, and the next interval. Candidates who can only say "cycles are bad" have not thought about what the scheduler is actually doing.

  • A step must poll for an input file until it appears. Doesn't that need a loop in the graph?
    No. The polling lives inside a single task that waits and re-checks, with a timeout, or the task fails and is retried on a delay. Either way the repetition happens in time, as repeated executions of one node, and the declared graph stays acyclic. Edges never carry an "if not ready, go back" meaning.
  • When is a cycle detected — when the pipeline is written, or when it runs?
    Almost always when the definition is parsed or registered. A depth-first search over the declared edges is cheap, so the orchestrator rejects the pipeline with an error naming the nodes in the cycle before it can ever be scheduled. You get a development-time failure rather than a run that hangs.
  • Pipeline A waits on a dataset produced by pipeline B, and B waits on one produced by A. Is that a cycle?
    Logically yes, and it deadlocks — but most orchestrators check only within one graph, so nothing rejects it. Both pipelines sit waiting until their sensors or timeouts fire, and you diagnose it from missed freshness rather than an error. Cross-pipeline dependencies need a deliberate review, or a shared registry of who produces what.

saying these in an interview costs you the question

  • Says cycles just cause slow runs rather than non-termination
  • Claims retries or polling require an edge looping back
  • Thinks the graph must be a tree with one parent per task
  • Cannot explain what acyclicity gives the scheduler
  • Believes the orchestrator detects cycles across separate pipelines

context

open as a page

What is the difference between ETL and ELT in a data pipeline?

level: juniorimportance: must knowfreq 82%

basics

~20 s

ETL transforms data in a separate processing tier before writing it to the target. ELT loads source-shaped data into the target first and transforms it there, using the target's own compute, keeping the raw input queryable.

open as a page

What makes a scheduled batch task idempotent, and why do orchestrators require it?

level: juniorimportance: must knowfreq 75%

basics

~20 s

A task is idempotent when running it again for the same input window leaves the target in the same state as one successful run. Orchestrators retry, replay and re-run tasks constantly, so anything else duplicates data.

open as a page

What does the pipeline cron schedule `*/15 9-17 * * 1-5` actually fire on?

level: juniorimportance: must knowfreq 70%

basics

~20 s

It fires every fifteen minutes, at :00, :15, :30 and :45, during hours 09 through 17 inclusive, Monday to Friday. That is 36 firings per weekday, the last one at 17:45, and none at weekends.

open as a page

In a workflow orchestrator, when one upstream branch of a fan-in task fails, what decides whether it runs?

level: middleimportance: must knowfreq 68%

basics

~20 s

The downstream task's trigger policy — the rule saying which combination of upstream terminal states lets it start. The default everywhere is all upstreams succeeded, so one failed branch blocks the join and the run ends incomplete rather than publishing partial data.

open as a page

Why did cheap warehouse compute make ELT the default over ETL?

level: middleimportance: must knowfreq 62%

basics

~20 s

Elastic columnar warehouses over cheap object storage made loading raw and transforming with SQL cheaper and faster than provisioning a separate ETL cluster. Transformation became rented per query, written in SQL, and rerunnable against retained raw data.

open as a page

Why is a task that appends its window's rows unsafe to re-run, and what write pattern fixes it?

level: middleimportance: must knowfreq 68%

basics

~20 s

An append has no way to remove what a previous attempt wrote, so every re-run adds another copy. Fix it by making the write replace the slice: delete-then-insert for that window in one transaction, a merge on a natural key, or an atomic partition swap.

open as a page

A pipeline task finished successfully but wrote a tenth of the usual rows — which metrics catch that?

level: middleimportance: must knowfreq 62%

basics

~20 s

Exit status only proves the code ran. Catching this needs data-level metrics recorded per run: output row counts compared with recent history, input-to-output ratios, null and distinct rates on key columns, and dataset freshness — with thresholds that fail the run.

open as a page

In a data pipeline, which task failures are worth retrying and which will just fail again?

level: middleimportance: must knowfreq 72%

basics

~20 s

Retry failures whose cause may differ next attempt: timeouts, throttling, preempted workers, transient locks. Deterministic causes — bad input data, a schema mismatch, a missing permission, a code bug — fail identically every time, so retries only delay the alert and burn compute.

open as a page

In a scheduled data pipeline, why does the midnight run process yesterday's data rather than today's?

level: middleimportance: must knowfreq 78%

basics

~20 s

A scheduled run is named for the data interval it covers, not for the clock time it starts. A daily interval only closes at midnight, so the run firing then is responsible for the day that just ended.

open as a page

How do you stop a failed data-quality check from publishing bad rows to downstream pipeline consumers?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Write to a staging location first, run the checks against it, and publish only if they pass — the write-audit-publish pattern. Make the check a task that downstream work depends on, so a failure leaves the previous good version in place and stops dependent tasks from running.

open as a page

Why does alerting only on pipeline failures miss most data-freshness SLA breaches?

level: seniorimportance: must knowfreq 68%

basics

~20 s

Failure alerts fire on errors, but data goes stale in ways that produce no error: a run that succeeded over an empty source, a run still going and simply late, a paused pipeline that never started. Monitor the dataset's age, not the job's exit status.

open as a page

What is data lineage in a data pipeline, and what questions does it let a team answer?

level: juniorimportance: should knowfreq 58%

basics

~20 s

Data lineage is the recorded graph of which datasets each pipeline job read and which it wrote. It answers where a table's numbers came from, what breaks if an upstream column changes, and which reports a failed job affected.

open as a page

In an orchestrated data pipeline, what happens when a task with three configured retries fails?

level: juniorimportance: should knowfreq 60%

basics

~20 s

The orchestrator marks that attempt failed, waits the configured delay, then re-runs the whole task body from the start. Only when the retry budget is exhausted does the task count as failed and downstream work stop.

open as a page

In an orchestrated pipeline, what does a dependency edge between two tasks actually guarantee?

level: middleimportance: should knowfreq 64%

basics

~20 s

Only ordering: the downstream task will not start until the upstream one reaches a terminal state the downstream accepts. The edge moves no data, shares no memory, and spans no transaction — data travels through storage both tasks agree on.

open as a page

In an orchestrated pipeline with a conditional branch, why does a skipped task also skip everything downstream?

level: middleimportance: should knowfreq 50%

basics

~20 s

Skipped is a terminal state distinct from success, and the default rule for starting a task is that all upstreams succeeded. A skipped parent does not satisfy it, so the child skips too, and the skip cascades down the whole branch.

open as a page

In an ELT pipeline, why keep an immutable copy of the raw extracted data?

level: middleimportance: should knowfreq 52%

basics

~20 s

Transformation logic changes and has bugs. An untouched raw copy lets you rebuild every downstream table from what actually arrived, without re-extracting from sources that overwrite rows, purge history, throttle large reads or have since changed shape.

open as a page

How does a scheduler's automatic catchup of missed windows differ from an operator-triggered backfill?

level: middleimportance: should knowfreq 62%

basics

~20 s

Catchup is the scheduler filling in windows it never ran — every interval between the pipeline's start date and now — automatically as part of normal scheduling. A backfill is a human explicitly asking for a stated range to be re-run, usually because the logic or the source changed.

open as a page

Why should a batch task derive its target partition from the run's interval rather than the current date?

level: middleimportance: should knowfreq 58%

basics

~20 s

Because the wall clock moves and the interval does not. A task scoped by the current date reads and writes whatever today happens to be, so retries after midnight, replays and backfills all land in the wrong slice and collide with each other.

open as a page

How does column-level lineage differ from table-level lineage for impact analysis in a data platform?

level: middleimportance: should knowfreq 52%

basics

~20 s

Table-level lineage says table B was built from table A; column-level says B.total came from A.amount and A.qty. Table-level over-reports impact — every downstream consumer looks affected — while column-level narrows a schema change to the handful of fields that actually break.

open as a page

Why does a data pipeline task need an execution timeout as well as a retry policy?

level: middleimportance: should knowfreq 55%

basics

~20 s

Retries only fire on failure. A task that hangs never fails, so it runs forever, holds its worker slot and blocks everything downstream. A timeout converts hanging into a failure, which is what the retry policy and alerting can then act on.

open as a page

Why do pipeline data intervals use half-open bounds that include the start but exclude the end?

level: middleimportance: should knowfreq 50%

basics

~20 s

Half-open bounds make consecutive intervals tile the timeline with no gap and no overlap. A record whose timestamp lands exactly on a boundary belongs to exactly one run, so nothing is processed twice and nothing is skipped.

open as a page

How do you decide how much work belongs in a single orchestrated task?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Size a task around what you want to retry, observe, parallelise and rerun. Split where failure causes differ or where a step is not safely repeatable; keep together work that shares state or must commit as one visible effect.

open as a page

Why should an orchestrator trigger transformation work on external compute rather than run it itself?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Orchestrator workers are sized for coordination, not data. Pulling a dataset into a worker makes its memory the pipeline's ceiling, lets one heavy task starve co-located tasks, wastes the target engine's parallelism, and makes every retry replay the whole transfer.

open as a page

In an ELT stack, when should transformation run on an external engine rather than warehouse SQL?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Move work out only when the language, not the size, is wrong: per-row model scoring, library-dependent parsing, calls to external services, or data that never needs to enter the warehouse. Set-based joins and aggregations should stay pushed down.

open as a page

A backfill re-sent last month's customer emails — how do you make pipeline side effects replay-safe?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Separate computing data from delivering it. Put external effects — email, webhooks, partner file drops — in a step that only fires for current runs, gated on a backfill flag or on the window being recent, so a replay recomputes tables without re-notifying anyone.

open as a page

After a multi-step data pipeline fails midway, when is restarting only the failed steps unsafe?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Restarting from the failure point assumes the completed steps' outputs are still valid and the failed step left nothing behind. It is unsafe when the failed step wrote partially, when earlier steps captured a source snapshot that has since moved, or when a shared staging area was cleared.

open as a page

An hourly pipeline regularly takes 90 minutes to finish — what happens to the next scheduled run?

level: seniorimportance: should knowfreq 55%

basics

~20 s

It depends on the concurrency policy. Either two runs execute at once and contend for the same source and target, or the pipeline is capped at one active run and the queue backs up, so each run starts later than the last and lag grows without bound.

open as a page

Your ELT warehouse bill tripled as models multiplied — how do you decide where transformation belongs?

level: principalimportance: should knowfreq 36%

basics

~20 s

Attribute spend per model first — a handful usually dominate. Most of the overrun is shape and schedule, not location: full rebuilds that should be incremental, refreshes nobody reads, per-row logic. Move work off the warehouse only where measurement says the pricing genuinely misfits.

open as a page

How would you plan a three-year backfill so it neither starves the daily schedule nor explodes cost?

level: principalimportance: should knowfreq 38%

basics

~20 s

Size it first: windows times per-run cost and duration. Then chunk into coarser windows where the transform allows, run it on isolated compute with a hard concurrency cap, checkpoint progress so it resumes, build into a shadow target and swap after validation.

open as a page

showing 1–30 of 36