skip to content

Orchestration Concepts

The tool-agnostic ideas every scheduler implements: dependency graphs, data intervals, idempotent reruns, retries and SLAs, and lineage. Interviewers lean on these when they want to know whether you understand pipelines or only one vendor's UI, and the answers transfer directly between Airflow, Prefect, Dagster and managed services.

on this pageshow

explore

questions

page 2 of 2

How would you build end-to-end lineage across a platform where several different tools transform data?

level: principalimportance: should knowfreq 33%

basics

~20 s

Pick one interchange format and one metadata store, then feed it from three sources: runtime events emitted by each tool, SQL or query-log parsing where no integration exists, and manual edges for opaque hops. The hard part is agreeing dataset naming so the pieces connect.

open as a page

How do you decide whether a downstream pipeline runs on a clock schedule or on upstream data readiness?

level: principalimportance: should knowfreq 40%

basics

~20 s

Use a clock schedule only when the downstream can tolerate acting on whatever data exists at that moment. If correctness depends on the upstream being complete, trigger on an explicit readiness signal instead, because a guessed time offset fails silently when the upstream is late.

open as a page

In the OpenLineage standard, what does a run event carry and what are facets used for?

level: middleimportance: nice to knowfreq 28%

basics

~20 s

An OpenLineage run event ties three entities together — a job, one run of it, and the input and output datasets — plus an event type such as START, COMPLETE or FAIL and a timestamp. Facets are optional typed blocks that attach extra metadata to any of them.

open as a page

When should a pipeline generate its tasks dynamically at runtime instead of declaring a fixed graph?

level: seniorimportance: nice to knowfreq 36%

basics

~20 s

Only when the set of work units is genuinely unknown until the run starts and each unit deserves its own retry and visibility — for example one task per tenant discovered at run time. Otherwise a fixed graph is cheaper to read, compare and operate.

open as a page

What happens to a pipeline scheduled in a local timezone when that zone's daylight-saving change occurs?

level: seniorimportance: nice to knowfreq 35%

basics

~20 s

One local wall-clock hour disappears in spring and one repeats in autumn, so a schedule inside those hours may be skipped or fire twice, and that day's nominal daily interval is 23 or 25 hours instead of 24.

open as a page

When a shared upstream pipeline misses its SLA, should downstream pipelines block or run on stale data?

level: principalimportance: nice to knowfreq 34%

basics

~20 s

It depends on whether the consumer is harmed more by an old answer or a wrong one. Block where output is published externally or joins would mix vintages; proceed with an explicit staleness marker where consumers can tolerate age and see it.

open as a page

showing 31–36 of 36