How do you decide whether a downstream pipeline runs on a clock schedule or on upstream data readiness?
answer
- an offset encodes a guess as a contract
- late upstream means a green run with wrong data
- wait for evidence, not for a clock
- waiting tasks still consume capacity
- ask how a historical window gets reprocessed
basics
~20 sUse a clock schedule only when the downstream can tolerate acting on whatever data exists at that moment. If correctness depends on the upstream being complete, trigger on an explicit readiness signal instead, because a guessed time offset fails silently when the upstream is late.
solid answer
~60 sThe lazy default — "upstream usually finishes by 02:40, so schedule downstream at 03:00" — encodes a guess as a contract. When the upstream is late the downstream reads incomplete data and succeeds, which is the worst failure shape: wrong numbers with a green run. When the upstream is early the offset is pure wasted latency. The alternatives are to **poll** for readiness (wait for a completion marker, an expected partition, or a row count, with an explicit timeout and a defined action on timeout) or to **push** (the upstream emits a completion signal that starts the downstream). Polling is easy across ownership boundaries and works when the upstream cannot be changed, but a naive poller occupies a worker for the whole wait. Push gives the lowest latency and the tightest coupling, and needs a signal contract plus a way to trigger a specific historical window on demand. My decision factors are correctness sensitivity, whether the two pipelines share an owner, the freshness target, and whether backfills need to drive the downstream independently.
code
text · 13 linesA. clock offset
upstream 02:00 -> usually done 02:40
downstream 03:00 (fixed)
upstream late -> downstream reads partial data, reports SUCCESS
B. poll for readiness
downstream 02:05 -> wait for marker _SUCCESS in partition dt=<interval>
timeout 90 min -> FAIL loudly (do not run on partial input)
C. push on completion
upstream final step announces interval complete
downstream starts immediately for that interval
still needs: a manual trigger path for historical intervalsgo deeper
Know that a downstream job can either run at a fixed time or wait for the upstream to signal that its data is ready, and that a fixed offset is a guess.
Explain why the fixed offset fails silently rather than loudly, and describe how a readiness check works, including the need for a timeout and a defined action when it expires.
Weigh polling against push in production terms: worker occupancy and timeout policy on one side, signal contracts and historical triggering on the other, plus input-freshness alerting either way.
Own the cross-team convention — what a completion signal means, who publishes it, how reprocessing is driven, and when two coupled pipelines should simply become one.
## Why the time-offset default is a trap Coupling two pipelines by a clock offset means the dependency exists only in someone's head. Nothing in the system states that the downstream needs the upstream's output; there is only a number chosen because it looked safe on the day. Three things then go wrong. **Silent incorrectness.** If the upstream is late, the downstream still runs. It reads a partition that is missing or half-written, produces plausible-looking output, and reports success. Green runs with wrong numbers are far more expensive than a red run, because nobody investigates a green run and the bad figures propagate to every consumer downstream. **Wasted latency.** The offset must be sized for the worst case, so on every normal day the downstream sits idle for the padding. Chains of three or four pipelines each padded this way turn a 40-minute critical path into a four-hour one. **Silent drift.** Upstream runtime grows over months. The margin erodes with no signal until the day it goes negative, and the incident looks like a data-quality problem rather than a scheduling one. ## Option one: poll for readiness The downstream starts on its own schedule but its first step waits for evidence that the upstream's window is complete: a completion marker written as the upstream's last action, the existence of the expected output partition, a metadata row recording the interval as done, or a minimum row count. The design questions are: **what is the evidence** (a marker written last is far better than partition existence, which can be true while a write is still in progress); **how long to wait** (a timeout is mandatory — an unbounded wait is an outage that never pages); and **what a timeout means** (fail loudly, run on partial data with a flag, or skip the interval — decide deliberately, and prefer failing). The operational cost is worker occupancy. A poller that sleeps inside its worker slot holds capacity for the entire wait, and a dozen of them can deadlock a pool: every slot is occupied by something waiting for work that cannot start because there are no free slots. Mitigations are pollers that release the worker between checks, sensible poll intervals, and a dedicated capacity pool for waiting tasks. Polling's real virtue is that it needs nothing from the upstream except a signal it can observe, which makes it the pragmatic choice across team and platform boundaries. ## Option two: push on completion The upstream declares "window X is complete" and that declaration starts the downstream. Latency drops to near zero, the padding disappears, and the dependency becomes explicit and inspectable. What you take on: a **signal contract** (which windows are announced, exactly when, and what completeness means — is a partial load announced?), an **at-least-once story** (a duplicate signal must not corrupt output, which brings you back to idempotent partition-scoped writes), and **triggering for historical windows**. That last point is the one teams forget: a scheduled pipeline can always be asked to run an old interval, whereas an event-triggered one only runs when something fires. Reprocessing six months of history requires a way to drive the downstream for a past window without fabricating fake signals. ## Option three: stop pretending they are two pipelines If the two stages share an owner, a cadence and a failure domain, the honest design is often one pipeline with an ordinary dependency edge — no offset, no poller, no signal contract, and one place to look when it breaks. Splitting is justified when the stages have different cadences, different owners, different reprocessing needs, or genuinely independent failure handling. "They are separate pipelines" is frequently an accident of history rather than a decision. ## The decision factors I actually weigh - **Correctness sensitivity.** Can the downstream tolerate partial input? A dashboard refresh often can; a published financial mart cannot. Low tolerance rules out the clock offset immediately. - **Ownership boundary.** Same team: push or merge. Different teams or platforms: poll for an agreed signal, because it requires the least from them. - **Freshness target.** If the SLA has hours of slack, the offset's cost is invisible and its simplicity may win. If it is tight, padding is unaffordable. - **Backfill semantics.** How does a historical window get reprocessed? If the answer is "trigger the downstream for that interval", event-only triggering needs an explicit manual path. - **Fan-in.** A downstream depending on five upstreams should wait on all five, not on the latest guessed offset of the slowest. That is where offsets fail worst and where readiness signals pay for themselves several times over. - **Observability.** Whatever the mechanism, alert on the downstream's *input freshness*, not just its run outcome. That is the check that catches all three failure modes. ## How to answer Start by naming the failure mode of the default — silent success on incomplete data — because that is what makes this a design decision rather than a configuration one. Present polling and push as a coupling-versus-latency tradeoff, state the operational costs of each honestly (worker occupancy; signal contract and historical triggering), and finish with the factors that decide it: correctness tolerance, ownership boundary, freshness budget and how reprocessing is meant to work.
- Why is checking that the upstream's output partition exists a weak readiness signal?Existence can become true mid-write. A partition directory or table appears as soon as the first file lands, so a downstream can start against a half-written window. The stronger signal is a marker or metadata row the upstream writes as its final action, which is only true once everything else has committed.
- How do you stop a fleet of waiting tasks from starving the worker pool?Use pollers that release their worker between checks instead of sleeping in it, give waiting tasks their own capacity pool so they cannot exhaust the one execution needs, set poll intervals proportional to the expected wait, and always bound the wait with a timeout so a stuck upstream cannot hold a slot indefinitely.
- What does event-driven triggering make harder that a schedule gives you for free?Reprocessing history. A scheduled pipeline can be asked to run any past interval on demand; an event-triggered one only runs when a signal arrives. You need an explicit path to trigger the downstream for a specific historical window, and it must not require replaying or fabricating upstream signals.
- When would you keep the clock offset despite its weaknesses?When the downstream tolerates partial input, the freshness budget has hours of slack, and the upstream belongs to a team or platform that cannot emit any observable signal. Under those conditions the offset's simplicity is worth more than its precision — but it should still be paired with an input-freshness check so the silent case is caught.
saying these in an interview costs you the question
- Picks an offset because the upstream 'usually finishes by then'
- Waits for upstream readiness with no timeout at all
- Treats partition existence as proof the write completed
- Ignores that waiting tasks consume worker capacity
- Has no way to trigger the downstream for a historical window