When a shared upstream pipeline misses its SLA, should downstream pipelines block or run on stale data?
answer
- old is not the same as wrong
- who reads it, and what do they do with it
- a stalled queue then a stampede
- mixed vintage joins fail silently
- publish the vintage with the data
basics
~20 sIt depends on whether the consumer is harmed more by an old answer or a wrong one. Block where output is published externally or joins would mix vintages; proceed with an explicit staleness marker where consumers can tolerate age and see it.
solid answer
~50 sThere is no universal answer, so make it a per-dataset contract rather than a platform default. **Blocking** keeps everything internally consistent and prevents a mixed-vintage join — a fresh fact table against yesterday's dimension, where new keys silently fall to unknown or drop. Its cost is blast radius: one late upstream stalls every dependant, runs queue behind it, and when it clears you get a thundering backfill competing for the same warehouse. **Proceeding** keeps dashboards alive and confines the damage, but only if the staleness is *visible* — a last-updated stamp, a partition marker, a banner. Silent staleness is the worst outcome, because a plausible wrong number outlives the incident. My rule of thumb: block anything published externally or feeding an automated decision; let internal exploratory surfaces proceed, clearly marked; and always page the upstream owner rather than each dependant team.
code
text · 7 linesorders_raw fresh today 06:00
customer_dim STALE yesterday 02:00 (upstream missed its SLA)
daily_orders_mart = orders_raw JOIN customer_dim
-> today's new customers have no dimension row
-> inner join drops them; left join buckets them as 'unknown'
-> run is green, totals look plausible, revenue is understatedgo deeper
Understand that a pipeline can run successfully on out-of-date input, and that whether that is acceptable depends entirely on who consumes the result.
Explain the concrete danger of proceeding — a fresh fact table joined to a stale dimension drops or mislabels new keys with no error — and the value of a visible last-updated stamp.
Weigh the two failure modes in an incident: blocking stalls dependants and creates a catch-up stampede, proceeding risks silent wrongness. Show how you would decide per consumer and what you would notify.
Own it as contract and platform policy: each consuming dataset declares its maximum acceptable input age and its behaviour on breach, the platform enforces both as configuration, vintage travels with the data, and one page goes to the upstream owner.
## Why this is a policy question, not a technical one Both behaviours are trivially implementable — the orchestrator can gate on a freshness assertion, or not. The hard part is that blocking and proceeding fail in opposite directions, and which failure you prefer depends on the consumer. A principal-level answer classifies consumers first and only then talks about mechanism. ## What blocking actually buys and costs **Buys:** internal consistency. Everything downstream reflects one coherent state of the world, and nobody publishes a number derived from data that was not there. It also concentrates the incident: one owner, one root cause, one fix. **Costs:** blast radius and recovery shape. A late upstream that gates fifty downstream pipelines converts a single team's problem into a platform-wide stall. Blocked runs accumulate — one per schedule period — and when the upstream finally lands, all of them become runnable at once and contend for the same compute, which is how a two-hour delay turns into a full day of degraded warehouse performance. Blocking also trains people to disable the gate under pressure, which is worse than not having it. ## What proceeding buys and costs **Buys:** availability and a contained failure. A dashboard that shows yesterday's data with a visible date is often perfectly useful; forcing it to show nothing helps nobody. Teams whose input is genuinely tolerant keep working. **Costs:** silent wrongness, and it compounds. Two shapes matter: - **Mixed vintage.** A pipeline joins a fresh fact table to a dimension that failed to refresh. New keys have no match: depending on the join type they either drop rows or land in an "unknown" bucket. The totals still look plausible, the run is green, and the error is discovered weeks later by a business user. - **Propagated staleness.** A consumer of the consumer has no idea any of this happened. Staleness that is explicit at the first hop becomes invisible at the third unless the marker travels with the data. ## The classification I would actually use - **Block** when the output is published outside the company, feeds a regulatory or financial report, drives an automated action (pricing, sending, paying, provisioning), or joins fresh data to the stale dataset in a way that silently drops or misattributes rows. - **Proceed, marked** when the consumer is an internal analytical surface with a human in the loop, the output is fully recomputed on the next successful run, and the staleness is displayed where the human will see it. - **Proceed, degraded** when a partial answer has real value — serve the fresh portion, exclude and label the stale portion, rather than blending them. Notice that the axis is *what the consumer does with the output*, not *how stale the data is*. Six hours is nothing for a monthly trend and catastrophic for an intraday operational feed. ## Make staleness a first-class property of the data The technique that makes "proceed" defensible is carrying the vintage with the data instead of only in the orchestrator's state. Stamp each output with the effective input timestamps it was built from; expose that in the serving layer; and have consuming surfaces show it. A dashboard headed "as of 02:00 yesterday" is an honest artefact. The same dashboard with no stamp is a trap. This also fixes the third-hop problem: the marker propagates because it is a column, not a run state. ```text orders_raw fresh 06:00 customer_dim STALE yesterday 02:00 <- upstream missed SLA -> daily_orders_mart joins both new customers from today have no dim row rows drop or bucket to 'unknown'; totals look normal ``` ## Who gets paged One page, to the upstream owner, with the list of affected downstream datasets attached. Paging every dependant team produces a storm in which nobody owns the fix and everybody loses trust in the alerting. Dependants need *notification*, not a page — ideally delivered into the surface they use, plus an entry on a status page for the platform. ## The organisational half This decision cannot be made per-incident at 03:00 by whoever is on call, because that person does not know the tolerance of fifty consumers. It has to be declared in advance, by the owner of each consuming dataset, as part of that dataset's contract: *maximum acceptable input age, and what to do when it is exceeded*. The platform's job is to make both behaviours one line of configuration, to enforce the declared choice consistently, and to make the resulting staleness visible everywhere the data is used. ## Second-order effects worth naming When a critical upstream regularly misses its SLA, the right answer is usually neither gate nor proceed — it is to reduce the dependency: keep a last-known-good copy of slowly changing inputs so downstream can build against a defined previous vintage rather than an undefined one; split monolithic pipelines so a late partition blocks only what it feeds; and price the incident honestly so the upstream's reliability work gets funded. Blocking policy is a mitigation for a reliability problem, not a substitute for solving it.
- What happens operationally when a blocked upstream finally lands?Every gated run becomes eligible at once and they contend for the same warehouse or cluster, so a short delay turns into extended degradation. Plan for it: cap concurrency for catch-up work, prioritise the datasets with the tightest promises, and consider collapsing several queued periods into one wider run rather than executing each separately.
- How do you stop staleness becoming invisible three hops downstream?Carry the vintage as data, not as orchestrator state. Stamp each output with the effective timestamps of the inputs it was built from, propagate that through every derived table, and surface it in the consuming tool. A run-state flag stops at the first hop; a column travels as far as the data does.
- Who should be paged when a shared upstream misses its SLA?The upstream owner, once, with the affected downstream datasets listed. Dependant teams get a notification and a status-page entry, not a page. Paging every consumer creates a storm where nobody owns the fix, and it is the fastest way to teach an organisation to ignore its own alerts.
A newspaper that prints yesterday's edition with yesterday's date is old; one that prints yesterday's stories under today's date is wrong. Staleness is survivable, disguised staleness is not.
saying these in an interview costs you the question
- Applying one blocking policy to every pipeline on the platform
- Letting downstream proceed on stale input with no visible marker
- Assuming a stale dimension joined to fresh facts fails loudly
- Paging every downstream team instead of the upstream owner
- Ignoring the catch-up stampede when the blocked backlog is released