skip to content

Why does alerting only on pipeline failures miss most data-freshness SLA breaches?

level: seniorimportance: must knowfreq 68%

answer

  1. green job, empty table
  2. nothing failed, nothing arrived
  3. a paused pipeline cannot page you
  4. watch the data's age, not the exit code
  5. alert on absence, not on error

basics

~20 s

Failure alerts fire on errors, but data goes stale in ways that produce no error: a run that succeeded over an empty source, a run still going and simply late, a paused pipeline that never started. Monitor the dataset's age, not the job's exit status.

solid answer

~50 s

Failure alerting is job-centric; an SLA is dataset-centric, and the two disagree constantly. A run can succeed while loading zero rows because the source dropped a file. It can succeed on a partial extract. It can still be running at 09:00 with a 06:00 promise — late, not failed. It can be paused, disabled, or never scheduled at all, in which case there is no run to fail. Every one of those breaches the promise and none of them raises an error. The fix is to **alert on the absence of an expected outcome** rather than on error events: for each published dataset, a deadline monitor asserting the newest record or load timestamp is younger than X at time T, plus volume checks against recent history. Run that monitor **outside** the pipeline it watches, because a pipeline that never started cannot alert on itself.

code

sql · 8 lines
sql
-- runs on its own schedule, independent of the loading pipeline
select
  max(loaded_at) as newest_row,
  timestampdiff(hour, max(loaded_at), current_timestamp) as age_hours,
  count(*) as rows_today
from analytics.orders
where loaded_at >= current_date;
-- breach if age_hours > 3, or rows_today far outside the recent band

go deeper

for a junior

Be able to give one concrete example of a pipeline that succeeds while the data is stale — an empty source producing a zero-row load — and say why a failure alert stays quiet.

for a middle

Explain the difference between a completion SLA and a freshness SLA, and describe how a freshness check is actually measured against the target table.

for a senior

Walk an incident: green runs, stale mart, and the layered detection you would add — freshness, volume, deadline warnings — plus why the watchdog must be independent of the pipeline.

for a principal

Own the policy across the platform: which datasets carry consumer-facing promises, what pages versus files a ticket, and how staleness is surfaced to consumers so silence is never mistaken for correctness.

## Two different promises A pipeline failure alert says *my job errored*. An SLA says *this dataset will be no more than N hours old*, or *this table will be complete by 06:00*. Those are different statements about different objects, and mistaking one for the other is the most common structural gap in data platform monitoring. Job success is a necessary condition for freshness; it is nowhere near sufficient. ## The ways data goes stale without an error **Success with no data.** The upstream drop did not happen, the source path matched zero files, the API returned an empty page. The task ran, wrote nothing, and exited zero. Everything is green and the mart still shows yesterday. **Success with partial data.** A source system was mid-restart and served a fraction of the rows, or a paginated read silently stopped early. The load is technically correct — it loaded what it was given. **Late but not failed.** The run started on time and is still going at hour six of a promised two. Failure alerting will not fire until it fails, which it may never do. A retry chain is a special case of this: attempts are being spent, no terminal state has been reached, the dashboard is quiet. **Never ran.** Someone paused the pipeline during an incident and forgot to resume it. A deploy dropped the schedule. The scheduler component itself was down. The trigger condition never fired. There is no run object, therefore no failure, therefore no alert. This class is invisible by construction to anything that watches runs. **Succeeded into the wrong place.** The run wrote to a staging table and the swap step was skipped by a dependency rule, or wrote the correct data under yesterday's partition. Green run, stale target. ## Completion SLA versus freshness SLA Define both, because they catch different things. A **completion SLA** is a clock statement about the pipeline: *the daily load must finish by 06:00*. It is easy to reason about, maps to downstream schedules, and detects lateness. Its weakness is that it is still framed around a run, so "no run at all" needs to be an explicit case. A **freshness SLA** is a statement about the data: *the orders table is never more than three hours old*. It is measured by querying the target — `max(event_time)`, `max(loaded_at)`, or the newest partition — and comparing to now. It survives pipeline redesign, it does not care how many jobs produce the table, and it catches every silent case above, including the pipeline that no longer exists. Most mature setups carry both: freshness as the consumer-facing promise, completion as the operational deadline that gives you time to react before the promise breaks. ```sql -- freshness assertion, run independently of the pipeline it watches select max(loaded_at) as newest, timestampdiff(hour, max(loaded_at), now()) as age_hours from analytics.orders having age_hours > 3; -- non-empty result = SLA breached ``` ## Where the monitor must live Outside the pipeline. A check that runs as the last task of the run it validates cannot fire when the run does not start, when the scheduler is down, or when someone paused the schedule — exactly the failure modes that failure alerting already misses. The watchdog needs an independent clock and an independent trigger: a separate lightweight job, a monitoring system's absent-signal rule, or a data-observability tool. The principle is generic: **alert on the absence of an expected event, not on the arrival of an error event.** ## Layering the signals A workable set for a published dataset: - **Freshness** — newest timestamp older than the promise. This is the page-worthy one. - **Volume** — today's row count far outside the recent band. Catches the empty and partial cases while the run is still green. - **Completion deadline** — a warning at a time chosen to leave room to fix things, well before the consumer-facing promise. - **Run health** — the failure and retry-rate signals you already have. Useful for diagnosis, weak as a promise. ## Deciding what pages Not every staleness deserves a phone call at 03:00, and this is where teams either build trust or destroy it. The useful axis is *consumer consequence*, not *severity of the error*. A dataset feeding an externally published report or an operational decision justifies a page at the deadline; a warehouse table read by an internal dashboard reviewed each morning justifies a ticket and a visible staleness banner in the tool. Two things make the difference practical: set the alert threshold at the point where a human still has time to act, and make staleness visible in the consuming surface itself — a last-updated stamp on the dashboard turns a silent wrong number into an obvious old number, which is a far cheaper failure mode. ## The interview-shaped summary Green pipelines are a claim about jobs. SLAs are a claim about data. Monitor the data, put the monitor outside the thing it watches, alert on absence rather than error, and page according to what the consumer loses.

  • Where should a freshness check live — inside the pipeline or outside it?
    Outside, on an independent schedule. A check that runs as the last step of the pipeline is silent in exactly the cases that matter most: the run never started, the schedule was paused, the scheduler itself is down. The watchdog needs its own clock and its own trigger so it can fire when nothing happened at all.
  • How do you avoid paging for a run that is merely slightly late?
    Separate the operational deadline from the consumer promise. Warn at a time that still leaves room to fix things, page only at the promise boundary, and derive both from observed high-percentile finish times rather than the ideal. Publishing a visible last-updated stamp to consumers also converts many pages into self-service answers.
  • A run succeeds but loads zero rows. Which signal catches that fastest?
    A volume check against the recent band, since it fails the moment the load completes rather than waiting for the freshness clock to run out. Freshness will catch it too, but only once the dataset's age crosses the threshold — potentially hours later. Volume and freshness together bound both the silent-empty and the never-arrived cases.

A failure alert is a smoke detector; a freshness monitor is noticing the post has not arrived. Nothing is burning, and the letter you were promised still is not there.

saying these in an interview costs you the question

  • Treating a green pipeline run as proof the data is current
  • Alerting only on task failure and calling that SLA monitoring
  • Running the freshness check as the last task of the pipeline it validates
  • Ignoring the paused or never-scheduled pipeline as a failure mode
  • Confusing a completion deadline with a promise about data age

context