A nightly ETL job and a streaming ingest feed have no request-and-response pattern, so availability and latency SLIs do not fit. What SLIs would you define for them, and why is "the job exited successfully" not one?
answer
- measure the data, not the job
- the event becomes a record or a partition
- freshness bound comes from the consumer
- exit zero, zero rows written
- worst partition, not the average
basics
~20 sDefine indicators on the data, not the job: freshness (how recently the output was updated), coverage or completeness (share of input records processed), and correctness (share of outputs that validate). A zero-exit job can emit nothing at all, so exit status says nothing about what a consumer received.
solid answer
~50 sFor data systems the unit of a good event is a record, a partition or a window rather than a request, and the indicator types shift accordingly. Freshness is usually the primary one: the proportion of time the consumer-visible output was no older than a stated bound, with the bound taken from what the downstream consumer can tolerate. Alongside it, coverage — the share of input records that made it through — catches silent partial loss, and correctness catches output that arrived on time and is wrong, validated against reconciliation totals or a sampled checker. Exit status is a poor indicator because it measures the orchestrator's view, not the consumer's: a job can exit zero having read an empty source, written zero rows, or skipped a partition, and every downstream user sees stale data while the pipeline dashboard is green. I would also measure freshness per partition and report the worst one, since averaging hides the single tenant whose data stopped updating.
go deeper
Know that pipelines are measured on their output rather than on whether the job ran, and be able to name freshness as the main indicator: how old is the data a consumer sees.
Explain how the good-event unit shifts from a request to a record or partition, and give the three indicator types — freshness, coverage, correctness — with a concrete way to compute each.
Show the failure modes that motivate them: a zero-exit job writing zero rows, coverage computed from the pipeline's own state proving nothing, and an averaged freshness number hiding one stalled tenant. Say where the freshness bound comes from.
Own the cost side. A tighter freshness bound commits the organisation to more frequent runs, more infrastructure and more pages, so it should be negotiated against what consumers actually need. Be ready to argue for tiering partitions rather than applying one strict bound to everything.
## The unit of a good event changes Everything about service levels is built on counting good events out of valid events. For a request-serving system the event is obvious. For a data system you have to choose it: a record, a file, a partition, a micro-batch, or a unit of time during which the output was usable. Once the event is chosen, the familiar indicator types have direct analogues — but they are computed **on the output the consumer reads**, never on the machinery that produced it. ## Freshness Freshness is the primary indicator for almost every pipeline, and it answers: *how old is the data a consumer sees right now?* Two common formulations: - **Time-based**: the proportion of minutes in the window during which the output's age was under the bound. Suits a continuously consumed dataset. - **Read-based**: the proportion of consumer reads that returned data fresher than the bound. Stronger, because it weights by actual use — nobody cares that the dataset was stale at 04:00 if nothing read it then. The bound comes from the consumer, and this is where the decision has a cost. A finance dashboard refreshed daily tolerates a bound of a few hours; a fraud model scoring live transactions may need minutes; a recommendation index may tolerate a day. Setting a tighter bound than the consumer needs buys you nothing and commits you to running the pipeline more often, on more infrastructure, with more failures to respond to. Crucially, freshness is derived from a timestamp **carried in the data** — a watermark, a max event time, or a partition's completion marker — not from when the job last ran. A job that runs on schedule and writes nothing keeps its run timestamp current while the data ages. ## Coverage and completeness Coverage is the proportion of input events that reached the output: rows in versus rows out, adjusted for known filtering. It catches the failures freshness cannot — a partition skipped, a shard whose consumer group stalled, a malformed batch quietly routed to a dead-letter queue. The practical trap is that coverage requires an independent count of the input. If you compute both sides from the same pipeline state, you learn only that the pipeline is self-consistent. Counting from the source system, or reconciling against a control total the source publishes, is what makes the indicator meaningful. ## Correctness Correctness is the proportion of outputs that are right, not merely present and timely. It is the hardest to measure and usually the most valuable, because silent corruption is the failure mode data systems are worst at detecting. Practical implementations: - **Reconciliation**: aggregate totals in the output must match a control total from the source within a tolerance. - **Invariants**: no negative balances, no duplicate primary keys, referential integrity across joined datasets, distributions within historical bounds. - **Sampled recomputation**: independently recompute a small random sample by a second path and compare. A correctness indicator is what separates a pipeline that has an SLI from one that has a cron job with monitoring. ## Durability, for the storage layer Where the system's job is to keep data rather than move it, durability — the proportion of stored objects still retrievable and intact — becomes the indicator. It is measured by continuous background verification of checksums against a sample, because by the time a user discovers a lost object the indicator is far too late. ## Why exit status is not an SLI Job exit status measures the orchestrator's experience, not the consumer's, and it fails the proportionality test in both directions: - **Green while broken**: the job reads an empty or partially written source, applies a filter that now matches nothing, writes zero rows, and exits 0. Every consumer sees yesterday's data. - **Red while fine**: a retryable step failed and the third attempt succeeded, or a cleanup task exited non-zero after the data was already correctly published. Nobody downstream noticed anything. Exit status is a useful operational signal for the team that runs the pipeline. It is not a statement about what anyone received. ## Aggregation: report the worst partition, not the mean A fleet-level freshness number computed as an average across partitions or tenants hides the exact failure you most need to see: one tenant's data stopped updating three days ago while the other nine hundred are fine. The average moves by a tenth of a percent. Two defensible shapes: - **Worst-partition** freshness: the indicator is the age of the *stalest* partition, so any single stalled shard is visible. - **Per-partition ratio**: the proportion of partitions that were fresh, which degrades proportionally to how many tenants are affected. Which to choose is a real decision. Worst-partition is unforgiving and will page you for one abandoned test tenant; the per-partition ratio is proportional but lets a small number of important customers sit stale under the threshold. The usual resolution is to tier the partitions and apply the strict form only to the ones that matter. ## Putting it together for the two examples *Nightly ETL*: freshness — output partition for day D published within N hours of midnight, measured on the data's own completion marker; coverage — output row count reconciles with the source's control total within a tolerance; correctness — daily aggregates match the source system's reported totals. *Streaming ingest*: freshness — the proportion of time end-to-end lag stayed under the bound, measured from the event timestamp to the moment it became queryable; coverage — the proportion of produced records that became queryable, counted independently at the producer; correctness — duplicate rate and invariant violations per window.
- Where should the freshness timestamp come from?From the data itself — an event-time watermark, a max event timestamp, or a partition completion marker written only after the data is queryable. Wall-clock time of the last successful run is the wrong source, because a job that runs on schedule and writes nothing keeps that timestamp fresh forever. Ideally the consumer computes freshness at read time, so it reflects what was actually served.
- How do you choose between worst-partition freshness and a proportion-of-fresh-partitions indicator?By what a single stalled partition costs. Worst-partition is right when any one tenant going stale is a real incident, and it will page you for an abandoned test partition unless you tier which ones count. The proportion form degrades in step with blast radius and suits large fleets of equal-value partitions, but it lets a handful of important tenants sit stale under the threshold.
- A pipeline's coverage indicator computes input and output counts from the pipeline's own state and always reads 100%. What is wrong?Both sides come from the same source of truth, so the indicator only proves internal consistency — a batch never read is missing from both counts. Coverage needs an independent count from the producing system, such as a control total the source publishes or the broker's own offset high-water mark, against which the output is reconciled.
saying these in an interview costs you the question
- The job exited zero, so the pipeline met its SLO
- Freshness means the scheduler ran on time
- Average freshness across partitions is a fine indicator
- Data correctness cannot be measured, only spot-checked
- Availability and latency cover every kind of service