What do freshness and volume SLOs add to a data contract that schema checks miss?
answer
- the shape was right, the data wasn't
- an empty table is perfectly typed
- a job can succeed and write nothing
- staleness measured on event timestamps
- a row-count floor and a ceiling
basics
~20 sSchema checks prove the shape is right; freshness and volume SLOs prove the data actually arrived. A contract states a maximum staleness and an expected row-count range, so a perfectly-typed but empty or half-loaded dataset still raises an alert.
solid answer
~50 sEvery structural check can pass on a table that received nothing. A **freshness SLO** commits the producer to a maximum lag — the delay between when an event happened at the source and when its row is queryable in the target — measured on data timestamps, not on whether a job exited zero. A job can succeed, log nothing unusual, and write zero rows. A **volume SLO** states the expected row count for a window, usually as a floor and a ceiling with allowance for weekly seasonality; it catches partial loads, an upstream filter someone tightened, and a source outage that produced an empty but valid file. Together they convert the operational half of the contract into something you can alert on and report against, and they give consumers a defensible answer to "is this table safe to use right now?"
code
yaml · 11 linessla:
freshness:
measured_on: placed_at # event time, UTC — not ingest time
max_lag: 30m
quiet_hours_heartbeat: true # low-volume windows need a pulse, not a max()
volume:
window: 1d
compare_to: same_weekday_last_week
tolerance: -20%/+20%
absolute_floor: 1
on_breach: flag_degraded # alert | flag_degraded | fail_closedgo deeper
Know that a data contract promises more than column types: it also promises the data will be there and roughly how much of it, so an empty table counts as a violation.
Explain freshness as measured lag between event time and queryable time rather than job success, and volume as a floor plus a ceiling compared against a seasonally comparable window.
Demonstrate operating these without drowning in noise: heartbeats for low-volume datasets, tolerance bands negotiated with the producer, and a stated breach behaviour of alert, flag degraded, or fail closed per dataset.
Own where the numbers come from and what they commit the organisation to. Setting a freshness target is a capacity and on-call decision, and every fail-closed clause you sign trades availability of the data for trust in it.
## Why structure is only half a contract A schema check answers *is this the right shape?* It cannot answer *did the data show up?* Those are independent failure modes, and the second is the more common outage. A table with zero rows has correct types, correct nullability and correct field names; every static assertion passes. So does a table that received 40% of yesterday's orders because one source shard was unreachable. The contract's service-level clauses exist to cover exactly this gap. ## Freshness: measured on data, not on jobs The most consequential detail here is what you measure. The naive implementation alerts when the pipeline job fails. That misses the whole class of silent failures: the job runs, succeeds, and produces nothing or too little. The correct implementation measures **lag** — the difference between now and the maximum event timestamp (or source commit timestamp) present in the target. Define the clock carefully in the contract, because there are three different timestamps in play: - **Event time** — when the thing happened in the real world. - **Source commit time** — when it was durably written in the producing system. - **Ingest time** — when your pipeline wrote the row. A freshness commitment of "at most 30 minutes" is only meaningful once the contract says which pair it spans. Event-time-to-queryable is the promise consumers care about and the hardest to hold, because it includes any delay in the producing system itself. Source-commit-to-queryable is the promise your pipeline can actually control. Watch out for two traps. First, a low-volume dataset looks stale during quiet hours simply because nothing happened; contracts for those datasets need a heartbeat mechanism or a window-scoped rule rather than a naive max-timestamp check. Second, freshness measured against *ingest* time is self-fulfilling — if the pipeline stamps every row as it writes it, the lag metric can never detect that the source stopped sending. ## Volume: a floor, a ceiling, and seasonality A volume SLO states how many rows the contract expects per window. Both bounds matter: - The **floor** catches empty loads, partial loads, a filter that got tighter, an upstream shard that vanished, a truncated file. - The **ceiling** catches duplicate loads — the same file processed twice, a replay run twice, a broken deduplication step. Expressed as a fixed number, a volume SLO is noisy from day one, because real data has weekday/weekend and seasonal patterns. Practical contracts express the expectation relative to a comparable window: the same weekday last week, a rolling median with a tolerance band, or an explicit min for datasets where any drop is worth waking someone. The tolerance is a negotiation with the producer, and writing it into the contract is what makes the resulting alert legitimate rather than the data team's private opinion. ## Where these checks run Unlike the schema diff, these cannot run in CI — they are properties of live data, so they run continuously against the landed dataset and report against the contract version. That makes them detection rather than prevention, which is fine: nobody can prevent a source outage in a pull request. What the contract adds is that the threshold was agreed by the producer in advance, so a breach is a violation of a commitment rather than an argument about whether the number was ever reasonable. ## Consequences of a breach A good contract also says what happens on a breach, and the choice is a real design decision. Options range from **alert and continue serving stale data**, through **flag the dataset in the catalog as degraded** so consumers and downstream models can decide, to **fail closed**: hold the downstream build so no dashboard renders a partial day. Failing closed is right for financial reporting and wrong for an exploratory table, and stating which per dataset is part of what the contract buys. ## What these SLOs still cannot catch Be candid about the limit. Correct volume and low lag say nothing about correctness: the rows can arrive on time, in the right quantity, with the right types, and carry a value whose meaning changed because the producer altered a business rule. That is semantic drift, and it needs field-level distribution and reconciliation checks rather than the two operational SLOs discussed here. Freshness and volume are the cheapest, highest-yield pair — not the whole quality story.
- Why is alerting on pipeline job failure not the same as a freshness SLO?Because the worst failures succeed. A job can exit zero having written zero rows — the source returned an empty page, a filter excluded everything, a shard was skipped. Freshness measured as lag against the newest event timestamp in the target detects that; a job-status alert never fires.
- Why does a volume SLO need an upper bound as well as a lower one?The ceiling catches duplication. A file reprocessed twice, a replayed backfill, or a broken dedupe step doubles the rows while every schema check and the freshness lag stay perfectly healthy. Without a ceiling, over-delivery is invisible until someone notices revenue doubled.
Schema checks are the parcel's declared dimensions; freshness and volume are whether it arrived today and whether the box was full.
saying these in an interview costs you the question
- Treating a successful job run as proof the data is fresh
- Measuring freshness against ingest time the pipeline stamps itself
- Setting a fixed daily row count that ignores weekly seasonality
- Only setting a volume floor, so duplicate loads go unnoticed
- Claiming on-time correct-volume data must therefore be correct