skip to content

When no independent expected value exists for an aggregate, which invariants and bounds still catch a wrong result, and which errors survive?

level: middleimportance: should knowfreq 55%

answer

  1. assert properties, not the value
  2. structural, arithmetic, relational
  3. envelope from history, physics, the input
  4. a re-derivation must take a different route
  5. plausible wrong numbers pass every bound

basics

~20 s

Assert what must be true of any correct result rather than the value itself: uniqueness at the declared grain, parts summing to the whole, keys present in the reference set, measures inside their physical range, and volume inside a band drawn from history. A plausible wrong number survives all of them.

solid answer

~50 s

Invariants are properties no correct run can violate whatever the data says — one row per declared grain, no null in a key, every output key present in the reference set, a non-negative measure never negative, components summing to the total they were split from. Bounds are plausibility envelopes drawn from history or from physics: today's row count within a band of the trailing median for the same weekday, a session shorter than a day, a rate between zero and one. Both are assertions inside the job's own test or its post-run check, distinct from rules declared outside the job and checked against what it published, which belong to the data-contracts subject. The residual is the point: an invariant only fires outside the envelope, so a wrong number that looks reasonable passes, and an assertion that re-derives the value through the same code as the job proves nothing at all.

go deeper

for a junior

Recall the difference between asserting a value and asserting a property: uniqueness on the grain, no nulls in keys, and measures that cannot be negative never being negative.

for a middle

Explain where each bound comes from — history, physics or the job's own input — and why a bound derived from the input is stronger than one derived from belief.

for a senior

Show you know the residual: plausible wrong numbers, shared wrong beliefs and slow drift all pass, so invariants sit behind reconciliation rather than replacing it.

for a principal

Decide which invariants are mandatory for every published dataset versus per-team, and who owns a failing one at three in the morning when the alternative is publishing nothing.

## Two things called an assertion, and only one is yours Before writing any of these, separate them. An **assertion inside the job's own test or post-run check** is written by the job's authors, runs with the job, and fails the run. A **rule declared outside the job and checked against what it published** is a different mechanism with a different owner, a different lifecycle and a different audience — this subject names it only to hand it away. Everything below is the first kind. ## Invariants: true of any correct result, whatever the data says **Structural** - exactly one row per **declared grain**, the column set that identifies an output row; - no null in any key column; - every key in the output present in the reference set it was supposed to come from, and nothing invented; - every declared column populated where the logic says it must be. **Arithmetic** - a measure that cannot physically be negative is never negative; - components sum to the total they were decomposed from — the per-category amounts add back to the overall amount; - a share or rate lies within zero and one; percentage columns within zero and a hundred; - a value derived two ways agrees with itself, **provided the two derivations are genuinely different routes**. **Relational** - every output row traces to at least one input key; - the per-key multiplicity of a join result matches the multiplicity the relationship allows. ## Bounds: plausibility, not proof A bound is an envelope you expect a correct run to sit inside: - **From history.** Tonight's row count within a band around the trailing median for the same weekday; the grand total within a band of the previous period, widened for known seasonality. Drawn from the series, not from a number someone remembered. - **From physics or policy.** A user cannot accumulate more than twenty-four hours of session time in a day; a discount cannot exceed the price; a count of active accounts cannot exceed the count of accounts. - **From the input.** A filtered sum cannot exceed the unfiltered sum of the same measure; the distinct keys in the output cannot exceed the distinct keys in the input. The third class is the strongest, because it is derived from the run's own input rather than from belief. ## The tautology trap The most common dead assertion re-derives the quantity **using the job's own code**: the check imports the same function, runs it over the same rows, and compares. It will agree with itself through any bug. A useful re-derivation takes a different route — a simple full aggregate over the raw input compared against the incremental or pre-aggregated path the job actually uses, or a total computed from a different column that should agree by construction. If you cannot describe how the two routes differ, the assertion is decoration. ## An invariant that is model-dependent "A cumulative series never decreases" is a natural invariant and it is wrong as stated for several runtimes. A runtime that revises a group and **re-emits** it will publish a lower value after a correction, and a runtime that emits speculative partial results early will publish a sequence that is not monotone at all. The repair is to place the invariant where it is genuinely true: on the **latest value per key once its completeness claim has passed**, not on the sequence of emissions. Where the output is a series of increments from a rapid succession of small finite runs, the invariant belongs on their accumulation rather than on any one increment. ## What survives all of it - **A plausible wrong number.** Bounds fire on the implausible; a figure five per cent wrong sits comfortably inside every band you would dare to set. - **A shared wrong belief.** The invariant encodes the author's model. A bug that shares that model — the same misunderstanding of what a row means — passes. - **Slow drift.** A bound drawn from the trailing series moves with the series, so a defect that degrades gradually is absorbed into its own baseline. - **A correct result over the wrong input.** Every invariant holds; the job read yesterday's file. - **Compensating violations.** Two errors of opposite sign inside one bucket net to a value inside the bound. That residual is why invariants are the second line, behind totals reconciled against the input, and why a leaf-level answer that stops at "we assert the data looks sane" is half an answer: the honest half names what the assertions cannot see.

  • Which bound is strongest, and why?
    One derived from the run's own input — a filtered sum cannot exceed the unfiltered sum, output distinct keys cannot exceed input distinct keys. It is computed from the same data the job read, so it holds regardless of seasonality, growth or a changed upstream, unlike a band drawn from history.
  • An invariant fails on one run out of fifty. How do you respond?
    Treat it as a finding, not as noise: either the invariant is wrong about the domain, or the run is wrong. Both outcomes are worth a change — a corrected invariant or a fixed defect. Silently loosening it until it stops firing converts a working check into decoration.
  • Why is 'the total never decreases from one run to the next' a risky invariant?
    It is false for any runtime that revises and re-emits a group, and for outputs that legitimately restate a past period after a correction. Assert monotonicity on the settled latest value per key, and even then only where the domain genuinely forbids a decrease.

saying these in an interview costs you the question

  • Re-derives the value with the job's own code and calls it a check
  • Sets a bound from a remembered number rather than the observed series
  • Treats a passing invariant suite as proof the result is correct
  • Asserts a monotone series on emissions that a runtime may revise
  • Confuses assertions inside the job's test with rules declared outside it
  • Widens a band whenever it fires instead of explaining the run