How do you decide whether a 0.4% move between last night's published total and the previous run's is real or a defect?
answer
- two questions: input, or job
- reconcile each run to its own input
- read the shape, not the size
- concentrated is source, proportional is systematic
- explain the remainder, never widen the allowance
basics
~20 sDecompose the delta before judging it: attribute the move per key and per period, then check whether the input moved too. A few keys shifting is usually the source; every key shifting proportionally is usually the job. Explain the remainder rather than widening the tolerance.
solid answer
~40 sStart by separating the two candidate causes: the input changed, or the logic's behaviour changed. Reconcile both runs against their own inputs — if the input totals moved by the same 0.4%, the job is faithful and the source moved. Then look at the shape of the delta rather than its size: concentrated in a few keys points at restated or late-arriving source records; spread proportionally across every key points at a systematic change such as a unit, a rounding rule or a reference-data remap; confined to the boundary periods points at records crossing a period edge. Whatever remains unexplained stays unexplained in writing. The failure mode to avoid is widening the tolerance until the check passes, because the next genuine regression will then be absorbed silently by the same allowance.
go deeper
Recall that a difference between two runs is normal when the input differs, and that the first check is whether the input moved by the same amount as the output.
Explain how to decompose a delta per key and per period, and which shapes point at the source, at reference data and at the transformation.
Demonstrate the discipline: name the benign causes, rule them in with evidence, record the explanation, and refuse to widen an allowance to silence a run nobody understood.
Own the policy — what magnitude of unexplained movement blocks publication, who may authorise publishing anyway, and how a tolerance change is recorded so the ratchet does not loosen unnoticed.
## Two questions hiding inside one number "The total moved 0.4%" bundles two separate questions: **did the input change**, and **did the job's behaviour change**? Answer them independently and the diagnosis usually falls out. - Reconcile **each run against its own input**: rows read, the grand total of the additive measure, the distinct key count. If the input totals moved by the same proportion, the job is faithful and the movement is upstream. - If the input totals are flat and the output moved, the change is in the job, in its reference data, or in something it reads that is not the main input. This question assumes the input genuinely differs between the two runs — it is a new period of data. A job that produces different answers from **identical** input is a different subject, repeatability, with its own causes. ## Read the shape, not the size The magnitude of a delta says almost nothing; its distribution says most of what you need. | shape of the delta | usual cause | |---|---| | concentrated in a handful of keys | source-side: restated records, a late-arriving correction, one customer's genuine activity | | spread proportionally across every key | systematic: a unit or currency change, a rounding rule, a changed conversion factor | | confined to the first and last period of the range | boundary: records moved across a period edge by a timestamp or timezone change | | a step change at one date, flat either side | a reference-data remap, or a source that changed shape on that date | | only in keys that are new or disappearing | population change, or a filter whose behaviour moved | This is why the comparison has to be keyed and per-bucket. A top-line percentage cannot be read; a per-key attribution can. ## The benign explanations worth ruling in early 1. **Late-arriving source records for a period already published.** The earlier run saw fewer records for that period than the later one does; recomputing a past period would move the earlier figure to match. 2. **Restated source records.** An upstream system corrected values that both runs read for the same period. 3. **Reference data that changed.** A category hierarchy, a rate table or a mapping was updated, so the same input rows now roll up differently. 4. **Approximate arithmetic.** Reordered summation of approximate types moves the last places of a total. This explains a delta of the order of the type's precision — it does not explain a tenth of a per cent. 5. **Population movement.** Real activity changed. This is the null hypothesis and it needs evidence like everything else. ## The defect explanations - A **partial read**: one input source, day or object was missing, so the run computed over less than it should have. The input row count exposes it if the reconciliation is per source. - **Fan-out or loss at a join**, from a reference table that gained or lost a duplicate. - **A key construction change** — trimming, case, a null substitution — that silently re-partitions the population between keys. - **A filter or bound that started rejecting more**, with the rejects counted but nobody reading the count. ## Unbounded input drifts by design Against a continuous input, a comparison of two consecutive snapshots is *expected* to be non-empty, and the interesting property is whether the movement is **confined to the periods still open**. A runtime that revises and re-emits a group will keep moving recent buckets until their **completeness claim** — the running assertion that no record older than a stated moment will still arrive — has passed them; movement in buckets well behind that point is a finding. Where continuous work runs as a rapid succession of small finite runs, the same reading applies to the accumulation of those runs, not to any single one. ## Do not widen the tolerance to make it green The tolerance exists to absorb arithmetic reordering of approximate types, and it should be set from the magnitude of *that* effect, not from the largest drift anyone has seen. Every time it is widened to silence a run that nobody explained, it is also widened for the next genuine regression, and the ratchet only goes one way. The discipline that holds is: - an unexplained movement is a finding, not noise, and is recorded as one; - the explanation is written down next to the run, so the next person comparing the same measure inherits it; - if the same benign cause recurs, the fix is to make it visible — attribute the delta to it explicitly — rather than to raise the threshold that hides it; - a tolerance change is a decision with a named reason and a date, like any other change to how the number is produced. Nothing here is about what a production monitor watches continuously; that is a neighbouring subject. This is the judgement made when two runs' outputs are put side by side and someone has to say whether the difference is acceptable.
- The delta is concentrated in three keys out of twelve million. What is the likely cause?Something source-side about those keys: restated records, a late-arriving correction, or genuine activity. Pull their input rows for both runs and compare directly — with the population that small, the answer is usually visible in the records themselves rather than in the aggregate.
- Both runs reconcile cleanly against their own inputs and the output still moved. What now?Something the job reads besides the main input changed: reference data, a rate or mapping table, a configuration value baked into the logic, or a dependency the reconciliation does not cover. Extend the reconciliation to those inputs so the next occurrence is attributed rather than investigated from scratch.
- When is a non-empty diff between consecutive runs the expected result?Whenever the input differs — a new period, late-arriving records for an earlier one, or a continuous job still revising buckets whose completeness claim has not passed. The question is never whether it is empty but whether the movement is confined to where movement is legitimate and its size is accounted for.
saying these in an interview costs you the question
- Judges a delta by its percentage without attributing it per key
- Widens the tolerance until the comparison stops failing
- Assumes any movement in a total must be a defect
- Never checks whether the input itself moved by the same amount
- Treats an unexplained difference as noise and publishes anyway
- Flags movement in still-open periods of a continuous job as a fault