How would you prove a corrected job is right over a year of history before any consumer sees its output?
answer
- two outputs, one period
- the comparison is the deliverable
- counts, aggregates, then per-key diff
- intended, explainable, unexplained
- cut over by moving the pointer
basics
~20 sRun it beside the existing output: the same stored input, the same period, written to a side destination nobody reads, then compared against what the live job already produced. The comparison, not the run, is the deliverable.
solid answer
~50 sGive the corrected logic its own destination with the identical shape and layout, and replay it over a period the live job has already produced. Then compare at three resolutions: row and group counts, aggregates per period and per key group, and a per-key diff on a sample. Classify every difference into three buckets - intended, meaning the defect being fixed; explainable, meaning a difference a correct job would also produce; and unexplained, which is a stop condition until it moves into one of the other two. Only then cut over, by repointing consumers at the new location rather than overwriting the old one in place, keeping the previous output long enough to reverse the decision, and telling consumers their historical numbers have been restated. The habit worth demonstrating is that a parallel run without a defined comparison is just a second job burning money.
go deeper
Know that corrected logic is run to a separate location first, and that the point of the exercise is comparing two outputs rather than producing a new one.
Describe the three resolutions of comparison - counts, aggregates, per-key diff - and say what each catches and misses, rather than reaching straight for a full row-by-row equality check.
Predict the fix's effect before running, classify every difference as intended, explainable or unexplained, name the differences a correct job legitimately produces, and cut over by repointing consumers while keeping a way back.
Decide how much history is worth recomputing at all, who is told that published numbers have been restated, and whether the team's pipelines are built so that a side-by-side run is cheap or so that it is an expedition.
## The shape of a parallel run The method is simple and the discipline is where candidates separate. Stand up a second destination with the same schema, partitioning and naming as the live one, and point the corrected logic at it. Replay stored history through it for a period the live job has already covered. Nothing reads the second destination. At the end you hold two datasets over the same period produced by two versions of the logic, and the work now begins. The common failure is to treat the run itself as the evidence. A completed replay proves the job does not crash. It says nothing about whether its numbers are right, and the whole reason for a parallel run is that the numbers are the thing in doubt. ## Choosing the comparison window Pick a period long enough to contain the conditions you care about and short enough to compare properly. Useful choices: a period known to contain the defect, a period known **not** to contain it, a period containing a month boundary or a holiday, and a period during which the live job itself failed and was restarted. A corrected job that matches on a quiet fortnight has demonstrated very little. ## Three resolutions of comparison | Resolution | What it catches | What it misses | |---|---|---| | Row and group counts per period | wholesale loss, duplication, a shifted boundary | offsetting errors that preserve counts | | Aggregates per period and per key group | systematic drift, a rescaled or reclassified population | errors confined to a few keys | | Per-key diff on a sample, plus every key above a value threshold | the individual wrong row | anything outside the sample | Run all three. The first two are cheap and catch the large failures; the third is the only one that ever shows you **why**. ## Differences that are not defects A row-for-row identity check against the live output will fail on a perfectly correct job, and knowing why is the senior part of this question: - **The replay is more complete.** Records that arrived too late to be counted live are present in stored history from the start, so groups that were short live are whole in the replay. That difference is the archive being better than the original, not an error. - **Emission counts differ.** Where a runtime emits an updated running result on each arrival, the live run may have produced many rows per group and the replay produces one, because the group was complete almost immediately. Where a runtime emits once per group, counts match but the emission times do not. Compare final values per key, not the stream of rows that produced them. - **Ordering within a group is not reproducible.** Where the result depends on arrival order within a key - a last-value-wins field, an unordered concatenation - the two runs may legitimately disagree. If that matters, the logic has a determinism problem that the comparison has just found for you. - **The intended fix.** The defect being corrected must show up, in the direction and roughly the magnitude predicted before the run. A fix whose effect is smaller or larger than expected is a finding, not a success. Write these categories down before running the comparison. Deciding after the fact which differences are acceptable is how a wrong job gets shipped. ## The cutover and the way back Cut over by moving consumers, not data: repoint the name, view or pointer they read at the new location, in one step, so nobody ever reads a half-replaced dataset. Keep the previous output for long enough that reversing the decision is a second pointer move rather than another multi-day replay. And publish the restatement - which period changed, by how much, and why - because consumers who discover on their own that last year's numbers moved will not trust the next set either. One caution on cost: a parallel run doubles the compute for the period compared, and its writes land on a destination that has its own limits. The comparison window is a budget decision as much as a statistical one.
- The corrected replay differs from the live output on thousands of keys. How do you avoid rationalising the differences?Predict before the run: state which periods and which kinds of key the fix should move, in which direction and by roughly how much. Then classify every difference against that prediction as intended, explainable by a known mechanism such as late records now being present, or unexplained. Unexplained differences block the cutover until each is reclassified with evidence, not argument.
- Is a parallel run over the full history always necessary, or can a sample do?A sample is usually the right first pass, because it costs a fraction and finds most defects. A full run earns its cost when the output is a corrected dataset consumers will actually read, when the defect is concentrated in rare conditions a sample may miss, or when the destination's totals must reconcile exactly. Sample to decide, full run to publish.
saying these in an interview costs you the question
- A completed replay is itself the evidence that the logic is right
- Expects a correct fix to match the live output row for row
- Writes the corrected output straight over the live tables
- Decides after seeing the diff which differences are acceptable
- Deletes the previous output as soon as the new one is published