Two versions of a job's logic ran over the same input and thousands of grouping keys differ — how do you decide which differences are the change working?
answer
- a non-empty result is the point
- predict the moved set first
- classify, then attribute to an input
- count what moved unexpectedly
basics
~20 sWrite down which keys the change should move, and why, before reading anything. Then classify every differing key against that prediction: the number that matters is how many moved for a reason you did not predict, not how many moved.
solid answer
~50 sThis is the two-version comparison, whose rows are what you are trying to explain before cutting over — not a comparison against a stored expected result, which is supposed to come out empty. Here a non-empty result is the point, so counting rows tells you nothing. The move that works is **predict, then classify**. Before looking, state which keys the change should touch and roughly what share of the total that is. Then sort every key into four buckets: present only in the current output, only in the candidate's, present in both but differing, and equal. Attribute the differing ones by joining them back to an input attribute — if every moved key traces to the one upstream source you changed handling for, the change is doing what you think. If the moved keys are spread evenly across everything, you changed something you did not intend to.
code
python · 28 lines# Two arms of a side-by-side run, keyed by the grouping key.
# current[k] and candidate[k] hold the settled value for key k.
def classify(current, candidate, predicted_to_move, tolerance=1e-9):
buckets = {
"only_current": [],
"only_candidate": [],
"equal": [],
"moved_as_predicted": [],
"moved_unpredicted": [],
}
for key in set(current) | set(candidate):
a = current.get(key)
b = candidate.get(key)
if b is None:
buckets["only_current"].append(key)
elif a is None:
buckets["only_candidate"].append(key)
elif abs(a - b) <= tolerance * max(1.0, abs(a)):
buckets["equal"].append(key)
elif key in predicted_to_move:
buckets["moved_as_predicted"].append(key)
else:
buckets["moved_unpredicted"].append(key)
return buckets
# The first number to look at is len(buckets["moved_unpredicted"]).
# It is the one the change was not supposed to produce.go deeper
Recall that the two outputs must be compared by key, not by row position, because the order and the number of output files are not stable between executions.
Explain the four classes — only in one side, only in the other, differing, equal — and why the equal rows are the part that carries no information about the change.
Show that you write the prediction first and then treat the unpredicted movers as the real finding. Name the noise floor from the job's own non-determinism, and say what a healthy comparison window never exercises.
Decide what standard of explanation is owed before a cutover, and who signs it off. Requiring every row explained stalls every change; requiring only a total stalls nothing and catches nothing. Set the bar and make it the same one each time.
## The comparison is supposed to be non-empty Two senses of *difference* collide on this subject and they want opposite outcomes. A comparison of a job's output against a stored expected result must come out empty, or the job is broken. The two-version comparison discussed here must come out non-empty whenever the change was meant to change a number, and its rows are exactly the thing you have to explain before you are allowed to cut over. Counting them proves nothing on its own. A thousand differing keys can be a perfect change and three can be a disaster. The mechanics of comparing large outputs — comparing by key rather than by file position, tolerating floating-point drift, normalising before comparing — are general output-checking craft and apply here unchanged. What is specific to a change under way is everything that follows the comparison: attribution. ## Predict before you look Write down, before reading a single row: - **which keys** the change should touch, stated as a rule you can evaluate — for example, every key whose records include the one source whose unit handling changed - **roughly what share** of all keys that is, as an order of magnitude - **in which direction** the affected values should move, and roughly by how much - **what should not move at all** This is the whole method. A prediction written afterwards is not a prediction; it is a rationalisation, and it will accommodate whatever you find. ## Classify, then attribute Sort every key into four classes and read the counts against the prediction: 1. **Only in the current output.** The candidate dropped it. Intended for a change that filters; alarming otherwise. 2. **Only in the candidate's output.** The candidate created it. Intended for a change that stops over-filtering; alarming otherwise. 3. **In both, values differ.** The main body. Split it further by whether the key was in the predicted set. 4. **In both, values equal.** The part you already understand; it carries no information and should not be read. Then attribute class three. Join the differing keys back to something about their input — the upstream source, the period, a category, whether a field was null. A change doing what you think produces moved keys concentrated on the attribute you changed handling for. A change doing something else produces moved keys spread uniformly, which is the single clearest signal that you have touched a shared step rather than the one you meant to. ## When everything moved If you predicted two percent and a hundred percent moved, the first hypothesis is not that the logic is catastrophically wrong. It is that something shared and cosmetic changed underneath: a rounding rule, a precision, a period label rendered in a different time zone, a trailing-space or null-versus-empty convention, a currency or unit formatting step every row passes through. These move every row while changing nothing meaningful, and they are far more common than a logic change that genuinely reaches every key. Normalise for them and compare again before drawing conclusions. ## Reading a comparison too large to read You will not read a million rows, so read summaries in this order: counts by class; the distribution of relative change among the differing keys; the top contributors by absolute change, because those are the ones a reader will notice; and a small random sample of each class, because the top contributors are systematically unrepresentative. ## What the comparison still cannot tell you Two hard limits, and an interviewer is listening for both. - **It cannot say which version is right.** It is evidence for a conversation, not a verdict. The published version is the baseline, not the truth — it is frequently the thing being fixed. - **It cannot surface defects that only appear under failure or under rare data.** A healthy comparison window never exercises a worker lost mid-run with its piece recomputed, a write retried after a failure, a **redistribution** under one enormous key — the point where a step needs records currently held by other workers, so every worker writes its records out and every worker fetches the ones addressed to it — or the period-close path that runs once a month. Non-determinism inside the job itself is a further limit: if the logic breaks ties by arrival order, or sums floating-point values whose order depends on how records were delivered, then two executions of the *same* version can disagree, and that noise floor has to be known before any of the two-version rows mean anything. Runtime model changes what settles, too. Some engines emit a group's result once when it closes; others emit a speculative result early and corrections afterwards; one common model emits an updated running result on every input. Compare settled final values per key, or you will be comparing one arm's intermediate emissions against the other's final ones and explaining a difference that does not exist.
- You predicted two percent of keys would move and a hundred percent moved. What do you check first?Whether something shared and cosmetic changed: a rounding rule, a precision, a period label in a different time zone, a null-versus-empty convention. These pass every row through and move all of them while changing nothing meaningful, and they are far commoner than a logic change that genuinely reaches every key.
- The comparison has been clean for a week. Which defect classes has it still not ruled out?Anything that needs failure or rare data: a worker lost and its piece recomputed, a write retried after a crash, the period-close path, a category that only appears seasonally, one key large enough to behave differently from the rest. A healthy week exercises the happy path and nothing else.
- Why compare settled values per key rather than every row each arm emitted?Because emission behaviour varies by runtime. Some engines emit a group once when it closes, some emit an early speculative result then corrections, and one common model emits an updated running result per input. Comparing raw emissions pits one arm's intermediate output against the other's final output.
saying these in an interview costs you the question
- Expects the two-version comparison to come out empty
- Compares the two outputs line by line in file order
- Reads the differences before writing down which keys should move
- Concludes the published version is correct because it is published
- Says a clean week of comparison rules out every defect
- Judges the change by how many rows moved rather than which