A teammate replaces the access-log summary loop with a transformation chain; what do you check before agreeing the results match?
answer
- a refactor is a claim about equivalence
- every skip path, not just the first
- the skip path often did a second job
- empty and fully skipped input
- diff both versions on a captured sample
basics
~20 sCheck that the keep-stage predicate matches every skip path the loop had, that the same field is read, that empty and fully skipped input give the same answer, and that no side effect the loop performed while walking was quietly dropped.
solid answer
~50 sTreat it as a refactor with a claim attached: same input, same summary. The usual differences are not in the counting but at the edges. Every skip path in the loop has to be reproduced by the keep-stage predicate, and a loop often has more than one - a malformed line, an out-of-scope path, an early return. Anything the skip path *did* besides skipping, such as incrementing a rejected counter or emitting a warning, disappears silently into a keep-only stage unless it was re-added. Then check the inputs nobody tries: an empty log, a log where every line is skipped, a log with one line. Those are where a missing starting value, an absent key versus a zero, and a reported total differ. Finally run both over the same captured sample and compare the summaries rather than reasoning about them.
code
pseudocode · 11 linesfunction summarise(lines)
counts = empty table
rejected = 0
for each line in lines
if not isWellFormed(line)
rejected = rejected + 1
continue
status = statusOf(line)
counts[status] = valueOr(counts, status, 0) + 1
report("rejected", rejected)
return countsgo deeper
The thing to remember is that a rewrite has to reproduce every skip the loop had, not just the obvious one at the top of the body.
Be able to list the edges where the two versions diverge: multiple skip paths, an empty or fully skipped log, and an absent key versus a reported zero.
Demonstrate that you verify rather than reason - run both over a captured sample containing the awkward records, and name the skip-path side effect as the usual silent loss.
The lead's call is how much assurance a summary of this kind deserves: a pinned test, a shadow run for one release, or nothing, depending on what consumes the numbers.
The rewrite arrives in review as a tidy diff: a dozen lines of loop become four stages. It is a refactor, which means it carries a claim - that the two compute the same summary for every input. Reviewing it is a matter of knowing where that claim usually fails, because it almost never fails in the counting itself. ## Where the two versions actually diverge 1. **Selection.** Walk every path in the loop body that avoids the accumulation - a skip on a malformed line, a skip on a path the summary excludes, a condition nested inside another - and check that the keep-stage predicate is the conjunction of all of them. Loops accumulate skip paths over years, one incident at a time, and a rewrite typically reproduces the one at the top of the body. 2. **Side effects on the skip path.** This is the classic silent loss. The skipped branch often did a second job: counted rejects, emitted a warning, bumped a health metric. A keep-only stage discards without recording, so the number simply stops existing and nothing fails. 3. **The field and the key.** The loop read one field to form the key and the chain must read the same one, with the same normalisation - a status held as text and a status held as a number do not group together. 4. **Empty and fully skipped input.** An empty log, and a log where nothing survives selection. This is where a missing starting value, an absent key where the loop reported a zero, or an error on an empty collection shows up. 5. **Order dependence.** Did the loop rely on lines arriving in order - last write wins into a table, a first match, a running boundary? Stages usually preserve order, but not every combining step does, so the assumption should be stated rather than inherited. 6. **Early exit.** If the loop could return part-way through, the chain has to reproduce both the stopping condition and what was returned at that point. | In the loop | What to look for in the chain | |---|---| | several skip paths | one predicate that covers all of them | | a counter bumped while skipping | a stage or reduction that re-creates the number | | an early return | an equivalent stop, and the same value returned | | a zero written for a missing key | the same absent-versus-zero convention | ## Check it, do not reason about it The fastest review is not a careful read. Capture a real sample that contains the awkward records - malformed lines, an unusual status, a line at each boundary of any range the code cares about - run both implementations over it, and compare the summaries key by key. If the summary feeds anything that matters, it is cheap to run both for one release and alert on a difference; the discrepancies that survive a code review are exactly the ones a sample was never built for. A test that pins the empty case and the all-skipped case is worth more than a test over a thousand ordinary lines, because both versions agree on ordinary lines by construction. ## What a reviewer should not demand The review is about equivalence of results, and that is a narrower question than whether the rewrite was a good idea. Cost is a separate conversation with its own evidence. So is style. Conflating them is how a review of a correct rewrite turns into an argument, and how a genuinely lossy rewrite gets waved through because the discussion was about something else. ## The judgment being tested Asked this way, the question is not about transformation chains at all - it is about whether you review a refactor by its claim. A strong answer names the skip-path side effect and the empty input without prompting, proposes comparing both versions on a captured sample instead of reading harder, and separates "does it compute the same thing" from "was it worth doing".
- The summary is identical on every sample you have. What might still differ?Inputs your samples do not contain: an empty log, a status never seen before, a line at the edge of a range the code buckets on. Also anything the loop did that the summary does not show - a metric, a log line, a cache it warmed - which no comparison of results can detect.
- How do you decide whether an absent key or a zero is the right answer?By asking what the consumer does with it. A dashboard that plots a series usually wants the zero; a report that lists observed statuses wants the key absent. Whichever the loop produced is the current contract, so a rewrite that changes it is a behaviour change and should be reviewed as one.
saying these in an interview costs you the question
- Assumes a keep-only stage preserves what the skip path also did
- Checks only the ordinary path and never an empty or fully skipped log
- Reproduces the first skip condition and misses the others
- Treats an absent key and a zero count as interchangeable
- Argues about cost and style instead of whether results match