Why keep the naive all-pairs checker in the test suite after the optimized version ships?
answer
- code called by tests is not dead
- one version is readable, one is fast
- randomly generated inputs, two implementations
- compare the two answers, not a fixture
- the slow obvious version is the oracle
basics
~10 sThe naive version is a test oracle, not dead code: short enough to read and believe, it judges the fast version's answers on randomly generated schedules and catches boundary bugs that hand-written cases miss.
solid answer
~50 sIt is not dead code — it is the oracle a differential test compares against. The all-pairs conflict checker is short enough that a reviewer can read it and believe it, which makes it a reasonable encoding of the specification; the optimized version is fast but full of boundary conditions. So the suite generates random schedules, runs both, and asserts identical verdicts. That finds the cases nobody thinks to hand-write: intervals that merely touch at an endpoint, zero-length entries, several bookings starting at the same instant, one interval fully containing another. I keep it out of the production path, cap the generated sizes so the quadratic oracle stays fast in the suite, and log the seed on failure so any discovery is reproducible. It earns deletion when the optimized version is retired or when the oracle stops being obviously correct.
go deeper
Know that a slow, obviously correct implementation can be kept as a reference for tests to compare against, and that code exercised only by tests is not dead code. Being able to name the idea is enough at this level.
Explain the mechanism: generate random inputs, run both implementations, assert the answers match. Be specific about why the naive one is trustworthy and about which boundary cases random generation finds that hand-written fixtures miss.
Demonstrate the operating discipline — bounded input sizes, seeds logged, failures shrunk, the oracle kept out of production code — and be able to defend the second implementation to a reviewer who sees only duplication.
Own the cost-benefit: two implementations of one requirement is ongoing maintenance, justified while the fast path is risky and the oracle stays readable. Set the retirement trigger explicitly, and be clear that this technique cannot catch a shared misreading of the specification.
## The reviewer's objection, and why it is wrong here The pull request adds a fast interval-overlap checker for a scheduling service and *also* keeps the original all-pairs version that compares every booking against every other. The reviewer asks the reasonable question: why is superseded code in this diff? The answer is that the two implementations play different roles. The optimized checker is what runs in production. The naive checker is a **test oracle**: an independently written, obviously correct implementation of the same specification, used to judge the answers of the one you cannot eyeball. Code called from tests is not dead code. ## Why the naive version makes a good oracle A useful oracle has one property above all: you can read it and believe it. The all-pairs checker is a double loop and one overlap condition. A reviewer holds the whole thing in their head, checks the condition against the specification once, and is done. The optimized version is where the bugs live. It probably sorts the bookings and sweeps them, or maintains an ordered structure of active intervals, and its correctness now depends on tie-breaking between equal start instants, on whether a touching endpoint counts as an overlap, on what happens when an interval is fully contained in another, and on the removal step being exactly right. Those conditions are precisely the ones a human forgets to write a test for — because forgetting them is what caused the bug. ## Differential testing, concretely The test generates random schedules — random counts, random start instants drawn from a small range so collisions are frequent, random durations including zero — runs both implementations, and asserts the verdicts match. Run a few thousand of those per suite execution and the boundary conditions get hit thousands of times each, without anyone enumerating them. The deliberate part is the input generator. Drawing start times from a *narrow* range is what makes ties and containments common; uniformly random instants over a huge range produce schedules that almost never overlap and test nothing. Property-based testing tooling exists in most ecosystems for exactly this, but a twenty-line generator and a loop gets most of the value. When the two disagree, the test must report a reproducible case. Log the seed and shrink the failing input — drop bookings one at a time while the disagreement persists — so the report is a three-interval counterexample rather than a schedule of four hundred. ## The discipline that keeps this from rotting **Bound the sizes.** The oracle is quadratic. Generating thousand-booking schedules turns a fast suite into a slow one; keep generated inputs small, which is also where boundary bugs live anyway. **Keep it out of the production path.** The oracle belongs in test sources. If application code can call it, someone eventually will, and the performance work is quietly undone. **A disagreement is not automatically the optimized version's fault.** Sometimes the oracle is wrong, and the failing case is the first time anyone has read the specification carefully enough to notice the ambiguity. "Do bookings that touch at an endpoint conflict?" is a product question; the differential test surfaces it, but the answer comes from the specification, and both implementations then get corrected. **Agreement is not proof.** Both implementations can encode the same misreading of the requirement, and the test will happily pass forever. Differential testing checks that the fast version matches the simple version; it does not check that either one matches what the business meant. Keep example-based tests written from the specification alongside it. **Know when to delete it.** Retire the oracle when the optimized version is retired, when the specification changes so much that maintaining two implementations costs more than the bugs it catches, or when the oracle itself has grown enough special cases that a reviewer can no longer read it and believe it. That last one is the important trigger: an oracle you no longer trust is worse than none, because failures get blamed on the wrong side. ## What to write in the review reply One comment on the naive implementation is usually enough: state that it is a reference implementation used as a differential-test oracle, that it is not reachable from production code, and that it is intentionally the slow, obvious formulation and should stay that way. Reviewers object to the second implementation because it looks like a leftover; naming its job in a comment converts it from clutter into a documented testing strategy — and stops a future maintainer from "optimizing" the oracle, which would destroy the only reason it exists. This is the long tail of "brute force first": the naive approach you stated in five minutes at the whiteboard has a second life as the thing that proves the clever one correct.
- The random schedules almost never overlap and the test finds nothing. What is wrong with the generator?The start instants are drawn from too wide a range, so collisions are rare and the interesting boundaries never occur. Narrow the range so ties and containments are frequent, include zero-length and identical intervals deliberately, and keep the generated counts small. The generator's distribution, not the iteration count, decides what gets covered.
- The two implementations disagree. How do you decide which one is wrong?Neither, until you check the specification. Shrink the failing case to the smallest disagreeing input, then read the requirement for that boundary — touching endpoints and zero-length bookings are usually ambiguities nobody resolved. Fix whichever implementation encodes the wrong reading, sometimes both, and add the shrunk case as a named regression test.
- If both implementations agree on millions of random inputs, is the optimized one proven correct?No. It is proven consistent with the oracle, and both can share the same misreading of the requirement. Differential testing eliminates implementation bugs in the fast version, not specification bugs. Keep example-based tests derived directly from the requirement alongside the differential ones.
- When should the oracle finally be deleted?When the optimized version it guards is retired, or when the oracle has accumulated enough special cases that a reviewer can no longer read it and believe it. An oracle nobody trusts is worse than none, because a disagreement then gets blamed on the wrong implementation and real bugs get dismissed.
It is the hand-wound reference clock kept in the back room: too slow to sell, but it is what you set the fast ones against.
saying these in an interview costs you the question
- Calls the naive implementation dead code and deletes it
- Says passing unit tests on the fast version is sufficient
- Runs the quadratic oracle on production-sized generated inputs
- Generates inputs so spread out that overlaps never occur
- Treats agreement between the two as proof of correctness
- Reports a failure without a seed or a shrunk counterexample
- Optimizes the oracle for speed, destroying its readability