What does running two versions of a live job's logic over the same input, with only one publishing, prove that a test cannot?
answer
- two arms, one audience
- real input, not imagined cases
- measures difference, not correctness
- an identical output is not a pass
basics
~20 sRunning both versions over the same production input, with only one publishing, measures the change against real volume, real key distribution and real dirty records — evidence a test cannot give, because a test holds only the cases its author imagined.
solid answer
~50 sA **side-by-side run** means both versions of the transformation logic — same engine, same input, same period — executing at the same time, with only the current version's output reaching readers. A test proves the logic does what its author expected on the cases the author wrote down. The side-by-side run answers a different question: over the whole real input, on every key that actually exists, how many rows move and by how much. It surfaces the shapes nobody encoded — the key that holds a tenth of the rows, the one upstream source that sends a different unit, the record that only appears at period close. What it does not prove is that the new version is right. It tells you exactly how it differs from what is published today, which is the thing you then have to explain.
go deeper
Recall that the two versions must see the same input over the same period, and that only one of them may reach readers. If the inputs differ, the comparison says nothing about the change.
Explain why real input beats crafted input here: distribution, dirty records and volume. Then state the limit out loud — the comparison measures difference against today's output, not correctness.
Show that you plan the comparison window around the rare paths, and that you write down the expected set of moved keys before reading anything. Name what a healthy week of comparison still cannot exercise.
The tradeoff is evidence against cost and delay. Two arms double compute and postpone the change; say which classes of change earn that spend and which ship on tests alone, and make it a standing rule rather than a per-change argument.
## What a side-by-side run is A **side-by-side run** means two versions of the same transformation logic — the same engine, the same input, the same period — executing at the same time, with only one of them publishing. The current version keeps feeding the reports, tables and dashboards people already read. The candidate version reads the same records and writes its output somewhere nobody is looking. Nothing changes for any reader until someone has read the difference between the two outputs and decided to act on it. This is a different instrument from a test, and it answers a different question. Note that *version* here means a variant of the transformation logic you wrote, not a runtime or engine upgrade; upgrading the engine underneath a job is a separate subject entirely. ## What a test proves, and where it stops A test proves the logic produces the expected answer on the cases someone wrote down. That is valuable, it is cheap, and it should already exist before you reach for anything heavier. Its limit is not rigour — it is coverage of **input shapes**: - a test contains the records its author imagined; production contains the ones nobody imagined - a test runs on an input small enough to read; the change may only matter on the single key holding a tenth of the rows - a test asserts against a value the author chose; the side-by-side run compares against what is actually being published today - a test cannot tell you how many readers' numbers move, or by how much — and that is the question you will be asked before anyone lets you cut over | question | a test answers it | a side-by-side run answers it | |---|---|---| | does the logic handle the case I wrote down? | yes | not directly | | what happens on the records nobody anticipated? | no | yes | | how many published figures move, and by how much? | no | yes | | is the new version right? | only for the written-down cases | no | ## What real input carries that a crafted input does not Three things. **Distribution** — the true spread of keys, including the dominant one and the rare ones, which decides whether a change that is harmless on average is catastrophic on the key that matters. **Dirt** — nulls where the schema promised none, a unit that differs by upstream source, a record that arrived twice. **Volume** — which turns a one-in-ten-thousand difference into thousands of rows that are countable, sortable and attributable to a source. ## What it does not prove This is the half candidates skip, and the half an interviewer is listening for. 1. **It does not prove the new version is correct.** It measures the distance between two versions. Both can be wrong, and the current one frequently is — that is often why the change exists at all. 2. **An identical output is not a pass.** If the change was a fix, identical output means the fix never fired: either the condition it addresses does not occur in the period you compared, or the fix does not do what its author believes. 3. **It exercises the healthy path only.** A comparison over a week in which no machine died says nothing about what the two versions do when a piece of the input is recomputed after a worker is lost, or when a write is retried. 4. **It is not free.** Two arms consume roughly twice the compute and read the input twice, and the arm that must not publish has to be actively prevented from publishing — which is engineering work, not a setting. ## How long to run it, and what you read Run it for at least one full cycle of whatever the output's period is, and long enough to cross the rare paths: a period close, the day the slow upstream source delivers, the month with an extra weekend. Stop when the comparison stops producing new **kinds** of difference, not when it stops producing rows. Then read the difference by class rather than row by row, and write down beforehand which keys you expected to move, so the number you actually care about is the count of keys that moved and should not have. ## Where engines differ in what the second arm costs The shape of the candidate arm depends on the runtime model, and a claim that fits one model misleads on the next. Where the runtime executes a continuous workload as a rapid succession of small finite runs, the candidate arm is simply a second series of those runs and can be started and stopped freely. Where the runtime is record-at-a-time and the job carries a **retained set** — everything it still holds between records, such as counters, buffered join sides and the last value per key — the candidate arm begins empty, so its earliest outputs differ from the current arm's for that reason alone and have to be excluded from the comparison. For a periodic job over a finite input there is no warm-up at all, because each execution sees its whole period from beginning to end.
- The two arms produce identical output for a week. Is that good news?It depends on the intent. For a refactor meant to change nothing, it is exactly the result you wanted. For a bug fix it is bad news: the fix never fired. Check whether the compared period even contains the condition the fix addresses before concluding anything about the logic.
- How long should the two arms run before you decide?At minimum one full cycle of the output's own periodicity, including a period close and any rare path — the month-end run, the upstream source that only delivers occasionally. The honest stopping rule is that the comparison has stopped producing new kinds of difference, not that it has stopped producing rows.
- Is a side-by-side run worth it for a change that only renames a column?Usually not. The cost is real — double compute, double reads, and the work of keeping the second arm silent — so reserve it for changes that can move a published number. For a pure rename, a schema comparison plus the existing tests carry the same evidence far more cheaply.
saying these in an interview costs you the question
- Says a green test suite makes a side-by-side comparison unnecessary
- Treats an identical output as proof the new version is correct
- Lets the candidate version write to the published destination as well
- Compares the two versions over different input periods
- Assumes every difference found must be a defect in the new version
- Runs the comparison for one hour and calls it covered