A fix lands for a tax-filing wizard defect that only reproduces on the 6-hour nightly run. How do you confirm it before closing?
answer
- One clean cycle can mean nothing
- Agree the criterion before the fix lands
- Force the condition; shorten the loop
- Replay the report, not the reduced case
- Record build, cycles and the agreed criterion
basics
~20 sAgree a confirmation criterion before the fix lands, replay the reported steps on a named build containing the fix, and get enough evidence that a pass means something — a forced-condition harness or a stated number of clean cycles — verified by someone other than the fixer.
solid answer
~50 sConfirmation testing means re-executing the failing case on a build that carries the fix; the hard part here is that one clean 6-hour cycle is not evidence when the failure appeared in 3 of 137 runs — a single pass was the likely outcome even without a fix. So I would do four things. Agree the **confirmation criterion in advance** (a forced reproduction, or N consecutive clean cycles), because deciding after the fact turns into wishful thinking. Build a **targeted harness that forces the triggering condition** — here, driving the clock relationship directly — so the loop is seconds, not hours. Replay the **reported steps**, not the reduced case the fixer used, on a **named build identifier**, and have a **verifier who is not the fixer** do it. Then keep the defect in a verifying state, not closed, until the agreed evidence exists, and record what was actually run.
code
pseudocode · 18 lines# historical signal: 3 failures in 137 nightly runs, each ~6h
# diagnosis: submission compared against a peer clock that can run ahead
criterion = agreed_before_fix(
forced_iterations = 1400,
boundary_offsets = [-2s, -1s, 0s, +1s, +2s],
fallback = "12 consecutive clean nightly cycles"
)
function confirm(build_id, reported_steps):
assert build_contains_fix(build_id)
for offset in criterion.boundary_offsets:
repeat criterion.forced_iterations / 5:
set_peer_clock_offset(offset)
result = run(reported_steps) # the reported path, not the reduction
assert result.final_page_present
assert result.filing_date == expected_filing_date(result)
record(build_id, criterion, iterations_run, environment, verifier)go deeper
Know that confirming a fix means re-running the case that failed on a build that contains the fix, and that a defect is not closed on the fixer's word alone.
Be ready to explain why one clean run of a rarely-failing case proves little, and to distinguish confirmation testing of the failed case from regression testing of cases that already passed.
Show how you make a rare failure cheap to confirm: forcing the triggering condition, agreeing an explicit criterion in advance, replaying the reported path, and recording build and run counts as evidence.
Own the policy for confirmation under release pressure: what evidence a close requires, who may waive it, how unverified fixes are surfaced in the release decision, and how confirmation continues after ship.
## What confirmation testing is Confirmation testing (also called re-testing) is the narrow act of **re-executing the case that failed, on a build that contains the fix, to establish that the reported behaviour is now correct**. It is distinct from the regression testing you do *around* a fix, which asks a different question: did this change break something that used to work. A verifier owes both, but they are not the same activity and conflating them is a common interview stumble. Confirmation testing has three inputs that must all be pinned: the **case** (the reported steps, not a reduction of them), the **oracle** (what correct looks like, taken from the report and the specification), and the **build identifier** (which artefact was run). Any verification that cannot name all three is not evidence. ## Why a rare, long-cycle defect breaks the naive approach Suppose the wizard drops the final page of a return, and the history shows the failure in 3 of 137 nightly runs, each run taking about 6 hours. The naive confirmation is: fix lands, tonight's run is clean, close it. That reasoning is worthless. At roughly a 2% failure rate, a single clean run was the overwhelmingly likely outcome **whether or not the fix works**. The observation and the hypothesis are barely correlated, so the test distinguishes nothing. There are only three honest ways out, and a strong answer names them: 1. **Make the failure deterministic.** The best outcome of diagnosis is a case that fails 100% of the time before the fix and passes 100% after. Here the diagnosis was a clock-skew artefact between two components; a harness that sets the clock relationship directly turns a 6-hour lottery into a sub-second, repeatable case. Confirmation then takes one run, and the case joins the suite as a permanent guard. 2. **Buy statistical confidence deliberately.** If the condition genuinely cannot be forced, state a number in advance — say 12 consecutive clean cycles, or 1,400 forced iterations of the narrow path — chosen so that a fix-that-did-nothing would very likely have produced at least one failure. Write down the reasoning, not just the number. 3. **Confirm by mechanism.** Where neither is available, verification shifts partly to evidence that the specific mechanism is gone: a log line or counter that fired on every historical failure and now cannot fire, an assertion added at the point of the skew that would fail loudly if the condition recurred. This is weaker, and it should be labelled weaker rather than dressed up as a pass. ## Sequence the decision before the fix, not after The most valuable senior habit here is agreeing the **confirmation criterion at triage or at assignment**, while nobody is under release pressure. After a fix lands, every hour of waiting has a cost and the temptation is to redefine the criterion downward. Written in advance, the criterion is a commitment; written afterwards, it is a rationalisation. ## Replay the report, not the reduction An engineer diagnosing a defect necessarily reduces it — strips the wizard down to the one page and the one comparison. That reduced case is the right thing to fix against and the wrong thing to verify against. Verification replays the **reported** steps: the full path a user takes through the wizard, the same data, the same environment class. Fixes that repair the reduced case and miss the reported one are common, especially when the reduction dropped a step that also mattered — the return keeps its final page now, but carries the wrong filing date because the same skew feeds a second field. ## Independence The verifier is not the fixer. This is not a trust rule, it is an oracle rule: the fixer's model of the defect is the model that produced the fix, so re-running against that model can only confirm internal consistency. A second person reading the report cold applies the reporter's oracle instead. ## What gets recorded A verification note that says "tested, works" is not usable a month later. Record the build identifier, what was executed, how many cycles or iterations, on what environment and data, and the criterion that was agreed. When the defect reopens — and rare ones do — that record is what tells you whether the fix was wrong or the confirmation was. ## And if it cannot be confirmed in time Say so explicitly rather than closing on hope. The defect stays in a fixed-but-unverified state, the release decision is made knowing that, and the confirmation continues after release with a named owner. A defect closed as verified without evidence is worse than an open one, because it removes the team's own signal that anything is still outstanding.
- How does confirmation testing differ from the regression testing you run around the same fix?Confirmation testing re-executes the case that failed to show the reported behaviour is now correct. Regression testing executes cases that already passed, to show the change did not break them. They answer different questions and can disagree: a fix can confirm cleanly and still break a neighbouring behaviour. A verifier owes both, and the scope of the second is a risk judgment about what the change touched.
- The release date arrives and the agreed confirmation evidence does not exist yet. What do you do?Leave the defect in a fixed-but-unverified state and make that visible in the release decision, rather than closing it as verified. The release owner then chooses knowingly, and confirmation continues after release with a named owner and the same criterion. Closing without evidence destroys the team's own signal that something is outstanding, which is worse than shipping with an honest open record.
- The forced harness passes but the nightly run fails again a week later. What does that tell you?That the harness models a condition adjacent to the real one — the fix removed the mechanism you reproduced, not the mechanism in production. Treat the harness as falsified rather than the report: reopen, compare what the nightly environment does that the harness does not, and widen the forced condition until it fails on the pre-fix build. A harness that never failed before the fix was never a confirmation instrument.
Confirming a rare defect with one clean run is like declaring a leak fixed because the ceiling was dry on a day it did not rain.
saying these in an interview costs you the question
- Closes the defect after one clean long-running cycle
- Lets the fixer verify their own fix
- Verifies the reduced case instead of the reported steps
- Cannot say which build was actually run
- Decides the confirmation criterion after the fix lands
- Records only that the case passed, with no run count