skip to content

How do you turn an intermittent failure into a reliable reproduction?

level: middleimportance: should knowfreq 58%

answer

  1. Not random - a variable you missed
  2. Make failures leave artefacts
  3. Diff a failing run against a passing one
  4. Flip one candidate, count the attempts
  5. Confirm both on and off

basics

~20 s

Treat intermittence as a hidden variable, not randomness. Capture artefacts from a failing run, diff it against a passing one, then flip one candidate condition at a time and re-run a fixed batch until the hit rate moves.

solid answer

~50 s

An intermittent failure is one whose trigger you have not identified yet, so the work is variable isolation rather than repetition. First I make failing runs recordable: timestamps, logs, correlation identifiers, a recording, data snapshots. Then I diff a failing run against a passing one and list every difference as a candidate - concurrency and ordering, accumulated data, regional settings, cached state, account role, build. I fix everything at the passing values, flip one candidate, and run a fixed batch of attempts to get a hit rate. A candidate is confirmed only when it flips the rate in both directions: on, it fails most attempts; off, it fails none. If I run out of budget, I file what I have with the artefacts and the honest count - reproduced 3 times in 47 attempts under these conditions - rather than the word 'sometimes'.

code

pseudocode · 15 lines
pseudocode
candidates = [concurrent_accept, huge_trip_history, regional_units, warm_cache]
baseline   = fix_all(candidates, at = passing_run_values)

assert count_failures(baseline, attempts = 47) == 0

for c in candidates:
    trial = baseline with c set_to failing_run_value
    print(c, count_failures(trial, attempts = 47))

# concurrent_accept  -> 31 / 47
# huge_trip_history  ->  0 / 47
# regional_units     ->  0 / 47
# warm_cache         ->  0 / 47

confirm(concurrent_accept, on = 31 of 47, off = 0 of 47)

go deeper

for a junior

Recall that intermittent means a condition is varying that you have not spotted, and that the first move is capturing evidence - logs, timestamps, a recording, a data snapshot - so a failure you were not watching can still be examined afterwards.

for a middle

Explain the isolation mechanics: baseline everything at passing values, flip one candidate, run a fixed batch, read a hit rate, and confirm in both directions. Be ready to name the candidate families - concurrency, data shape, leftover state, environment, identity, build.

for a senior

Show judgement about the budget and about partial results: what a rate that improves but never reaches certainty tells you, when to stop and file with honest counts and eliminated hypotheses, and how you keep an intermittent finding from being quietly dismissed.

for a principal

Own the systemic side: whether your environments and data can even produce evidence for a failure nobody was watching, and how you stop teams from absorbing intermittence through retries and waits until the underlying defect is invisible.

### Intermittent means 'the trigger is a variable you have not found' Software does not roll dice. When a case fails on some runs and passes on others with no code change, something is varying between the runs and you have not identified it yet. Renaming that ignorance as randomness is the mistake that stops most investigations dead, because you cannot isolate something you have decided is unknowable. The whole method follows from refusing that framing. ### Step 1: make failure observable after the fact The first problem with an intermittent failure is that you are never watching when it happens. Before hunting, make sure a failure that occurs at 02:14 leaves evidence you can read at 09:00: - continuous logs with timestamps and correlation identifiers you can tie back to your attempt; - a snapshot of the relevant data immediately before and immediately after the attempt; - a screen or network recording that runs across all attempts, not just the one you expect to fail; - the exact build identifier and configuration for each attempt. This alone converts many 'cannot reproduce' cases into ordinary ones, because the artefacts show the difference you could not see live. ### Step 2: diff a failing run against a passing run With artefacts from both outcomes, enumerate every difference, however irrelevant it looks. Typical candidate families: - **Concurrency and ordering** - two operations landing inside the same second, a callback arriving before the state it updates was written, a retry overlapping the original. - **Data** - a record with an empty optional field, a very large history, a duplicate, a value at a boundary. - **Accumulated or leftover state** - a warmed cache, a session from a previous attempt, a queue that was not drained. - **Environment** - regional and unit settings, clock skew between machines, a different node behind a balancer. - **Identity** - role, permissions, team membership, account age. - **Build and switches** - a feature switch evaluated per account, a partially rolled-out change. ### Step 3: one variable at a time, measured as a rate Fix every candidate at its passing-run value to get a baseline that does not fail. Then flip exactly one candidate and run a fixed batch - twenty, forty, whatever your patience and the cycle time allow - and record how many attempts failed. Two disciplines make this work: **Change one thing.** If you flip two candidates and the rate jumps, you have learned nothing attributable. **Measure a rate, not an outcome.** With a base rate of 3 in 47, a single passing attempt proves nothing. Rates also let you recognise partial progress: a candidate that moves the rate from 3-in-47 to 19-in-47 is not the whole trigger but is certainly part of it, and it usually points at the family the rest lives in. Confirm in both directions before you believe it: with the condition on the case fails most attempts, with it off it fails none. A one-directional result is often coincidence dressed up as a finding. ### Step 4: know when to stop, and report honestly Isolation has a budget. When it runs out, the honest report is not 'intermittent' - it is the artefacts, the conditions you did control, the candidates you eliminated, and the count: reproduced 3 times in 47 attempts on this build with this data. Saying which hypotheses you ruled out is what stops the next person repeating your dead ends, and it is the difference between a report that gets investigated and one that gets closed as not reproducible. ### A worked example On a ride-hailing dispatcher, a completed trip occasionally settles against the wrong driver. Manual attempts fail 3 times in 47. Recording every attempt shows the failing ones share a pattern: two drivers accepted the same request inside the same clock second, and the second acceptance arrived while the first was still being written. Fixing everything else and forcing that overlap moves the rate to 31 of 47; removing it gives 0 of 47. The case is now reproducible on demand, and the interesting part is what the isolation revealed - the losing acceptance still writes a settlement record, so the corruption is silent: no screen ever shows a conflict and nothing in the product flags it. That finding, not the original 'sometimes settles wrong', is what gets the defect fixed.

  • What do you write in the report when the isolation budget runs out and the failure is still intermittent?
    The artefacts from the failing runs, the conditions I did hold fixed, the hit rate as a count of attempts, and the candidates I eliminated with the evidence that eliminated them. Naming the dead ends is the part people skip, and it is what stops the next investigator repeating three hours of my work. I also say plainly where I stopped, so nobody reads the report as a completed investigation.
  • A candidate condition moves the failure rate from 3-in-47 to 19-in-47 but never to certainty. How do you read that?
    As a partial trigger: that condition is involved but something else is still varying. It is a strong result, because it narrows the search to the family the condition belongs to - if overlapping acceptances raise the rate, the rest of the trigger is probably also about ordering or concurrency. I keep it fixed at the failing value and start the one-variable-at-a-time pass again inside that family.
  • Why is running the same unchanged case many more times a poor response to an intermittent failure?
    It gathers observations without gathering information. More attempts sharpen your estimate of a rate you already roughly know, but they cannot tell you which condition drives it, because nothing is being varied deliberately. Repetition is only useful once it is attached to a controlled change - the same batch run with one candidate flipped is what turns attempts into evidence.

It is the same move as tracing an intermittent rattle in a car: you do not drive further hoping to hear it, you strap on a recorder and then change one thing at a time - load, speed, surface - until the rattle appears on command.

saying these in an interview costs you the question

  • Calls it random and stops investigating
  • Adds a wait or a retry until it passes
  • Changes several conditions between runs
  • Judges a candidate from a single attempt
  • Reports 'sometimes fails' with no count or artefacts
  • Confirms a condition on but never checks it off

context