How do you write up an intermittent defect that reproduces only some of the time?
answer
- A ratio needs its denominator
- Say which build and which tier
- Conditions that raise and lower the rate
- Capture from the failing run itself
- Trend before the event, uptime at failure
basics
~20 sReport failures over a stated number of attempts on a named build, the conditions that raise or lower that rate, evidence captured from a failing run, and what you ruled out. "Fails sometimes" is not a report.
solid answer
~50 sTurn the word "sometimes" into a number: run the scenario a fixed number of times and report failures over attempts, naming the build, tier, data set, concurrency and whether the harness retried, since retries hide the true per-attempt rate. Then convert randomness into conditions by varying one factor at a time and reporting the rate under each - serial versus parallel, fresh versus reused data, short versus long uptime. Capture artefacts from a failing occurrence specifically: correlation identifier, timestamped logs, trace, and the resource metrics for the window before the failure, since ageing and exhaustion show as a trend rather than an event. Preserve the failed environment before anything cleans it up. State what you ruled out, state impact separately from frequency, and state up front how many clean runs would count as a fix.
code
pseudocode · 9 linesattempts = 200
failures = 0
for i in range(attempts):
resetEnvironment()
run = execute(scenario, retries = 0)
if run.failed:
failures = failures + 1
keepArtifacts(run.id, logs = true, trace = true, resourceMetrics = true)
report(failures, attempts, build = "7.4.219", tier = "shared", concurrency = 8)go deeper
Know that "sometimes fails" is not reportable: give a failure count over a stated number of attempts, and say which build and which environment you ran them on.
Explain how to establish a rate, how retries in the harness distort the rate you observe, and why evidence has to come from a failing occurrence rather than a similar passing one.
Show that you convert randomness into conditions - varying one factor at a time, reporting a rate per condition, preserving the failed environment, and reading resource trends in the window before the failure.
Own the policy: what the organisation does with intermittent failures nobody has localised yet, how a fix is confirmed statistically rather than on one green run, and why impact rather than frequency should set urgency.
## Why "sometimes" is not a report An intermittent defect is one that does not reproduce on every attempt under conditions you believe are identical. The two things that make such a report actionable are a **count** and the **denominator** it was counted over. "Fails sometimes" cannot be prioritised, cannot be confirmed as fixed, and cannot be told apart from an issue already on file. "14 failures in 200 sequential runs on build 7.4.219" does all three: it is a measurement, it sets the bar for confirming a fix, and it can be compared with the next measurement. ## Measure the rate, and state what you measured over Run the scenario a fixed number of times under stated conditions and report both numbers. State the build, the tier, whether runs were serial or parallel, the machine class, the data set, and - critically - whether the harness retried. Retries are a common trap: if the pipeline retries twice on failure, a 7% per-attempt rate surfaces as roughly one report in two thousand, so the rate you quote must say which of the two it is. ## Separate the two kinds of "intermittent" Distinguish a product defect that behaves nondeterministically from an unreliable check that fails on unchanged code for reasons of its own - timing assumptions, ordering dependence, shared state, contention in the harness. The evidence differs: a product defect leaves traces in the system under test, while an unreliable check usually leaves them in the harness. Say which you believe it is and what supports the belief, and if you cannot yet tell, say that too, because the two go to different owners. ## Turn randomness into conditions Vary one factor at a time and report its effect on the *rate*, not merely on the outcome. A small conditions table is the most valuable thing an intermittent report can carry: one worker gives 0 in 50 and eight workers give 14 in 200; fresh data gives 0 in 50 and a reused fixture gives 12 in 60; four hours of uptime gives 0 in 100 and sixty hours gives 9 in 40. Even a crude table converts "random" into "load- and age-dependent", and that is already a diagnosis. ## Capture evidence from the failing occurrence You cannot predict which run will fail, so instrument every run and keep the artefacts of the ones that do: logs at a level you can actually read, a correlation identifier per run, timestamps with time zone, a recording or trace retained on failure only, and resource metrics for the window *before* the failure. Never attach a screenshot from a passing run because it "looks the same". And do not let anything reset the failed environment before you have captured it - an environment currently sitting in the failed state is the most valuable artefact you will get, and an automatic cleanup step destroys it routinely. ## Report the trend, not only the event Many intermittent failures are exhaustion or ageing in disguise, and those have a signature the failure moment alone does not show. Worked example: the accrual endpoint of a loyalty-points ledger returns an error on about 1 run in 37, apparently at random. The failing runs cluster late in a process's life - pool usage climbs roughly 3 connections per hour and never falls, and every failure lands once usage reaches the ceiling of 24. Written up as "1 in 37, random", it ages in the backlog. Written up as "rate rises with process uptime; pool usage climbs monotonically at about 3 per hour and every failure occurs at the ceiling of 24; a restart resets the rate to zero for the next seven hours", it is a resource leak with a location - and it explains why the failure always shows on the last deploy of a 3-week train, the one that has been up longest. ## What else the report needs State what you ruled out and how: the same rate on a second machine (not machine-specific), the same rate with client caching disabled (not client state), zero occurrences at concurrency 1 (concurrency-dependent). State impact separately from frequency - a 1-in-37 failure on a path that silently drops points is a data-integrity problem, while a 1-in-3 cosmetic flicker is not, and a weak candidate conflates the two. Finally, state your confirmation criterion up front: how many clean runs, under which conditions, you would accept as evidence that a fix worked. Without that number agreed in advance, an intermittent defect is closed on the first passing run and reopens weeks later.
- How many clean runs would you accept as evidence that an intermittent defect is fixed?Enough that the previously measured rate would almost certainly have shown itself, under the conditions that made it worst. At roughly 7% per attempt, twenty clean runs prove very little - about a one-in-four chance of seeing nothing even with no fix at all - while a few hundred is convincing. Agree the number and the conditions before the fix lands, otherwise the first passing run closes the report and it reopens weeks later.
- The failure reproduces on the shared tier but never locally. What do you report?Report the non-reproduction as data, not as an excuse. List the measured rate on each side, and what differs between them: build, configuration, data volume, concurrency, process uptime, network shape. Then attack those differences one at a time and report which one moves the rate. "Only on the shared tier" is not a diagnosis, but "reproduces at eight concurrent workers and never at one, on either environment" is, and it is reachable from the same evidence.
- How do you avoid losing evidence when you cannot predict which run will fail?Instrument every run and retain artefacts only from failures: per-run correlation identifiers, logs at a readable level, traces or recordings kept on failure, and resource metrics sampled throughout. Disable or record the harness retry policy so the observed rate means something. Most importantly, block automatic cleanup on failure - a machine still sitting in the failed state, with its memory, handles and open connections intact, is the most informative artefact available and is routinely destroyed within seconds.
saying these in an interview costs you the question
- Writes "happens sometimes" with no attempt count
- Attaches evidence from a passing run instead of a failing one
- Assumes intermittent automatically means low impact
- Blames the environment without ruling anything out
- Records only the failure moment, never the preceding trend
- Lets cleanup wipe the failed environment before capture