skip to content

Why must a harness capture failure evidence during the failing run rather than let an engineer reproduce it?

level: middleimportance: must knowfreq 58%

answer

  1. Think about when the evidence stops existing
  2. Teardown happens before anyone reads the report
  3. A re-run is a second, different execution
  4. Intermittent failures may not recur at all
  5. Wire capture into the harness's failure path

basics

~20 s

The failing execution is the only one holding the evidence. Teardown discards the product state the case saw, an intermittent failure may not recur, and a re-run meets different data and timing. Capture therefore belongs in the harness's failure path.

solid answer

~50 s

A failure is a property of one execution, not of the case. By the time anyone reads the report the process is gone: the product state the case observed has been torn down, the parallel test worker's slot has been reused, and the clock has moved on. Re-running does not recover that evidence — it produces a *second, different* execution against data that has since changed, and an intermittent failure may simply pass. So the harness treats the moment of failure as the only chance it gets: on the failing path it records the expected and actual values, the ordered log of interactions and responses the case produced, a snapshot of the observable interface or process at that instant, and the identifiers that tie the case to the system's own records. Wiring that into the harness rather than into each case is what makes it reliable.

code

pseudocode · 21 lines
pseudocode
# harness lifecycle: runs after every case, whatever the outcome
on_case_finished(case, outcome):
    if outcome.passed:
        return                                  # green cases pay nothing

    bundle = new EvidenceBundle(caseId = case.id, runId = run.id)

    guarded(bundle, "assertion", deadline = 1s):
        bundle.put(outcome.expected, outcome.actual, outcome.divergedAt)
    guarded(bundle, "steps", deadline = 1s):
        bundle.put(case.interactionLog.drain())      # ordered, with durations
    guarded(bundle, "view", deadline = 5s):
        bundle.put(appSession.renderCurrentView())
    guarded(bundle, "identifiers", deadline = 1s):
        bundle.put(case.correlationId, run.workerIndex, clock.nowWithZone())

    evidenceStore.write(bundle)      # written BEFORE teardown releases anything
    teardown(case)

# guarded() records that a capture failed instead of throwing,
# so a broken capture never replaces the case's real failure

go deeper

for a junior

Be ready to say what you personally need to see when a case you wrote fails: what was expected, what arrived, and which steps ran before it. Know that the report has to answer that without you running anything.

for a middle

Explain the mechanics: a harness lifecycle step fires on the failing path, gathers evidence while the product state still exists, and writes it out before teardown. Be able to name what is ephemeral and what can be reconstructed later.

for a senior

An interviewer at this level expects the argument from real incidents — why a re-run is a different execution, why intermittent failures make in-run capture non-negotiable, and how you stop the capture itself from hanging, throwing, or perturbing the timing it records.

for a principal

Own the standard rather than the code. Capture is a property the whole harness confers, so a case that fails undiagnosably is a framework defect, not an author's oversight. Be ready to say what evidence level you mandate and what it costs to run.

## A failure belongs to one execution, not to the case When a case fails, what failed is not "the case" in the abstract. One **execution** of it failed: at one instant, against one state of the product, on one machine, with one set of data. Everything that would explain the failure lives inside that execution — the values in scope, the interface as it stood, the responses the system returned, the elapsed times, the identifiers that tie the case to the system's own records. The moment the case finishes, most of that starts to disappear. Teardown rolls back or discards the data the case created. The session that was operating the application is closed and whatever it held is released. The parallel test worker takes the next case and reuses its slot. On a disposable executor the whole working directory goes away with the machine that had it. The report a human eventually reads is written after all of that has happened. So the harness gets exactly one chance to preserve the evidence, and it is inside the failing run. ## What a re-run actually gives you The instinct to "just run it again and watch" treats a re-run as a replay. It is not one; it is a second, independent execution that happens to use the same instructions. | What you hope a re-run reproduces | What actually changes | |---|---| | The same data | New records, new identifiers, a store that has moved on | | The same timing | Different load, different scheduling, different latency | | The same build | Possibly a redeployed system with different configuration | | The same target | A different deployed instance, or a local one that does not match it | | The same outcome | An intermittent failure may simply pass | That last row is the decisive one. A failure that appears in one run of forty is the failure you most need evidence for, and it is exactly the one a re-run will not hand you. A suite whose diagnosis strategy is "run it again" can only diagnose its deterministic failures — the ones that were easy anyway. ## What the harness captures on the failing path 1. **The comparison itself.** The value the case expected, the value that arrived, and the path at which they diverged. Not a bare statement that two things differed. 2. **The interaction log.** The ordered steps the case performed, each with its own outcome and its own elapsed time, so a step that quietly failed earlier is visible. 3. **A snapshot of observable state.** A rendering of the interface, a dump of the relevant process, or a serialisation of the responses — whichever matches what the case was driving. 4. **The identifiers.** Whatever lets the reader find the system's own record of the same work rather than searching by time. 5. **The circumstances.** The clock and time zone, the target the case was pointed at, the worker that ran it, and how far into the run it happened. ## Where the capture is wired, and why that is the design decision Capture belongs to the **harness**, not to the case. A harness lifecycle step that runs after every case and inspects the outcome sees every failure in the suite. Capture written inside individual cases covers only the cases whose author remembered to add it — and the cases that fail undiagnosably are disproportionately the ones nobody has looked at recently. Making capture something every case inherits also makes it uniform: every failure produces the same bundle in the same shape, so a reader does not learn a new format per case. Two constraints on that code, and both are learned the hard way: - **It must not run before the outcome is known.** Gathering evidence for a case that is about to pass is pure cost, paid on every green case of every run. - **It must be defensive.** Capture runs when the system is already unhealthy: the target may be unresponsive, the interface half-drawn, the process out of memory. Each capture step needs its own error boundary and its own deadline, so a slow or failing capture degrades the bundle instead of hanging the worker — and never replaces the case's real failure with a diagnostic error of its own. A third, subtler constraint: capture should read from what the case already accumulated rather than driving the application again. Asking the product for more information after a timing failure can perturb exactly the timing you are trying to explain. ## The test of whether it works The criterion is behavioural, not architectural. Can somebody who was not present, who does not have access to that environment, and who did not write the case, say what went wrong from the bundle alone? If the honest answer is "I would have to run it", the capture is incomplete however many files it produced. The same criterion guards the opposite failure. A hundred files nobody opens is not diagnosability; it is noise with a storage bill. The goal is not maximum evidence — it is enough evidence, arranged so the first thing the reader sees is the thing that broke.

  • What should the harness do when the evidence capture itself throws?
    Keep going, and record the gap. Capture code runs when the system is already unhealthy, so it must never replace the original failure with its own error. Wrap each capture step separately, keep whatever succeeded, and attach a short note naming the capture that could not be taken — so nobody reads a missing snapshot as a missing failure.
  • How do you keep the capture from changing the outcome it is recording?
    Run it only after the outcome is decided, and read from what the case already accumulated rather than driving the application again. Asking the product for more information after a timing failure can perturb the very timing you are explaining. Give each capture a hard deadline so a hung system produces one failure, not a stalled run.
  • What do you capture for a case that fails by timing out rather than by asserting?
    The same bundle plus the three things a timeout lacks: the condition that was being waited on, the last value actually observed, and the elapsed time against the deadline. Without those, a timeout reports only that something did not happen, which is the least diagnosable failure a suite can produce.

Photographing a crash site: the skid marks and the positions of the vehicles exist for minutes, not for the length of the investigation. If nobody records them before the road is cleared, no amount of driving that road again brings them back.

saying these in an interview costs you the question

  • Says the failure can just be reproduced locally to see what happened
  • Adds capture calls inside individual cases instead of the harness
  • Believes a re-run of the failed case reproduces the same conditions
  • Captures nothing beyond the bare failure text and calls it enough
  • Lets an error in the capture code replace the original failure
  • Gathers evidence after teardown has already released the product state