When an automated test fails in a pipeline run, what should its failure message and captured artefacts tell you before you re-run it?
answer
- Could you diagnose it without re-running?
- Expected versus actual, not a boolean
- Name the case and the row
- Capture artefacts before teardown runs
- Record locale, build, seed, image
basics
~20 sA failure should name the case, state expected versus actual in real values, and point at the failing step, with artefacts captured at that moment - logs, a screen or payload capture, the environment - so nobody has to re-run it to learn what happened.
solid answer
~50 sA failing case should be diagnosable from the report alone, which means four things. The report names the case and, for a data-driven case, the input row that failed, so you do not have to open the source to know what ran. The assertion states expected and actual as domain values rather than a bare boolean; a check that can only report `expected true but was false` throws away the evidence at the exact moment you need it. Artefacts are captured at the moment of failure, before teardown restores or destroys state: the relevant logs, a screen or payload capture, the request and response, and an identifier that ties the case to the system's own logs. And the run records its conditions - build identifier, runner image, time zone, locale, any random seed - because those are often the cause. Capture on failure, keep retention longer than the triage cycle, and keep the passing path cheap.
code
pseudocode · 15 lines# weak: the report can only say "expected true but was false"
assert isEqual(field.renderedText, expectedText)
# better: the failure names the case, the field and both values
assert field.renderedText == expectedText,
message = "visit-date for subject " + subjectId +
": expected '" + expectedText +
"' but rendered '" + field.renderedText +
"' (runner locale=" + runEnvironment.locale + ")"
# capture evidence on the failure hook, before teardown runs
on_case_failed(case):
attach(case, "screen", capture_screen())
attach(case, "server-log", fetch_log(correlationId = case.correlationId))
attach(case, "conditions", runEnvironment.asRecord())go deeper
Be ready to say what a good failure message contains: the case, the input row, expected and actual as real values, and the failing step. Interviewers ask this to see whether you have ever had to triage someone else's red run.
Explain the mechanics - where a runner lets you hook the failure outcome, why capture must happen before teardown, and how to keep a shared assertion helper from erasing the caller's context in the message.
Show the operating judgement: what to capture on failure versus on pass, retention windows tied to the triage cycle, artefact size caps, and correlation identifiers that link a case to the system's own logs.
Own the economics. Argue that triage cost per red run, multiplied by red-run frequency, is usually the largest running cost of a suite, and set the expectations and defaults that keep it bounded as the suite grows.
## Why this belongs to suite maintenance A test report is a message written to a stranger - usually to yourself, three weeks later, at the top of a morning triage queue. The quality measure is blunt: can someone who did not write the case decide, from the report alone, whether the product is broken, the case is wrong, or the environment moved? If the honest answer is "re-run it and watch the screen", the failure is not diagnosable, and every future failure of that case costs the same again. That recurring cost is why diagnosability is a maintenance concern rather than a nicety of case authoring. On a suite of a few thousand cases, the dominant running expense is rarely machine time; it is human triage time multiplied by the number of red runs. A suite whose failures each cost twenty minutes to interpret is a suite people will eventually stop reading, and a suite nobody reads has already been deleted - just without anyone deciding to. ## The four things a report must carry **1. Identity.** The report names the case in words that describe the behaviour, and for a parameterised case it names the specific row: not `case[7]` but the input values that produced the failure. If the case name is a method name and the data comes from a table, the report must render the row, or triage begins with reverse-engineering which row index 7 was. **2. Expected versus actual, in domain terms.** This is the single most common defect in a young suite. A boolean assertion collapses the comparison before the framework can describe it, so the report says only that something was false. Assert on the values themselves, or pass the values into the message. The rule of thumb: if the message would not let you write the fix without opening a debugger, it is not carrying enough. **3. The failing step.** In a multi-step journey, the report should say which step broke and what the preceding steps established. A journey that fails at its final assertion with no trace of the eight actions before it forces you to reconstruct the path. **4. The conditions.** Build identifier, the runner image or agent, the time zone, the locale, the data set version, any random seed. These belong in the run record, not in each case, and they are the facts most often responsible when a case passes on a workstation and fails nightly. ## Artefacts and their timing Artefacts are the evidence you cannot reconstruct: application logs for the window of the failure, a screen or rendered-payload capture, the request and response bodies, a dump of the record the case created. Two rules govern them. First, **capture at the moment of failure, before teardown**. Teardown exists to restore a clean state, which means it destroys exactly the evidence you want. Most runners provide a hook that observes the outcome of a case; that hook, not the end of the run, is where capture belongs. Second, **make the artefact findable and correlatable**. Name it after the case and the run, publish it with the report rather than leaving it on a worker that is about to be recycled, and include a correlation identifier that the system under test also logs, so a tester can jump from the failing case straight to the server-side trace of that one request. ## A worked failure A nightly suite over a clinical-trial data capture form had a case checking the visit-date field. One morning it failed with `expected true but was false`. Re-running it on a workstation passed, so it was re-run in the pipeline, where it failed again - each attempt costing a full 6-hour nightly cycle. It took three days to learn that the runner image had changed its default locale, that the field now rendered the date in a locale-dependent format, and that the case was comparing the rendered string against text built for a different locale. A report that had printed `expected '2026-03-14' but rendered '14/03/2026'` and recorded the runner locale in the run header would have made it a two-minute read. Nothing about the product changed; only the report's poverty was expensive. ## Keeping the cost down Diagnosability is not "capture everything". Full logs and a screen capture on every passing case inflate run time, fill storage, and bury the one artefact that matters. Capture on failure; if you want a baseline for comparison, sample the passing path rather than recording all of it. Set a retention window slightly longer than your triage cycle so a Friday failure is still explorable on Monday, and cap artefact size so a runaway log dump cannot take the pipeline down with it. The test to apply to any suite: pick a red run from last month at random and try to explain it using only what the report kept. Whatever you had to go and re-derive is the gap.
- A team captures a screen image and a full log dump for every case, passing or failing. What goes wrong?Run time and storage grow with the suite, the useful artefact is buried among thousands of identical ones, and retention limits are hit so quickly that evidence expires before anyone triages it. Capture on failure, sample the passing path if you want a baseline, cap artefact size, and set retention slightly longer than the triage cycle.
- How do you keep the failure message useful when the assertion lives in a shared helper used by hundreds of cases?Pass the context into the helper - the field, the record identifier, the row - and render it in the message, or have the helper compare values rather than return a boolean. A helper that asserts a bare condition erases the caller's context: the report then names the helper's line, which is identical for every case that uses it.
- What does a correlation identifier add to a failure that logs alone do not?It lets you move from the case to the system's own record of exactly that request, instead of guessing which of many concurrent entries in a shared log belongs to your run. On a sharded or parallel suite this is the difference between a targeted lookup and reading someone else's traffic.
A failure report is an incident photograph. Taken at the moment of the crash it settles the question; taken after the road has been swept it proves only that something happened here.
saying these in an interview costs you the question
- Says the diagnosis step is to re-run it and watch
- Asserts a bare boolean, so the report only says false
- Captures artefacts after teardown has restored the state
- Assumes pipeline logs and artefacts will still exist next week
- Thinks a stack trace alone identifies which data row failed
- Records nothing about locale, build or runner image