skip to content

Why should an automated suite record a pass on a second attempt as a distinct outcome rather than a plain pass?

level: seniorimportance: should knowfreq 46%

answer

  1. Green can mean two different things
  2. A second-attempt pass is evidence
  3. The merge deletes the first failure
  4. A count alone leaves it unfixable
  5. Ask which cases were red first

basics

~20 s

A pass that needed a second attempt is evidence that a case is unstable. Merging it into a plain pass destroys the only proof that instability exists, so the suite keeps reporting green while its reliability quietly decays.

solid answer

~50 s

The second attempt is the only artefact proving the case failed at all. Once it is labelled a plain pass, the run is indistinguishable from one where everything passed first time, and four things go with it: **which** cases are unstable, by name; the first attempt's failure evidence, which is the only description of the fault; the re-execution time nobody can now attribute; and the ability to tell a genuine intermittent product defect from a badly written case, because the failure that deserved triage was silently absorbed. The bar is that a run must still answer *which cases would have been red without retrying, and what did their first attempts say?* Recording it is not enough either — if the answer lives only in a log nobody opens, a run with four retried passes still renders identically to a clean one.

code

pseudocode · 10 lines
pseudocode
# the summary line a merged record can never print
retried = [c for c in cases if c.attempts > 1 and c.final == PASS]

print("verdict:", "green" if no_failures else "red")
print("passed only after re-execution:", count(retried))

for c in retried:
    print(c.name)
    print("   attempt 1:", c.attempt(1).failure)   # kept, not overwritten
    print("   attempt 2:", c.attempt(2).status)

go deeper

for a junior

Know that a failed case can be executed again, and that passing on the second attempt is not the same result as passing on the first. Be able to find where a run reports that difference.

for a middle

Explain what the record has to carry for a retried pass — the attempt count, and the first attempt's failure evidence kept rather than overwritten — and why a count on its own leaves the case unfixable.

for a senior

Say what a team loses in practice once green covers both outcomes: unnamed unstable cases, unattributed run time, and a real intermittent defect absorbed as noise. Bring a concrete example you lived through.

for a principal

Own the argument that a run's verdict is a product with consumers, and that merging two outcomes into one label is a decision to withhold information from them. State what you would make visible by default and who gets to change it.

A run that ends green can mean two very different things: every case passed the first time it was executed, or some cases failed and then passed when executed again. Those are different states of the product and different states of the suite. A results record that spells them the same way has thrown away the only evidence that the second one ever happened. ## Two events wearing one label | What the run reports | What actually happened | What a reader can act on | |---|---|---| | Green, no further detail | every case passed on its first attempt | nothing needed | | Green, no further detail | four cases passed only on a second attempt | nothing — the evidence was discarded | | Green, four retried passes named | the same four cases | the case names, and each first attempt's failure | The first two rows are indistinguishable to everyone downstream. That is the defect: the suite's own reliability has changed and its output has not. ## What is lost when the two are merged - **The identity of the unstable cases.** Nobody can name them, so nobody can fix them. Instability that cannot be named is instability that will still be there next quarter. - **The first attempt's failure evidence.** A count says a case is unstable; the failed attempt's message and captured output are the only things that say *why*. Once the passing attempt overwrites the record, that description is unrecoverable. - **Attribution of the cost.** Retried cases inflate the run's duration. When the retries are invisible, the suite is simply "getting slower" and no one can point at what is paying for it. - **The difference between an unstable case and a real intermittent defect.** A genuine product fault that appears one run in ten is absorbed into the same silent retry as a badly written wait, and the one failure that deserved triage is the one that is discarded. - **The input other people's suite-health reporting depends on.** Whatever a team publishes about its own pipelines is downstream of this record. If the run never writes it, there is nothing to publish and nothing to trend. - **The meaning of green.** Once a team half-knows that green sometimes required a second try, green stops being a decision input. It is read as *probably fine*, which is the same as not read at all. ## What must survive a successful retry The bar is low and teams still miss it. After a case fails and then passes, the run must still be able to answer: 1. **Which cases would have been red without any retrying?** By name, not as a number. 2. **What did each of those first attempts actually report?** The failure evidence, attached to the case it belongs to, not summarised into a tally. 3. **How much of the run's time went into re-execution?** Per attempt, so the cost is attributable to the cases that caused it. Note the shape of these: they are all questions about the failed attempt, and the failed attempt is precisely the thing a naive implementation discards, because the last attempt is the one that decided the case's status. ## Recording is not the same as reporting A record that exists only in a verbose log nobody opens has been written and not communicated, and the practical effect is identical to never writing it. The distinction has to appear where people actually look after a run — on the verdict or summary they check before merging or deploying. **A run with four retried passes must not render identically to a clean one.** That is the whole mechanism by which the record changes anyone's behaviour. ## The decay this prevents A suite that absorbs retries silently drifts in one direction only. Cases that pass sometimes accumulate, one at a time, none of them ever crossing a threshold, because the threshold is per case and each case is individually within it. The retry policy stays constant while the reliability underneath it falls. Eventually a case exhausts its attempts and reports red, and the team treats it as a new problem — when in fact it has been failing intermittently for months and the suite has been quietly paying for it the whole time. Recording a retried pass as its own outcome is what converts that slow drift into something visible early, while it is still one or two cases and still cheap to fix. It costs a field and a line in a summary. Merging it costs the ability to tell whether the suite is getting better or worse at all.

  • A run reports twelve retried passes. What is the first thing you ask of that number?
    Which cases, and what did each first attempt report? The count establishes that instability exists; the case names and their failure evidence are what make it actionable. A number on its own is a mood, and it will still be twelve next quarter because nobody could act on it.
  • Where does the distinction have to appear to actually change anyone's behaviour?
    On the surface people read after a run — the verdict or summary checked before merging or releasing — not only in verbose output. A retried-pass record that exists but is never rendered has the same practical effect as no record: the run still looks identical to a clean one, so nobody investigates.

A run that merges the two outcomes is a scoreboard reporting the final score and erasing every foul. The result is accurate and the game is unreconstructable.

saying these in an interview costs you the question

  • Says green is green, however many attempts it took
  • Overwrites the case result with the winning attempt
  • Counts retries but discards the first failure's evidence
  • Treats a retried pass as proof the product is fine
  • Buries the record where nobody reads it