skip to content

Why is re-running a test case until the build turns green not a repair when the product's answer comes from a generative step?

level: juniorimportance: should knowfreq 55%

answer

  1. another draw, not a diagnosis
  2. the rule quietly becomes at-least-once
  3. you stopped on the result you liked
  4. a red under the rule is a finding
  5. rate moved, criteria wrong, or input too hard

basics

~20 s

Re-running draws another sample from the same varying feature; it changes the reported result, not the feature. Under a pass rule that already tolerates the expected variation, a repeated red is evidence the success rate is genuinely below the bar.

solid answer

~40 s

Variation here belongs to the feature, so another execution is another draw rather than a diagnosis. Re-running until green quietly rewrites the rule into *at least one attempt succeeded*, which is the weakest claim a suite can make - a feature succeeding half its attempts clears that almost every time. It also hides the direction of travel: the rate can fall for weeks while somebody keeps pressing the button. Fix the repeat count and the required passes before the runs, and treat a red under that rule as a finding. A repeated red says one of three things: the feature's success rate has moved, the criteria demand something it never reliably did, or this input is beyond it. Each is a decision someone has to make, not noise to be re-sampled away.

code

pseudocode · 10 lines
pseudocode
# anti-pattern: the stopping condition is the result
while build.is_red:
    rerun(case)               # each attempt draws a fresh sample
    # the stated "4 of 5" has silently become
    # "at least one of however many attempts cleared it"

# instead: the count is decided first and every attempt counts
result = run_with_repeats(case, repeats = 5, required = 4)
if not result.passed:
    open_finding(case, met = result.met, of = 5, failed = result.misses)

go deeper

for a junior

Recall that pressing re-run draws a fresh sample rather than changing anything about the feature, and that the result you keep is simply the one you decided to stop on.

for a middle

Explain how unbounded re-runs collapse a stated rule into at-least-one-success, and what an automatic pipeline retry does to a rule that was already sized for the expected variation.

for a senior

Show what you do with a repeated red instead: separate a success rate that has moved from criteria that were always too strict and from an input the feature cannot handle, then act on the one you found.

for a principal

Own the policy: whether re-runs are permitted at all on cases governed by a rate rule, who may change a required-passes number, and how that change is recorded so a later reader sees the bar move.

## A re-run is another draw When a product feature's answer is produced by a generative step, two executions of the same test case on unchanged code can disagree. That is the feature's nature, and it is what makes the re-run button feel reasonable: the build is red, nothing was changed, pressing the button turns it green, the build moves on. What actually happened is that a second sample was drawn from the same distribution and the first was discarded. The feature is unchanged. Its success rate on that input is unchanged. The only thing that changed is which draw you decided to report - and you decided by looking at the result, which is the definition of selecting your evidence. ## What unbounded re-running does to the rule Suppose the case runs five repeats and requires four to meet the criteria. That rule was chosen to tolerate the feature's ordinary variation and to go red when the rate slips. Now allow the build to be re-run until it is green. The effective rule is no longer *four of five*; it is *at least one of these attempts at the rule succeeded*, and the number of attempts is however many times somebody was willing to press the button. - A four-in-five rule is cleared about 92% of the time by a feature that truly succeeds 9 attempts in 10, and about 74% of the time by one that has fallen to 8 in 10. - Allow two re-runs and that same 8-in-10 feature clears the build about 98% of the time; allow four and it is effectively never red. - The signal the rule existed to produce has been spent, and nobody made a decision to spend it. Two related mechanisms do the same damage more quietly. An **automatic retry** in the pipeline stacks a hidden rule on the stated one: the case passes if any retried attempt clears it. And **adjusting the required passes after seeing the result** - dropping from four to three because three is what happened - converts the rule from a measurement into a description. ## What a repeated red is evidence of Once the rule already allows the expected variation, reds that persist across independent builds are not noise. Three explanations account for nearly all of them, and each has a different owner. | Explanation | What it looks like | What it needs | | --- | --- | --- | | the success rate moved | the same inputs meet the criteria in a smaller fraction of repeats than they used to | find what changed around the feature - supplied data, the wording it is given, provider behaviour - and treat it as a product change | | the criteria demand what was never reliably delivered | the case has sat at the edge since it was written, missing on the same criterion each time | the criteria are the defect: assert what the feature actually promises, or narrow the case | | the input is beyond the feature | misses concentrate on one input while sibling inputs stay clean | a deliberate decision: fix the product path, withdraw the promise, or record the input as unsupported | Note what is not on that list: *it was just unlucky*. A rule sized for the ordinary variation is supposed to absorb bad luck, so when it goes red repeatedly, luck is the least likely explanation. Note also that a re-run cannot distinguish any of the three - from the button they look identical. ## Repetition that is legitimate Running the case again is not forbidden. Running it again **in order to obtain a different answer** is. The difference is decided in advance: 1. **Every attempt counts.** Decide the repeat count before the runs and fold all of the attempts into the reported figure, including the ones you did not like. 2. **The stopping condition is not the result.** You stop after the planned attempts, not at the moment the outcome turns green. 3. **The extra attempts have a question behind them.** Tightening an estimate, reproducing under controlled conditions, or checking whether misses cluster on one input are all reasons. "Make the build green" is not. The rule of thumb that survives an interview: a red on a case governed by a rate rule is a **finding**, exactly as a red on an exactly-asserted case is. It gets read, attributed and decided upon. What it never gets is another press of the button.

  • A pipeline automatically retries a failing case twice before reporting. What does that do to a four-in-five rule?
    It stacks a hidden rule on top of the stated one: the case now passes if any of three attempts at the five-repeat rule succeeds, so the effective bar is far below four in five and nobody chose it. Either switch the automatic retry off for cases governed by a rate rule, or fold the retried attempts into the repeat count so that every attempt is counted exactly once.
  • When is running the case again legitimate?
    When you are gathering evidence rather than shopping for an outcome. Running a larger, pre-decided number of repeats to tighten the estimate, or reproducing under controlled conditions to see whether misses cluster on one input, are both reasonable. The difference is that the count is fixed in advance and every attempt counts, instead of stopping the moment the result turns green.

saying these in an interview costs you the question

  • Presses re-run until green and calls the case fixed
  • Treats every red on a varying feature as noise
  • Adds automatic retries so the rate never surfaces
  • Lowers the required passes after seeing which repeats failed
  • Assumes a red always means the product itself is broken