skip to content

How many repeats should a test case run against a generative product feature, and what does a four-in-five pass rule claim?

level: middleimportance: must knowfreq 60%

answer

  1. one execution samples, it does not measure
  2. repeats estimate a rate
  3. count follows the drop you must detect
  4. a rate claim, never a user guarantee
  5. fix the count and the bar beforehand

basics

~20 s

Repeat count follows the success-rate difference you need to detect; three to five repeats separate only coarse differences. A four-in-five rule claims the feature's per-attempt success rate is high, not that any single user attempt will succeed.

solid answer

~50 s

A feature whose output comes from a generative step returns different answers to the same input, so one execution of its test case samples one attempt and a single pass says very little. Repeats turn that draw into an estimate of the feature's success rate on that input. Derive the count from the drop you must notice and the cost of an attempt: three to five repeats separate only coarse moves, so spend a larger count where variation matters. State the rule before the runs - *four of five repeats must meet the criteria* - and read it for what it is, a claim about a **rate** rather than a guarantee to a user. A feature that truly succeeds nine attempts in ten clears that rule about 92% of the time; one that has slipped to eight in ten still clears it about 74%.

code

pseudocode · 14 lines
pseudocode
REPEATS  = 5
REQUIRED = 4          # both fixed before the runs, part of the case definition

met    = 0
misses = []
for i in 1..REPEATS:
    output = feature.answer(input)        # the generative step runs inside the product
    if meets_criteria(output):
        met = met + 1
    else:
        misses.append({ repeat: i, failed: first_unmet_criterion(output) })

record(caseId, repeats = REPEATS, met = met, rule = "4 of 5", misses = misses)
result = (met >= REQUIRED) ? "pass-under-rule" : "fail"

go deeper

for a junior

Be ready to say why running the case once tells you little when the answer comes from a generative step: one execution is a single sample, and the next one may differ on identical input.

for a middle

An interviewer expects you to derive the repeat count rather than guess it - the drop in success rate you must notice, the cost of an attempt, and why three to five repeats separate only coarse moves.

for a senior

Show that the count and the required passes are fixed before the runs, and say out loud what a passing rate rule does and does not license about the next user attempt on that input.

for a principal

Own where repeats are spent: which inputs deserve a large count, what success rate the team is willing to ship, and how that budget is defended when the cost of an attempt rises.

## Why a single execution decides nothing A product feature whose output is produced by a generative step does not return the same answer twice. Ask it the same question on Monday and on Tuesday, with the same data behind it and nothing deployed in between, and the wording changes, the ordering changes, and sometimes the substance changes. That variation is a **property of the feature**, as much as its latency or its cost is. It is not a defect in the test, and the person writing the test will not engineer it away. The consequence for a test case is uncomfortable: one execution of the case samples **one attempt** out of a distribution of possible attempts. If the feature meets the criteria on four attempts in five, then a suite that runs the case once reports green about 80% of the time and red about 20% of the time - on identical code, identical data and identical input. Neither result is a lie and neither is informative. A green says this particular draw was fine; a red says this particular draw was not. Repeats are the way out. Running the case `n` times and counting how many attempts met the criteria turns a single draw into an **estimate of a rate**: how often the feature does what you promised on that input. ## Choosing the repeat count The count is derived, not chosen by taste. Three things decide it: - **The success rate you are prepared to ship.** You cannot judge a rate you have not named. A criterion such as "the summary must contain the deadline the customer stated" at 95% is a different feature from the same criterion at 70%. - **The drop you need to notice.** Detecting a fall from 90% to 50% takes very few attempts. Detecting 90% to 85% takes a great many. Precision improves only with the square root of the count, so the tenth repeat buys far less than the second. - **The cost of an attempt.** The generative step is normally the slowest and most expensive part of the case; five repeats make the case five times as expensive in money and in wall-clock time. The practical result is that a small count - three to five - is all most cases can afford, and a small count separates only coarse differences. So spend the repeats where the variation matters: a larger count on a handful of representative inputs, a small fixed count elsewhere, and no repeats at all on the parts of the feature that never pass through the generative step. ## What a rule over the repeats actually claims Fix the rule **before** the runs - the repeat count `n`, and the number of attempts `k` that must meet the criteria - then read it literally. A `k`-of-`n` rule is a statement about a **rate**, and about nothing else. | Rule over 5 repeats | What it claims | What it does not claim | | --- | --- | --- | | all 5 meet the criteria | the per-attempt success rate is high enough that a clean sweep is ordinary | that the sweep repeats; a feature succeeding 9 attempts in 10 misses this rule about 41% of the time | | at least 4 of 5 | one miss in five is within tolerance on this input | that the next user attempt succeeds; it may be the miss | | at least 1 of 5 | almost nothing | anything useful; a feature succeeding half its attempts clears it about 97% of the time | Three consequences follow, and an interviewer is usually listening for them. 1. **It is not a per-user guarantee.** A case passing at four in five means about one user attempt in five on that input is expected to miss. If that is unacceptable, the rule is not the thing to change - the feature, the criteria, or the product behaviour around the generative step is. 2. **The estimate is wide.** A feature that truly succeeds nine attempts in ten clears a four-in-five rule about 92% of the time; one that has slipped to eight in ten still clears it about 74% of the time. One green under that rule barely separates the two, which is why the per-repeat counts are worth keeping and reading against previous builds rather than reading a single build alone. 3. **It certifies only the criteria you wrote.** The rule counts how often the criteria were met. If the criteria assert only that the answer is non-empty and mentions the order number, then a rate of five in five certifies a low bar precisely. Adding repeats to a weak criterion is a category error; strengthen the criterion instead. Finally, treat `n` and `k` as part of the case's definition rather than as a dial. Raising `k` because a build went green, or lowering it because a build went red, converts the rule from a measurement into a description of whatever just happened.

  • Your four-in-five rule passes a build in which every repeat produced a different answer. What has it not checked?
    The rule counts only how often the criteria were met, so it says nothing about whether the criteria assert what you actually promised. If they check for a non-empty answer and a required field, five structurally valid but substantively different answers all count as passes. Repeats measure the rate at which a bar is cleared; raising the repeat count cannot raise the bar.
  • Why not simply raise the repeat count until the estimate is tight?
    Cost and wall-clock time. The generative step is usually the slowest and most expensive part of the case, and repeats multiply it directly. Precision improves only with the square root of the count, so the tenth repeat buys much less than the second. Concentrate a larger count on the few inputs whose variation matters and keep a small fixed count everywhere else.
  • The team wants the repeat count to be the same for every case in the suite. What is wrong with that?
    It spends the budget where nothing varies and starves the places that do. Parts of the feature that never pass through the generative step need no repeats at all, and a handful of inputs where the answer is marginal need many more than the house number. A uniform count is a way of not deciding which behaviour the team actually cares about.

saying these in an interview costs you the question

  • Treats one green execution as proof the feature works
  • Picks the repeat count by habit rather than the difference it must detect
  • Reads a four-in-five pass as a promise to every user
  • Raises or lowers the required passes after seeing the result
  • Thinks more repeats can compensate for criteria that assert the wrong thing