skip to content

When is a retry inside an automated test a legitimate model of the system's contract, and when is it masking a race?

level: principalimportance: should knowfreq 44%

answer

  1. Three mechanisms share one word
  2. Ask whether a real consumer retries
  3. Model a contract, never establish one
  4. Retrying obliges you to assert the guarantee
  5. An undeclared retry is discarded evidence

basics

~20 s

A retry is legitimate when it mirrors a contract the system really offers — at-least-once delivery, bounded eventual consistency, a published idempotent operation — and the case asserts that guarantee. Otherwise the retry is discarding evidence of a race.

solid answer

~50 s

First separate the mechanisms: a poll re-checking a condition is not a retry, re-issuing the action is, and re-running the whole case is a third thing. For the latter two, the test is: **does a real consumer retry here?** If the contract is at-least-once, bounded eventual consistency, or a published idempotent operation, a retry models reality — but only if the case then asserts the guarantee, such as one record after two submissions or convergence inside the stated window. Without that assertion, a retry silently converts a correctness property into an availability one. It is masking a race when the first attempt failed nondeterministically, when success depends on the second attempt landing late, or when the operation is specified to happen once. As a lead I require every retry to cite the contract clause it models, publish retry rate as a trended metric, and open an undeclared retry as a defect.

go deeper

for a junior

Know that re-running something until it passes is not the same as waiting for it to finish, and that a case which needed a second attempt has told you something you should not ignore.

for a middle

Explain the difference between a poll re-checking a condition and re-issuing the action, and why re-issuing puts the system's behaviour under duplication into the scope of the test.

for a senior

Show that you can tell a contract-modelling retry from a race-masking one using the system's own guarantees, and that when you keep a retry you add the assertion — uniqueness, idempotence, bounded convergence — that proves the guarantee holds.

for a principal

Own the policy: retries declared with the contract clause they model, the matching assertion mandatory, attempts published as a trended metric, and a standing rule that a system's correctness may never come to depend on tests retrying.

### Separate three things people all call "retry" Most of the confusion in this discussion is vocabulary. Three distinct mechanisms wear the same word, and only one of them is contentious. 1. **A poll re-checking its condition.** The wait evaluates a predicate repeatedly until it holds. Nothing is re-executed; the test observes the same single action several times. This is not a retry in any meaningful sense, and treating it as one is what leads people to say "all waiting is retrying". 2. **Re-issuing the action.** The test submits again, re-triggers the job, or re-sends the message after nothing appeared. Now a second attempt exists in the system, and the system's behaviour under a second attempt is squarely in scope. 3. **Re-running the whole case.** The case failed, so it is executed again from the start and reported on its second outcome. Levels 2 and 3 are where the judgement lives, and the question to ask of each is the same: **does a real consumer of this system do that?** ### The legitimacy test: does the retry model the contract? A retry inside a test is legitimate when it **mirrors a contract the system actually offers**, and the test is honest about it. Real contracts that justify a retry: - **At-least-once delivery.** Consumers are expected to re-send when they do not see an acknowledgement, so a test that never retries is testing a system nobody runs. - **Documented eventual consistency with a bounded convergence window.** The contract is "you may read stale for up to N"; a read that retries within N is exercising the contract, and a read that retries beyond it is hiding a violation. - **A published idempotent operation.** The API states that re-submission with the same key is safe. A retrying test is exercising the guarantee — provided it *asserts* the guarantee, not merely benefits from it. That proviso is the whole discipline. When the retry models a contract, the case must assert the invariant the contract promises: exactly one record after two submissions, a stable identifier, no duplicated side effect, convergence inside the stated window. A retry without that assertion is indistinguishable from a retry that is hiding something. ### When the retry is a defect A retry is masking a race whenever the system does **not** promise that a second attempt is meaningful, and the second attempt succeeds anyway. The tells: - The first attempt failed **nondeterministically** and nobody can say which interleaving produced it. - The retry's success rate depends on timing — it works when the second attempt lands late, and occasionally fails when it lands early. - The contract says the operation happens once, and the test needed it to happen twice. In all three, a real user hitting the same interleaving gets the failure the test just discarded. The test found a genuine product bug and then deleted the evidence. That is why the strong position is not "retries are bad" but "**an undeclared retry is a bug report you threw away**". ### A worked example On a smart-meter reading-feed suite, a case submitted a batch of 8,412 readings and waited for the per-meter daily total. When the wait timed out, a helper re-submitted the batch. The feed is genuinely at-least-once, so a resubmit looks defensible — and the case passed. What it hid was an ordering assumption in the deduplication path: duplicates were suppressed only when the second copy arrived after the first had committed. Landing early, both copies were aggregated, and the daily total was overstated. Nobody saw it, because the retry rate was never published: it drifted from roughly one run in forty-three to one in nine over a quarter while the board stayed green. The case had been asserting that a total eventually appeared, but never that submitting twice produced one set of readings. ### What a lead actually puts in place - **Retries are declared, never incidental.** Each retry in the suite cites the contract clause it models. A retry with no citation is opened as a defect, not tuned. - **The contract assertion is mandatory.** If a case retries, it asserts idempotence, uniqueness or convergence — otherwise the retry is silently converting a correctness property into an availability one. - **Retry rate is a published metric, not a hidden implementation detail.** Attempts per case, per run, trended over time. A rising rate is a production-facing signal about the system long before it is a suite problem. - **Bounded and attributed.** One extra attempt, inside the documented window, with the reason recorded in the run output — so a reader can see that a retry occurred and why it was permitted. - **A standing rule about direction.** Retries may model a contract; they may never *establish* one. The moment a system's correctness depends on the test retrying, the contract has moved, and that is an engineering decision made by people, in writing. ### The tradeoff to own out loud Forbidding retries entirely makes suites brittle against systems whose real contracts are probabilistic, and pushes teams into padding waits instead — the same masking, in a less visible form. Permitting them freely converts every race into an availability question that never reaches anyone. The defensible middle is narrow and administrative rather than technical: retries allowed where a contract says so, always paired with the assertion that proves the contract, always counted.

  • A team argues that all waiting is retrying, so the distinction is meaningless. How do you respond?
    A poll re-evaluates a predicate over a single action that has already been issued; nothing new enters the system, and the observation is idempotent by construction. Re-issuing the action puts a second attempt into the system and makes its behaviour under duplication part of what is being tested. Collapsing the two hides exactly the case that needs judgement, which is why the vocabulary is worth defending in a design review.
  • What single metric would you publish to keep retries honest across many teams?
    Attempts per case, trended per run. It is cheap for a shared helper to emit and it surfaces the drift that hides inside a green board — a case quietly moving from succeeding first time to needing a second attempt every few runs. Trended alongside wait durations it also distinguishes a system getting slower from a system getting less deterministic, which are different conversations with different owners.
  • Is there a case for allowing a retry around genuinely external nondeterminism, such as an unreliable shared dependency?
    Yes, provided it is scoped and recorded. If production tolerates that dependency's failures by retrying, the test is modelling reality and should retry the same way and assert the same recovery. What makes it safe is the boundary: the retry wraps only the interaction with that dependency, not the behaviour under test, and the run output shows it fired. An unscoped retry around the whole case cannot distinguish the dependency's failure from the product's.

saying these in an interview costs you the question

  • Treats every retry as harmless because the run went green
  • Retries an operation specified to succeed once
  • Retries without asserting the idempotence it relies on
  • Never counts or publishes how often retries fire
  • Calls all polling a retry to avoid the distinction
  • Lets a system's correctness depend on tests retrying

context