A feature retries its generative step once before falling back. What should the test case assert about that retry?
answer
- A second attempt is not free
- Two waits, two charges, one person
- Bound attempts, not just outcomes
- Some failures must never be attempted twice
- Count attempts where someone reads them
basics
~20 sAssert the retry's price, not only its success: the total wait the person can now face, a hard ceiling on attempts, that only transient failures are retried, and that the extra calls are counted somewhere people read.
solid answer
~40 sA retry is a second full wait and a second charge, so the case has to pin both. **Wait**: if each attempt is bounded at four seconds, assert an outcome appears inside the feature's overall budget — roughly eight seconds plus the pause between attempts — not merely that the second attempt eventually succeeded. **Attempts**: assert a hard ceiling, and assert it holds when every attempt fails. **Eligibility**: assert a deterministic rejection such as oversized material is attempted exactly once, while a refusal for volume is attempted again after the suggested pause. **Visibility**: assert the attempt count is recorded, so a doubling of calls during a bad hour shows up as a number rather than as a surprise on a bill.
code
pseudocode · 16 linesPOLICY at the feature edge:
attempt_budget_ms = 4000
max_attempts = 2
retry_when = refused_for_volume OR no_answer_in_time
never_retry_when = input_too_large
pause_before_retry = suggested_wait OR 500 ms
case worst_case_wait:
given every attempt fails
assert fallback shown within 2 * attempt_budget_ms + pause
assert attempts_made == max_attempts
assert attempts_made recorded for this interaction
case not_retryable:
given input_too_large
assert attempts_made == 1go deeper
Know that a second attempt means the person may wait twice as long and the product pays twice. At this level, be able to say that a retry is a decision with a price rather than an automatic kindness.
Explain which failures are worth a second attempt and which are not, and how a pause between attempts differs from trying again immediately. Be ready to work out the worst-case wait implied by a given attempt budget and ceiling.
Show the assertions: the whole wait bounded when everything fails, attempts capped, non-retryable failures attempted once, attempt counts recorded. Talk about how the policy behaves in aggregate while the dependency is degraded, not only in one interaction.
Own the tradeoff between a better success rate and doubled spend plus doubled latency at the worst moment. Deciding how many attempts a feature buys, and whether that budget is better spent on a queued promise, is a policy call with real money attached.
A retry looks like a free improvement: the same call, tried again, sometimes works. It is not free. It spends the person's patience a second time and the product's money a second time, and both of those are properties a test case can pin down rather than leave to whatever the implementation happens to do. ## The two prices of a second attempt **Wait.** If each attempt is bounded at four seconds and one retry is allowed, the worst case someone can experience is two full attempts plus the pause between them — around nine seconds before anything appears. Each attempt looks reasonable in isolation; the sum is what the person actually lives with. A case that asserts only the per-attempt deadline will happily pass a feature whose real worst case is twice what anyone agreed to. **Money and load.** Every failing interaction now costs two calls instead of one, and it does so exactly when the dependency is least able to absorb them. Retrying immediately after a refusal for volume is the pathological version: the feature answers "too many requests" with one more request, and a busy period becomes a self-inflicted one. A pause between attempts, ideally honouring whatever wait the refusal suggested, is the difference between a retry policy and a stampede. ## What the case asserts 1. **A bound on the whole wait.** Force every attempt to fail and assert a fallback view is on screen inside the feature's overall budget, not inside one attempt's budget. 2. **A ceiling on attempts.** Assert the number of attempts made when everything fails, so an unbounded loop hiding under a deadline cannot survive. 3. **Eligibility.** Assert that a deterministic rejection — material larger than the step accepts — is attempted exactly once, while a refusal for volume is attempted again after a pause. Trying again at something that cannot succeed is pure cost. 4. **A pause that exists.** Assert there is a gap between attempts, and that a suggested wait, where one is offered, is respected rather than ignored. 5. **Visibility.** Assert the attempt count is recorded for the interaction, separately from interactions that succeeded first time. | Failure | Attempt again? | Why | | --- | --- | --- | | Refused for volume | Yes, after a pause | The identical request usually succeeds shortly after | | No answer inside the deadline | Once, carefully | It may have been transient, and may also still be running | | Provider unavailable | Rarely inside one session | The condition normally outlasts the interaction | | Input too large | Never | The outcome is deterministic and will not change | ## The one that is easy to get wrong Trying again after a missed deadline is the subtle case. The first attempt may still be in flight and may still complete somewhere, which means the feature can end up paying for two pieces of work, and waiting through two budgets to show one fallback. That is why the number of attempts, and the conditions under which each is allowed, belong in a stated policy the case reads rather than in a helper nobody opens. ## Making the cost visible The reason to record attempts is not tidiness; an invisible retry is an invisible bill. A count of attempts per interaction, split by outcome, answers three questions a team is eventually asked: how often does the second attempt actually succeed, what did retries cost during the last bad hour, and would the same money have bought more as a queued promise. Without that number the attempt ceiling tends to be set once, by whoever wrote the code, and never revisited — and nobody can tell whether raising it from one to three would help anyone or simply double the spend. The same number is what makes a regression visible. If the share of interactions needing a second attempt climbs from two percent to thirty over a quarter, something upstream has changed: a longer input, a busier allowance, a slower dependency. Without the count, that change shows up first as a support complaint about slowness. ## A workable default to argue from - One retry, not several, unless the result is delivered asynchronously and nobody is watching a screen. - Transient failures only, never a deterministic rejection. - A short pause before the second attempt, honouring any suggested wait. - An overall budget the product owns, shorter than anything the feature depends on. - Attempts counted and reviewable beside spend. That is a starting position rather than a rule. A feature whose result arrives later can afford more attempts; a feature someone is watching can afford fewer. What matters is that the numbers live somewhere a test case can assert them, because a retry policy that exists only in code is a policy nobody has agreed to.
- Why is retrying on every failure worse on exactly the day the provider is struggling?Because it multiplies traffic when the dependency is least able to take it, and every failing interaction costs two calls instead of one. A refusal for volume answered with an immediate second call turns a busy period into a self-inflicted one. The case should assert a pause before the second attempt and a ceiling on attempts, not just for one person but as the feature's stated policy.
- How would you assert that the retry has not quietly eaten the whole patience budget?Fix the feature's overall budget in the case, force every attempt to fail, and assert a fallback view is on screen before that budget expires. Asserting per-attempt deadlines alone is how a two-attempt feature ends up with double the wait nobody agreed to, because each attempt looks perfectly reasonable in isolation.
- What makes a retry visible to the people paying for it?A recorded count of attempts per interaction, separated from interactions that succeeded first time, surfaced where volume and spend are reviewed. Without it the second call is invisible until the bill arrives, and nobody can answer whether the retry bought anything — which is the number that decides whether one attempt becomes zero or three.
saying these in an interview costs you the question
- Retries any failure, including a rejected oversized input
- Counts only the successful attempt's wait
- Retries immediately with no pause between attempts
- Leaves the number of attempts unbounded under a deadline
- Treats the extra calls as too small to measure