skip to content

In an automated suite, what does a retry budget for the whole run protect that a per-case retry cap does not?

level: seniorimportance: nice to knowfreq 20%

answer

  1. Two bounds with different scopes
  2. A cap is per case only
  3. Every case may spend its cap
  4. A budget bounds the whole run
  5. Exhaustion must show on the verdict

basics

~20 s

A per-case cap bounds one case's attempts and nothing bounds how many cases retry. A run-level budget caps total re-execution across the suite, so a suite-wide breakage reports red quickly instead of paying the full cap on every case.

solid answer

~50 s

A per-case cap is a **local** bound: it answers *how many attempts may this case have?* and it is fully satisfied when every case in the suite consumes all of it. Two failures escape it. A systemic breakage — a dependency down, a bad deployment — fails every case, and each one dutifully spends its allowance, so the run costs several times its normal execution to deliver exactly the same red verdict. And creeping instability spreads case by case without crossing any per-case threshold at all. A run budget is a **global** bound — an absolute count of re-executions, a share of cases, or a share of wall-clock, chosen from the team's own history. When it is exhausted, retrying stops, later failures are reported as failures, and the verdict says the budget ran out. Use both: the cap bounds one case's latency, the budget bounds the run.

code

pseudocode · 15 lines
pseudocode
run_budget = 20        # total re-executions allowed this run
case_cap   = 3         # attempts allowed to any single case

for case in suite:
    attempts = 1
    run_case(case)
    while case.failed and attempts < case_cap and run_budget > 0:
        run_budget = run_budget - 1
        attempts = attempts + 1
        run_case(case)
    record(case, attempts)

if run_budget == 0:
    # visible on the verdict, not only in a log nobody opens
    mark_run("retry budget exhausted")

go deeper

for a junior

Know that re-execution is bounded on purpose, and that there are two different bounds: how many attempts one case may have, and how much re-execution a whole run may do.

for a middle

Explain why a per-case cap is satisfied even when every case in the suite consumes it, and what that permits during a suite-wide breakage. Be able to describe a bound whose scope is the run instead.

for a senior

Work a scenario out loud: a suite-wide dependency failure with only a per-case cap, then the same failure with a run budget. Say what exhaustion does to the reported verdict and why that visibility matters.

for a principal

Own the budget as an agreement about how long a red answer may take and how much instability the team absorbs before it becomes scheduled work. Say who sets the number, how it ratchets, and what would make you change it.

Two different bounds are commonly called "the retry limit", and they do different jobs. A **per-case cap** answers *how many attempts may this one case have?* A **per-run budget** answers *how much re-execution may this whole run do in total?* The first is local, the second is global, and neither substitutes for the other. ## Two bounds, two scopes | | Per-case cap | Per-run budget | |---|---|---| | Question it answers | how many attempts for this case | how much retrying for this run | | Scope | one case | the whole suite execution | | Stops | one hopeless case spinning forever | the run paying without limit | | Worst case it still permits | every case using its full allowance | retrying halts once the budget is spent | | What it protects | one case's latency | the run's total cost and the honesty of the verdict | The crucial row is the fourth. **A per-case cap is fully satisfied when every case in the suite consumes all of it.** It contains no notion of the suite at all, so there is no number of failing cases at which it decides something is wrong. ## The failure a cap cannot see Suppose a suite of 800 cases with a cap of three attempts per case, and a run that normally takes twelve minutes. A shared dependency becomes unreachable a minute in. Every case now fails, and every case dutifully spends its full allowance: the run executes up to 2,400 case-attempts instead of 800 and delivers exactly the same red verdict roughly three times later than it needed to. The cap was never violated — it was honoured, 800 separate times. A run budget of, say, twenty-five re-executions turns that into a red verdict inside the first minute or two, with the run explicitly stating that its budget was exhausted. Nothing about the diagnosis changed; what changed is how long everyone waited for it and how much capacity the run consumed producing it. The second failure the cap cannot see is slower. Instability spreading from two cases to sixty crosses no per-case threshold at all — each of those sixty is individually within its allowance. Only a bound whose scope is the run notices that the aggregate moved. ## Shapes a run budget can take The unit is a team decision, not a given. Common shapes: - **An absolute count** of re-executions permitted per run. Simplest to reason about and to explain. - **A share of executed cases** — for example, re-execution may touch at most two percent of the cases in the run. Scales as the suite grows. - **A share of wall-clock** — re-execution may consume at most a stated fraction of the run's normal duration. Ties the bound directly to the feedback the team is buying. - **A deadline for the whole run** that any retrying has to fit inside. The bluntest form, and the one that never needs re-tuning as the suite grows. Every one of those numbers is defined by the team from its own history. There is no universal figure, and adopting one from elsewhere produces a bound that either never fires or fires constantly. ## What exhaustion must do The budget only means something if running out has consequences: 1. **Stop re-executing.** No further attempts, for any case, for the rest of the run. 2. **Report subsequent failures as failures.** They are failures; the budget's exhaustion does not change what happened. 3. **Say on the verdict that the budget was exhausted.** Otherwise the resulting red reads as one ordinary failing case, when it is actually a statement that the run gave up on absorbing instability. This is the step teams skip, and skipping it makes a suite-wide event look like a single broken case. 4. **Never quietly continue.** A bound that is ignored the moment it binds is not a bound; it is a comment. ## Choosing and maintaining the number Set the budget **above what a calm period actually consumes** — otherwise it fires on the status quo and everyone learns to ignore it — and low enough that a suite-wide breakage trips it early rather than after the whole suite has paid. Then ratchet it down as the suite gets steadier, so the bound keeps tracking the reality rather than the history. There is one move that destroys it: raising the budget because it fired. That converts a cost control into decoration, and it is exactly the response that feels reasonable in the moment, because the run is red and the budget looks like the reason. It is not the reason. It is the messenger. Finally, a budget bounds cost; it does not fix anything, and on its own it cannot say **which** cases consumed it. That question belongs to the per-case record. The two bounds are complementary: the cap keeps one case from spinning, the budget keeps the run from spending, and the per-case record is what turns either of them into work someone can actually do.

  • What should a run do at the moment its retry budget is exhausted?
    Stop re-executing, report every later failure as a failure, and state on the verdict itself that the budget ran out — otherwise the red reads as one ordinary failing case rather than a suite-wide event. The one thing it must not do is treat the budget as advice and keep retrying, which converts the bound into a comment.
  • How would you set the number for a run budget?
    Above what a calm period actually consumes, so it fires on something new rather than on the status quo, and low enough that a suite-wide breakage trips it in the first minute or two. Then ratchet it down as the suite gets steadier. Raising it because it fired is the move that turns the budget into decoration.

A per-case cap is a limit on what one person may withdraw; a run budget is the balance of the account. Every withdrawal can be inside the limit while the account still empties.

saying these in an interview costs you the question

  • Thinks a per-case cap bounds the run's total retrying
  • Raises the budget whenever it is exhausted
  • Keeps re-executing after the budget is gone
  • Cannot say what exhaustion does to the verdict
  • Copies a budget number from somewhere else