Where can a retry sit in an automated suite — around an action, around a case, or around a whole run — and what does each placement conceal?
answer
- Ask what unit runs again
- Three widths: action, case, whole run
- Each step outward deletes attribution
- Setup runs again at case level
- Stacked placements multiply, not add
basics
~20 sA retry can wrap a single action, a whole case, or an entire run. The wider the wrapper, the more it conceals: an action-level retry hides that a step was unstable, a run-level retry hides which case failed at all.
solid answer
~40 sThree placements, three blast radii. **Around an action** — a helper re-issues one call or interaction until it succeeds; the case never learns it happened, so that step's instability is invisible in the results. **Around a case** — the runner re-executes the failing case from its setup; the record shows the case needed another attempt but cannot say which step was unstable. **Around a whole run** — everything is launched again, which hides which case failed and re-pays the full cost, so it is only justified when the run went down wholesale rather than on one case's assertion. Narrower is cheaper and more diagnosable, but only where the retried unit is safe to execute twice: case-level retry re-runs setup, and any step with a side effect will perform it again.
code
pseudocode · 21 lines# placement A - inside one action; the case never sees it
function click_when_ready(target):
for attempt in 1..3:
if try_click(target):
return OK
pause(200 ms)
fail("target never accepted the interaction")
# placement B - around the case; setup runs again
for case in suite:
result = run_case(case) # setup + steps + teardown
if result.failed and case.attempts < 2:
first = result # keep it, do not overwrite
result = run_case(case)
result.passed_on_retry = true
result.first_attempt = first
# placement C - around the whole run; attribution is gone
result = run_suite()
if result.preparation_failed:
result = run_suite() # everything, from scratchgo deeper
Recall that failed work can be executed again and that a harness can be told to do it at more than one level. Before changing anything, know which level your own suite is using.
Explain the three units — one action, one case, one run — and what each re-executes, including that a case-level retry runs setup again. Be able to name a step that is unsafe to repeat and say why.
Demonstrate the diagnostic tradeoff: every step outward deletes a level of attribution and multiplies the time a genuine failure takes to report. Pick a placement for a described scenario and defend what it hides.
Own the policy across suites: which placements are permitted, who may add one, and how the combined multiplication is bounded so a red answer still arrives inside the time the organisation has agreed to wait.
A retry is a decision about **which unit of work is executed again after it fails**. The unit is the whole design: it decides what the results record can still tell you, how much a failure costs in time, and which side effects get repeated. Three units are available in almost any harness, and they are not interchangeable. ## The three placements at a glance | Placement | What runs again | What it conceals | What it costs | |---|---|---|---| | **Around an action** | one interaction, call or poll | that the step was ever unstable | milliseconds to seconds | | **Around a case** | the case from its setup onward | which step inside the case failed | one case's runtime | | **Around a whole run** | every case, plus the run's own preparation | which case failed, and how many did | the full run | ## Around an action The narrowest placement lives inside a helper: a `click_when_ready` helper re-issues an interaction until it is accepted, a request helper re-sends after a transport-level error, a poll re-reads a value until it settles. The case above it never learns any of this happened. That is the point of the placement, and also its danger. **Nothing above the helper can see the instability**, so a step that needs four attempts on every run looks identical in the results to one that succeeds immediately. Two rules keep this placement honest: - The retried operation must be **safe to execute more than once**. A read, a poll, or a write the system itself defines as repeatable is fine; a step that creates a record, sends a message, or spends a single-use credential is not. Wrapping one of those does not recover the step — it performs it twice, and a later assertion then fails on the duplicate rather than on the original fault. - The helper must **surface how many attempts it used** — a counter, a line in the case's output, a field on its result — or the placement is a blind spot by construction. ## Around a case The runner re-executes the failing case end to end. Setup runs again, every step runs again, and whatever the failed attempt left behind has to be cleaned up or tolerated. This placement is coarse: the record says *this case needed two attempts*, not *the third step needed two attempts*. It is also the placement whose cost is easiest to reason about, because the unit — one case — is the same unit the suite already schedules and reports on. Two consequences are easy to miss: 1. **Setup is re-executed.** A case whose setup claims a fixed data record, consumes a one-time credential, or seeds a value the first attempt already mutated will fail on the second attempt for a brand new reason, and the reported failure will be about setup rather than about the original defect. 2. **The first attempt's evidence must survive.** Once the second attempt passes, the interesting artefact is the first attempt's failure. A harness that overwrites the case's result with the winning attempt discards exactly the thing that would have made the case fixable. ## Around a whole run The widest placement re-launches everything. It conceals the most: after a whole-run retry the record shows one green run, with no case-level attribution at all unless each attempt's results are kept separately. It also re-pays the entire cost, which for a long suite is the difference between a fifteen-minute answer and a forty-minute one. There is one situation where it is nevertheless the right unit: a failure that took the run down **wholesale and non-specifically** — a shared dependency unreachable, preparation that never completed, the run aborted before cases were scheduled. There the individual results carry no information, and re-running cases one at a time would be re-running noise. A whole-run retry justified by anything narrower is a way of buying a green result rather than producing a signal. ## Choosing the placement Work outward from the narrowest unit that can actually recover: 1. **What exactly is unreliable?** If one operation is, retry the operation. If the case's own sequencing is, retry the case. If the run's substrate is, retry the run. 2. **Is repeating that unit safe?** Side effects decide this, not preference. 3. **What will the record still say?** Every step outward deletes one level of attribution. If the placement leaves nobody able to name what was unstable, it is too wide. 4. **What does the failure path cost?** A wide retry multiplies the time to red for a genuine failure, because a real defect now has to fail as many times as the placement allows before anyone hears about it. ## Placements compose by multiplication The last trap is that the three stack. Three attempts inside a helper, inside a case retried twice, inside a run retried once, is up to twelve executions of that operation and four executions of the case. Each number lives in a different configuration file, so nobody reads them together and the total is invisible. Whenever more than one placement is active, the worst case is the **product, not the sum** — write it down, and check it against how long the team has agreed to wait for a red answer.
- Which steps must never sit inside an action-level retry?Any step that changes state in a way a second execution cannot undo — creating a record, sending a message, incrementing a counter, spending a single-use credential. Retrying one of those does not recover the step, it performs it twice, and a later assertion then fails on the duplicate rather than on the original fault. Retry reads, polls, and operations the system itself defines as safe to repeat.
- If two placements are active at once, how many executions can one operation get?The product of the limits, not the sum. Three attempts inside a helper, inside a case retried twice, inside a run retried twice, is up to twelve executions of that operation and four of the case. Nobody sees the total because each limit lives in a different configuration file, so write the worst case down and compare it against how long a red answer is allowed to take.
Retrying one action is re-asking a single question; retrying a case is starting the whole conversation over; retrying a run is rebooking the meeting. Each step outward loses more of the record of what actually went wrong.
saying these in an interview costs you the question
- Treats re-running the whole run as the cheap option
- Wraps a record-creating step in a retry loop
- Cannot say which unit the retry re-executes
- Assumes a case-level retry skips setup
- Retries until it passes, with no cap
- Says a passing retry proves the failure was noise