skip to content

What are suspension and resumption criteria for a test cycle, and when do you invoke them?

level: seniorimportance: nice to knowfreq 24%

answer

  1. Rules for stopping, agreed in advance
  2. Stop when failures stop carrying information
  3. One defect blocking most of the plan
  4. Restarting needs a build and a re-run list
  5. The pause has a cost you can state

basics

~20 s

Suspension criteria are pre-agreed conditions for stopping a test cycle in flight — a build too broken to yield information. Resumption criteria state what must be true to restart, including which earlier results the interruption invalidated.

solid answer

~50 s

A **suspension criterion** stops a cycle mid-flight when further execution would produce no usable information: the smoke subset fails on the delivered build, a single defect blocks more than an agreed share of the plan, the environment has been unusable beyond an agreed duration, or builds are arriving so fast that no result survives long enough to be reported. A **resumption criterion** says what must hold before execution restarts — typically a fixed build with a published change list, the blocking defect confirmed fixed, the environment verified, and an explicit list of results the interruption invalidated and must be re-run. The pair exists so that stopping is a decision the team made in advance rather than an argument someone loses under pressure. Without them, teams keep grinding, generating failures that all trace to one cause and burning the credibility of the report.

go deeper

for a junior

You are unlikely to be asked this, but know that a cycle can legitimately be halted when the build is too broken to yield information, and that restarting means re-running some of what you already did.

for a middle

Be able to name a couple of quantified triggers — smoke failing on the delivered build, one defect blocking most of the remaining plan — and to explain why they are agreed before the cycle rather than argued during it.

for a senior

Show that you have made the call: the threshold you used, where the effort was redirected during the pause, and how you assembled the list of results the interruption invalidated. Being able to price the pause is the mark of experience here.

for a principal

Own the incentive design. Set thresholds loose enough that they are rarely met and mechanical enough that invoking one is never read as politics, and make sure a suspension is published with its cause so the pause is a shared fact.

## The gap these fill Entry criteria decide whether a cycle may start; exit criteria decide whether a release may ship. Neither says anything about the most common awkward moment: the cycle started legitimately and then the ground moved. **Suspension** and **resumption criteria** are the pre-agreed rules for that moment. The reason they are written in advance is the same reason entry criteria are: at the moment they are needed, the person proposing a stop looks like the person creating a problem. A criterion agreed a week earlier converts that from an individual's resistance into a rule the team already accepted. ## What a suspension criterion looks like Useful ones are quantified and observable. Common families: - **Smoke failure.** The smoke subset fails on the delivered build. Continuing means executing hundreds of cases against something that cannot start correctly. - **Blockage share.** A single defect blocks more than an agreed fraction of the remaining plan — say a third. Below that threshold you route around it; above it you are mostly measuring the same defect repeatedly. - **Environment unavailability.** The environment or a critical collaborator has been down beyond an agreed window, so results are unreproducible. - **Build churn.** New builds are landing faster than a cycle can complete against one of them, so no result can be attributed to a build. - **Defect arrival rate.** New defects are arriving faster than the team can triage them, which usually means the change was not ready and further execution just deepens a backlog. Notice that all five are about *information yield*, not about difficulty. A hard release with many real defects is not a reason to suspend; a release where the failures no longer tell you anything new is. ## What a resumption criterion looks like Resumption is not simply the negation of suspension, and this is where teams get it wrong. Three parts: 1. **The cause is confirmed fixed** on a new, identified build — confirmed by a check, not by a developer's assertion that the change is in. 2. **The environment is verified**, usually by re-running the smoke subset on that build. 3. **The invalidation list is settled.** The interruption did not merely pause the cycle; it retroactively devalued results. Anything that ran against the broken state, and anything the fix could plausibly touch, is marked for re-execution. This is the expensive part, and the reason a suspension has a real cost that should be weighed rather than treated as free. ## A worked case An 11-person team is two days into a cycle on a document e-signing flow. Midway through day two, a change to the envelope-limit check lands: an off-by-one at the signer boundary now rejects an envelope with exactly the maximum permitted signers instead of the one above it. Because almost every multi-signer scenario in the suite builds an envelope at or near that limit, **106 of the 174 remaining cases** cannot reach their starting state. The agreed suspension criterion — a single defect blocking more than a third of the remaining plan — is met, so the cycle is suspended rather than ground through. Two things follow. Execution effort redirects to areas untouched by the boundary rather than idling, and the team publishes the suspension with its cause, so the pause is visible to everyone waiting on the cycle. When the fix arrives the next morning, resumption requires the new build identified, the boundary case confirmed on it, smoke green, and 23 earlier results re-run because they exercised the same limit logic. Total cost of the pause: roughly half a day plus 23 re-runs — which the team can state, instead of a vague sense that the cycle went badly. ## The failure modes **No criteria at all** is the common one. Testing continues out of momentum, dozens of failures are filed that all trace to one cause, developers lose trust in the reports, and the cycle ends with a pass rate nobody believes. **Criteria too tight** is the opposite: any inconvenience triggers a stop, so the team stalls repeatedly and the criteria get ignored, which is worse than not having them. **Suspension as escalation** is subtler and genuinely damaging. If stopping is used as leverage in an argument with development, the mechanism stops being a shared rule and becomes a weapon; the next time a real suspension is warranted, it is read as politics. The protection against this is that the criteria are quantified, agreed jointly and applied mechanically — you invoke them because the number was met, and you say so. **Silent resumption** is the last one: work restarts because the blocker "seems fixed", nobody names the invalidated results, and stale passes from before the interruption are carried into the exit report. That is precisely how a defect the cycle already touched escapes to production. ## Why this is worth knowing Most interviews will not ask for this vocabulary by name, and not knowing the terms says little about a candidate. What the question really probes is whether you have ever been in a cycle that should have stopped and did not — and whether you drew a rule from it. The vocabulary is a convenience; the discipline of deciding in advance what makes execution pointless is the substance.

  • What is the cost of suspending a cycle, and how do you quantify it?
    The visible cost is idle execution time, but the real cost is the invalidation list — results that ran against the broken state and must be re-run once the cause is fixed. Quantify it as elapsed pause plus the number of re-executions, and report both. A team that can state 'half a day plus 23 re-runs' can weigh a suspension against grinding on; a team that cannot is guessing in both directions.
  • How do you stop a suspension criterion from being used as leverage against the development team?
    Quantify it, agree it jointly before the cycle, and apply it mechanically — you invoke it because a stated threshold was crossed and you say which one. The moment a stop is a judgement call made in a dispute, it reads as politics, and the next legitimate suspension gets fought rather than accepted. Publishing the cause with the suspension does most of this work on its own.

A referee stopping play for a hazard on the pitch: the rule is written before the match precisely so the call is not read as favouring one side.

saying these in an interview costs you the question

  • Says testing should never stop, only slow down
  • Suspends on any inconvenience rather than an agreed threshold
  • Treats resumption as simply undoing the pause
  • Restarts without naming which results were invalidated
  • Uses a stop as leverage in a dispute with developers
  • Cannot state what a suspension cost in time or re-runs

context