skip to content

An automated regression suite has doubled in size and upkeep cost - which cases do you delete, and how do you justify it?

level: principalimportance: should knowfreq 44%

answer

  1. Every case charges rent
  2. Duplication across levels is the pool
  3. Never failed is weak evidence
  4. Write down the risk you accept
  5. A budget stops the regrowth

basics

~20 s

Treat every case as paying rent in runtime, upkeep edits and triage time, and earning it by covering a risk nothing cheaper covers. Delete duplicates of lower-level checks, change-detector cases and cases guarding removed behaviour - in reviewable batches, with the accepted risk written down.

solid answer

~50 s

I start from the budget and the debt, not from individual cases. Every case charges rent: runtime, upkeep edits when unrelated code moves, and triage time when it fails. It earns that rent by covering a risk nothing cheaper already covers. So the candidates are cases duplicating a check that exists at a cheaper level, change-detector cases that fail only because implementation moved, cases asserting behaviour the product no longer has, and long-quarantined cases nobody owns. I keep the cheapest level that can detect each fault, plus one end-to-end pass over the critical path. I would not delete on "it has never failed" alone - a regression guard is insurance, and the evidence linking coverage numbers to escaped defects is contested. Deletions go in reviewable batches with the accepted risk recorded, and then I stop the regrowth with a wall-clock budget and an owner per area.

go deeper

for a junior

Know that suites need pruning at all, and that the safest thing to remove is a case checking something a faster, lower-level case already checks. Do not delete a case you find inconvenient without asking what it guarded.

for a middle

Be able to name the candidate categories - duplicated across levels, change-detector, guarding removed behaviour, long-quarantined - and explain why upkeep edits per change are a better cost signal than raw case count.

for a senior

Show judgement about evidence: why never having failed is weak grounds, why a coverage percentage settles nothing, and how you would confirm that a cheaper case still catches the fault before removing the expensive one.

for a principal

Own the whole loop - the rent ledger, the budget per tier, the entry cost at expensive levels, the written record of accepted risk, and the scoreboard change needed so that deleting a redundant case is not punished.

## Test debt is debt, and the interest is the problem A suite that has doubled is not automatically worse. What makes it a liability is the interest it charges: every unrelated change forces edits in cases that never intended to depend on it, every red run costs triage time, and the nightly cycle grows until it no longer fits the night. On one clinical-trial data capture suite the nightly run reached 6 hours 14 minutes, which meant a failure found at 02:00 could not be re-verified before the next working day - the feedback loop had quietly stopped being a loop. So the framing to bring to this question is a rent ledger, not a purity argument. Each case pays: runtime, environment cost, upkeep edits per unit of product change, and triage minutes per red run. Each case earns: a risk covered that nothing cheaper covers, weighted by the blast radius of the behaviour and by whatever defect history the case actually has. Deletion is what you do when the rent exceeds the earnings, and the interview is testing whether you can hold both sides of that ledger at once rather than only the side you find congenial. ## The four honest candidates **Duplicated coverage across levels.** The same rule verified at unit, service and full-journey level. This is the largest and safest pool. One suite had 214 cases asserting the same date-format rule at three levels; removing 61 of the top-level duplicates cut 71 minutes from the nightly run and lost no detectable coverage, because a lower, faster case failed on the same fault. The rule of thumb: keep the cheapest level that can actually detect the fault, plus one end-to-end pass so the pieces are known to be wired together. **Change-detector cases.** Cases that fail whenever implementation moves even though behaviour has not. They are pure interest - all upkeep, no signal - and they are usually recognisable from their edit history: touched in every refactor, never once red for a real defect. **Cases guarding behaviour that no longer exists.** Removed features, retired flows, an old error format kept alive only because a case asserts it. These are found by reading, not by metrics. **Long-quarantined, unowned cases.** A case that has not gated anything for months has already stopped covering the behaviour; the only question is whether anyone will make that explicit. ## The argument you must not overreach on "This case has never caught a defect, so it is worthless" is seductive and wrong as a sole justification. A regression guard is insurance against a change that has not been made yet; its value shows up in the counterfactual, which no report contains. Nor should you lean on a coverage percentage in either direction - the relationship between coverage numbers and escaped defects is genuinely contested in the literature, and a number that cannot distinguish an executed line from a verified behaviour will not settle a deletion argument. Techniques that estimate detection power more directly, such as seeding faults and measuring how many cases notice, give a better signal, but they are expensive to run and are a supporting argument rather than a licence. ## Governance: how deletions survive contact with the team Bulk deletion by one enthusiastic engineer is how a team loses trust in the whole exercise. Three habits keep it safe. Delete in **reviewable batches** with one written rationale per batch, so a reviewer can argue with a category rather than with 61 individual diffs. Record **what risk is accepted and by whom** - one line per batch naming the behaviour no longer guarded and whether anything cheaper still covers it. And **measure afterwards**: run time recovered, upkeep edits per change, red runs per week, and any escaped defect that lands in the area you pruned. If a deletion was wrong, the record tells you which decision to revisit instead of leaving the team with a vague sense that pruning is dangerous. ## Stopping the regrowth Pruning without changing the inflow buys a year at most. Three levers hold the line. A **wall-clock budget per tier**, stated up front: this tier must fit its window, and adding a case at that level means removing one or making it faster. An **entry cost at the expensive levels**, so a new full-journey case has to argue why the check cannot live lower. And **ownership per area**, because an unowned suite is one nobody prunes. Watch the incentives too. A team measured on a coverage percentage will not delete cases, because any drop in the number reads as a regression regardless of what was removed. If you want pruning, you have to stop scoring the thing that punishes it, and score run time, escape rate and upkeep cost instead. ## What a strong answer sounds like It is explicit that deletion transfers risk rather than eliminating cost, it names the categories rather than a single heuristic, it refuses to over-claim on contested evidence, and it ends on the inflow rather than on the cleanup - because the interviewer is asking how you would run this for three years, not how you would spend one afternoon.

  • How do you find duplicated coverage across levels without reading every case?
    Work from the behaviour rather than the code: list the rules for an area and map which levels assert each one, which usually exposes clusters immediately. Edit history helps - cases touched together in every change are often checking the same thing - and seeding a fault and observing how many cases go red gives direct evidence of overlap, at a cost you would only pay for a high-value area.
  • What do you measure after a pruning round to know it was safe?
    Run time recovered and upkeep edits per unit of change on the value side, and on the risk side any defect that escapes into the area you pruned, plus whether the remaining cases still fail when you seed a fault there. Keep the written record of accepted risk so an escape points at a specific decision rather than at pruning in general.
  • A team is measured on a coverage percentage. How does that affect this work?
    It effectively forbids it, because any deletion that lowers the number reads as a regression however redundant the case was. If pruning is the goal, the scoreboard has to change first - run time against the budget, escape rate, and upkeep cost per change are all measures that reward removing a duplicate rather than punishing it.
  • Who should be allowed to delete an automated case?
    The area's owners, through the normal review path, in batches with a written rationale - not a single engineer acting unilaterally and not a committee that never convenes. The reviewable batch is what makes the decision arguable at the level of a category, which is where the real disagreement lives.

Pruning a suite is closer to editing a manuscript than to clearing a warehouse: you are not removing the least-used sentences, you are removing the ones that say what another sentence already said better.

saying these in an interview costs you the question

  • Deletes every case that has never failed
  • Refuses to delete anything because coverage might drop
  • Removes cases in one enormous unreviewable change
  • Cites a coverage percentage as proof of redundancy
  • Prunes once and changes nothing about the inflow
  • Treats deletion as free rather than as accepted risk

context