skip to content

How do you choose the regression subset to run for a change when the full pack no longer fits the feedback budget?

level: seniorimportance: should knowfreq 51%

answer

  1. The pack no longer fits the wait
  2. Map the change set to cases
  3. Tags, coverage records, failure history
  4. Something must still run everything
  5. Count what the subset missed

basics

~20 s

Select by change impact: the cases traceable to the changed components or requirements, cases asserting on what the change writes, cases that failed recently in that area, and the new case for the change. Keep the full pack running on a slower schedule as the safety net.

solid answer

~50 s

Start from a mapping between the change set and the pack. Three sources are usable together: **traceability** from requirement or component identifiers recorded on each case; **coverage-derived mapping**, recorded from a prior instrumented run, of which cases exercised which units; and **history**, meaning cases that failed recently or that historically fail when this area changes. Always add a small core that runs regardless of mapping, plus any case written for the change itself, then cut to the budget. Say out loud that selection is a bet: a code-derived mapping cannot see configuration, reference-data or dependency changes, and a stale traceability record silently degrades. The safety net is the periodic full run, and the calibration signal is the escape rate — defects the full run finds that the selected subset should have caught. Feed those back into the mapping instead of widening the subset by feel.

code

pseudocode · 14 lines
pseudocode
def select_subset(change, pack, budget = minutes(11)):
    candidates = set()
    candidates |= pack.tagged_to(components_of(change.files))       # traceability
    candidates |= pack.covering_units(change.changed_units)         # coverage-derived
    candidates |= pack.asserting_on(shared_artefacts_written_by(change))
    candidates |= pack.failed_within(last_runs = 10)                # history
    candidates |= pack.always_on_core
    candidates |= change.new_cases

    if change.touches_reference_data or change.touches_config:
        candidates |= pack.area("pricing")     # no changed unit to map from

    ordered = sort_by(candidates, key = historical_failure_rate, desc = True)
    return take_while_under(ordered, budget)

go deeper

for a junior

Know that not every case runs on every change, and be able to name one basis for choosing — the area that changed. Understand that a slower run over everything still happens, so the subset trades feedback speed for immediate completeness.

for a middle

Explain the mapping sources — component or requirement tags on cases, coverage records from an instrumented run, recent failure history — and the always-on core. Be ready to say why a code-derived mapping misses configuration and reference-data changes.

for a senior

Demonstrate that you treat selection as a measured bet: an escape count at the boundary, mapping rules added in response to real escapes, explicit handling of non-code changes, and the periodic full run kept green rather than kept nominal.

for a principal

Own the trade between the feedback budget and what the organisation accepts finding hours later, and the decision to invest in mapping quality rather than more machines. Be ready to defend the budget number itself with escape and detection data.

## The situation A regression pack of, say, 3,140 cases takes a 6-hour nightly run — 6h 12m on a good night. Nobody is going to wait for it on every merge. The team wants feedback in about 11 minutes, which is roughly 180 cases. So on each change you run a subset, and the whole craft is in choosing which 180. ## Four inputs to the selection **1. Traceability.** Each case carries identifiers: the requirement it verifies, the component or feature area it belongs to. A change set names files, which map to components. Intersect the two. This is the cheapest mapping to build and the easiest to let rot — an untagged case is invisible to selection forever, and a case tagged to a component that was renamed silently drops out. Treat the tags as production data with an owner, or do not rely on them. **2. Coverage-derived mapping.** From a prior instrumented run, record which cases executed which code units. Selection becomes: take the changed units, look up the cases that executed them. This is the most precise input, and it has three costs — instrumented runs are slow, the mapping must be re-baselined as the code moves, and it is blind to anything that is not code. **3. History.** Cases that failed in the last several runs, and cases that historically fail when this area changes, are cheap to compute and empirically worth including. Ordering the selected subset by historical failure rate also shortens the time to the first failure, which matters when a developer is watching the run — though it does not change what the subset covers, only when you learn about it. **4. Always-on core.** A short list that runs regardless of mapping: the critical paths, the cases covering the most recent production incidents, and anything asserting on shared downstream artefacts. This is your hedge against every mapping being wrong. Then add the case written for the change itself, and cut to the budget. ## Where the mapping fails, concretely On the fare calculator for a public-transport network, a change to the refund-ledger writer was selected correctly by the code-derived mapping — but the mapping returned only cases tagged to the fare-total component, because the cases asserting on **ledger entry counts** lived under a different component tag. The change made each capped day write nine ledger entries instead of one. The duplicated side effect passed the 11-minute subset and was caught 6 hours later by the nightly run. The fix is not "run more cases". It is a mapping rule: a change that touches a writer of a shared artefact selects the cases that assert on that artefact, whatever component they are tagged to. Escapes are how you find the rules you were missing. ## What selection cannot see - **Non-code changes.** A tariff-table update, a configuration edit, a feature-flag default, a dependency upgrade. There is no changed unit to map from, so these need explicit rules — a reference-data change selects the whole pricing area, a dependency bump selects the always-on core plus the integration cases. - **Indirect coupling.** Shared caches, shared state, ordering assumptions, timing. Nothing static maps these. - **Anything the mapping was never given.** Untagged cases, cases added since the last instrumented baseline. ## The safety net is not optional Selection changes *when* something is found, not *whether* the pack can find it — provided the full pack still runs on a schedule. Teams that stop the periodic full run once selection is working have converted a latency optimisation into a coverage loss, and the loss is invisible until an escape reaches users. Keep the full run, and keep it green: a nightly run with 23 long-standing failures nobody triages is the same as not having one. ## Measuring whether the selection is any good Two numbers, both cheap: - **Escape rate at the boundary.** For every defect the full run catches, ask whether the selected subset could have caught it inside its budget. If yes, it is a mapping defect — add the rule. - **Time to detection.** How long between the change landing and someone knowing. If most defects are found 6 hours later, the subset is too narrow or the budget is too tight; that is a conversation about the budget, not a reason to guess. A subset chosen by feel and never measured is indistinguishable from a subset chosen at random, and in an interview the measurement half is what separates a real answer from a plausible one.

  • A team stops the nightly full run because per-change selection is catching everything. What do you say?
    That selection was a latency optimisation and they have quietly turned it into a coverage decision. Every mapping has blind spots — untagged cases, indirect coupling, non-code changes — and the periodic full run is the only thing that finds them before users do. It is also the source of the escape data that tells you the mapping is degrading. If the full run is too expensive to keep nightly, move it to a slower cadence deliberately and say what that costs, rather than deleting it because it has been quiet.
  • How does ordering the selected subset by historical failure rate help, and what does it not change?
    It shortens the time to the first failure, so a developer watching an 11-minute run often learns within the first two minutes. That matters for feedback and for cutting a run short once it has already told you something. What it does not change is coverage: the same cases run, so a defect only the last case would catch is found at the same moment either way. Treat it as a latency trick, never as a substitute for widening a subset that is missing cases.
  • A configuration-only change lands. Which cases should the selection pick?
    None, if the mapping is code-derived — which is exactly the failure. Configuration and reference-data changes need explicit rules, because there is no changed unit to trace from: map the changed key or table to the functional area it drives and select that area, plus the always-on core. Keep those rules alongside the selection logic and review them when the configuration surface changes, since they are hand-maintained and rot in the same way traceability tags do.

saying these in an interview costs you the question

  • Runs the entire pack on every change and calls it thorough
  • Selects cases by intuition with no traceability or history
  • Drops the periodic full run once selection appears to work
  • Assumes a code-derived mapping covers configuration and data changes
  • Never measures what the selected subset missed

context