skip to content

How do you decide the tier structure for a regression pack, and what confidence does each tier actually buy?

level: principalimportance: should knowfreq 43%

answer

  1. Feedback time bought against confidence
  2. Every tier needs a runtime budget
  3. Earliest tier that can hold it
  4. Write what green means
  5. Escapes and time to a human

basics

~20 s

Work backwards from the feedback time each decision can tolerate, then place every signal in the earliest tier that can produce it reliably inside that tier's budget. Every tier needs an owner, a blocking action and a written claim about what green means.

solid answer

~50 s

Tiering is a cost-of-delay decision, not a tooling one. Set a runtime budget per trigger — a short gating subset on every build, a change-selected subset on every merge, the full pack on a slower cadence, and a deep tier for signals that are slow by nature — and treat each budget as a hard constraint. Assign a signal to the earliest tier that can produce it deterministically within budget; signals that genuinely need hours stay late, and faking them early buys noise. Then write down, per tier, what a green run does and does not mean, so nobody reads "green subset" as "safe change". Govern it with two measurements: escapes at each boundary — defects a later tier caught that an earlier one could have — and time to a human acting, not just time to a red result. Move cases when those numbers say so, and delete any tier that blocks nothing, because an advisory tier stops being read within weeks.

code

pseudocode · 18 lines
pseudocode
tiers = [
  { name: "gating",   trigger: "every build",  budget: seconds(280), blocking: true },
  { name: "selected", trigger: "every merge",  budget: minutes(11),  blocking: true },
  { name: "full",     trigger: "nightly",      budget: hours(7),     blocking: true },
  { name: "deep",     trigger: "weekly",       budget: hours(4),     blocking: true },
]

def place(signal):
    for tier in tiers:                       # earliest tier that can hold it
        if signal.runtime <= tier.budget and signal.deterministic:
            return tier
    raise "no tier can produce this reliably"

def review_placement(defect):
    caught = defect.caught_by_tier
    for tier in tiers_before(caught):
        if could_have_caught(tier, defect):
            record_escape(boundary = (tier, caught))   # movement signal

go deeper

for a junior

Know that a regression pack is usually split so that a short set runs constantly and longer sets run less often, and that the split exists to keep feedback fast. Be able to say why running everything on every change is impractical.

for a middle

Explain the placement rule and the budgets: each tier has a runtime it must fit, and a signal goes in the earliest tier that can produce it reliably. Be ready to say why some checks genuinely cannot be moved earlier.

for a senior

Show that you govern the structure with data — escapes at each boundary, time until someone actually triages a failure — and that you fix placement rather than widening a tier after every incident. Expect to describe a case you moved between tiers and why.

for a principal

Own the honest claim per tier, the budget negotiation with delivery, and the risks the structure deliberately leaves to be found late. Be ready to argue against an advisory tier and to defend the numbers with detection and escape evidence rather than anecdote.

## The trade being made Every tier boundary trades **feedback time** against **confidence**, and both sides have a real cost. Waiting costs: the author has moved on and must reload the context; more changes pile onto the same build, so a failure is harder to attribute; a defect found six hours later is found against a batch rather than a change. Confidence costs: whatever the earlier tier does not cover is found later, or by users. So the design question is not "how do we test more" but "what is the latest we can afford to learn each class of thing, and what does learning it earlier cost per run?" ## A concrete structure For a fare calculator serving a public-transport network, a defensible shape: | Tier | Trigger | Budget | What it claims | |---|---|---|---| | Gating subset | every build | 4m 40s | The build stands up and its critical paths are alive | | Selected subset | every merge | 11m | No regression in the areas this change maps to | | Full pack | nightly | 6h 12m | No known regression anywhere in the pack | | Deep tier | weekly and pre-release | 3h 20m | Long-running behaviour: month-long ticketing replays, data migrations, sustained-load behaviour | The numbers are the point. A tier without a budget grows until it collides with the next one, and then the structure is decided by runtime rather than by design. ## The placement rule Assign each signal to **the earliest tier that can produce it deterministically inside that tier's budget**. Three consequences follow: - Signals that are slow by nature — a replay of a month of ticketing history, a migration rehearsal — belong late. Compressing them into a fast tier produces a weaker check wearing the same name, and people then believe the fast tier covers it. - A signal that a fast tier *could* produce and does not is a placement defect, and it shows up as escapes. - Duplicating one signal across four tiers is a decision to pay for it four times. Decide which tier **owns** each signal, and let the later tiers own what the earlier ones cannot. ## The written claim Each tier needs one sentence saying what a green run means. This sounds bureaucratic and is the highest-leverage artefact in the whole structure, because the default reading of any green tier is "safe". A merge-tier claim reads: *no regression in the areas this change maps to; nothing about areas the mapping missed, about configuration changes, or about long-running behaviour.* Once that is written, nobody has to argue about it during an incident. ## Governing it with two numbers **Escapes per boundary.** For each defect, record which tier caught it and ask whether an earlier tier could have, inside budget. A repeated pattern is a movement signal. A real one from the fare system: changes to a shared ledger writer produced a duplicated side effect — each capped day writing nine refund entries instead of one — and the merge tier missed it three times in a quarter because the ledger-integrity cases sat only in the nightly tier. The response was to promote a 51-second ledger-integrity check into the merge tier and re-baseline that tier's budget, not to widen the merge tier by feel. **Time to a person acting.** A nightly tier that goes red at 02:40 and is triaged at 15:00 has a detection time of hours and a response time of most of a day. The response time is the number that matters, and it is the one teams never measure. A tier whose failures wait a day is, for change-attribution purposes, barely better than the next tier down. Also watch the **false-alarm rate**. A tier that cries wolf loses its blocking power politically long before anyone proposes removing it — the bypass becomes routine, then normal. ## Anti-patterns worth naming in an interview - **The advisory tier.** Reports, never blocks. Read for three weeks, ignored thereafter, and still quoted as confidence in planning meetings. Either give it a blocking action or stop counting it. - **The inherited structure.** Adopted from another team or from a default configuration, never measured against this system's escape data. - **Confidence inflation.** "The merge run is green" reported as "the change is safe". This is what the written claim exists to prevent. - **The unowned tier.** No named owner means no triage, and an untriaged red tier decays to permanently red, at which point it provides nothing while still costing its runtime. ## What a principal is really being asked Not the tier list — everyone can recite one. The interviewer wants to hear that you set budgets deliberately, place signals by the earliest-reliable rule, write down what each tier's green means, and move work between tiers on evidence rather than on the last incident's emotional weight. And that you can say plainly which risks the structure deliberately leaves to be found late, because a structure that claims to leave none is not being honest about its budgets.

  • Someone proposes a new non-blocking tier so the team can 'see the signal without being blocked'. What is your answer?
    Ask what action a red result obliges anyone to take. If the answer is none, the tier will be read for a few weeks and then ignored, while still costing its runtime and still being quoted as confidence in planning. The honest options are to make it blocking with an owner and a triage expectation, or to run it as a time-boxed experiment with a date on which it either becomes blocking or is removed. A permanent advisory tier is decoration with a budget.
  • How do you decide the runtime budget for the merge tier rather than inheriting a number?
    Start from behaviour: how long the team will actually watch a run before switching context, typically minutes rather than tens of minutes. Then check the number against evidence — if most defects are being caught by the nightly tier and the escape review says the merge tier could have caught them inside a slightly larger budget, that is a concrete argument for raising it. Budgets should move on escape and detection data, not on whoever last got burned.
  • Two tiers both run the same expensive scenario. Is that duplication always waste?
    Usually, but not always. If the later tier runs it against a fuller environment or fuller data, the two are different signals wearing one name and both may be justified — as long as the difference is written down. If they are genuinely identical, decide which tier owns the signal and let the other stop paying for it. The test is whether a green result in the earlier tier makes the later run's result predictable; if it does, you are paying twice for one answer.
  • How do you present the structure to people outside engineering?
    As what is known when, not as a list of runs. Say which classes of problem are known within minutes of a change, which are known by the next morning, and which are only known weekly or before a release — and name what is deliberately left to be found late. That framing lets a delivery decision be made against real detection times, and it exposes disagreements about acceptable risk while there is still time to change the structure.

saying these in an interview costs you the question

  • Copies a tier structure from another team without measuring anything
  • Adds an advisory tier and calls it extra confidence
  • Claims a green change-selected run means the change is safe
  • Sizes tiers by whatever the tooling does by default
  • Ignores how long a nightly failure waits before anyone triages it
  • Says every tier should run everything for maximum safety

context