Your verification suite has grown to 40 minutes and currently runs in full on every push. How would you decide which work runs on which trigger?
answer
- who is waiting for this result
- measure runs, minutes and defect yield first
- relocating risk, not deleting it
- the revert path licenses the demotion
- the nightly nobody watches
basics
~20 sTier the work by who is waiting for it: fast checks on every proposal update, the full gate where changes land, and slow or drift-detecting suites on a schedule. Each move trades earlier feedback for compute, or compute for later discovery of a failure.
solid answer
~50 sStart by measuring, not rearranging: runs per day, minutes per run, and which jobs actually catch defects. Then tier by audience. Work with a human waiting — lint, type checks, unit tests, the build — stays on every proposal update and should target roughly ten minutes. Work that protects everyone else — the full integration suite — runs where the change lands, so the trunk keeps its guarantee even if proposals do not. Work with nobody waiting — long browser matrices, performance runs, dependency and vulnerability checks — moves to a schedule, because its job is catching drift rather than gating a change. Manual runs cover the expensive and irreversible. The price of every demotion is that a failure is found later, so each one needs a compensating control: a fast revert path, someone actually watching the scheduled results, and enough observability to catch what escapes. If your trunk cannot be reverted quickly, you have not earned the right to move checks off the merge gate.
go deeper
Know that not every check has to run on every push, and that fast feedback for the author comes first while slow suites can run later.
Explain the tiers concretely — fast checks on proposals, the full suite where changes land, slow and drift checks on a schedule — and what each move costs in detection time.
Show you have measured a real pipeline: runs by trigger, per-job duration and queue wait, and which jobs actually caught defects, then defended a specific demotion with the control that compensates for it.
Own the tradeoff as policy: the per-merge economics across teams, the revert and rollback capability that licenses moving checks off the gate, ownership of scheduled results, and periodic re-review of what you removed.
## Measure before you move anything The instinct is to start cutting. Resist it for one afternoon and collect four numbers: 1. **Runs per day, by trigger.** Where the minutes actually go — usually proposal updates, not merges. 2. **Per-job duration and per-job queue wait.** If jobs sit queued, you have a capacity problem and rearranging triggers will not help. 3. **Defect yield per job.** Over the last few hundred runs, which jobs ever failed *for a real reason*? A twelve-minute job that has never caught anything is not a gate, it is a tax. A job that fails constantly for flakiness is worse: it has trained people to re-run without reading. 4. **Time-to-detect and time-to-revert.** How long between a bad change landing and someone knowing, and how long to undo it. The fourth number is the one that licenses everything else. ## The tiering framework Ask of each job: *who is waiting for this result, and what do they do with it?* **Tier 1 — proposal updates (a human is blocked).** Lint, type checks, unit tests, the build, and any check that gives an author a specific, actionable message. Target something like ten minutes end to end, because past roughly fifteen minutes people context-switch and the review loop stretches by hours. This tier is optimised for latency. **Tier 2 — where the change lands (the team is exposed).** The full integration suite, contract tests against real dependencies, the artifact build that will actually be deployed. Optimised for coverage. Running this here rather than on every proposal update is the single largest saving available, because proposals update many times and merge once. **Tier 3 — scheduled (nobody is waiting).** Long browser or device matrices, performance and load runs, dependency and vulnerability checks, rebuilds against refreshed base images, flake-hunting repeat runs. Optimised for thoroughness and for detecting changes that no commit caused. Costs are bounded and predictable. **Tier 4 — manual or event-driven.** Deploys, migrations, expensive experiments, one-off matrix expansions. Optimised for deliberateness: the gate is a person. ## What each demotion actually costs Moving a check from tier 1 to tier 2 means broken changes can be approved and merged, and the trunk breaks instead of the proposal. That is acceptable *only* if reverting is routine — a one-click revert, a deploy that can be rolled back, and a culture that reverts first and diagnoses afterwards. Moving from tier 2 to tier 3 means a defect can live in the trunk for hours, so it is acceptable for classes of failure that are not user-visible in that window, or where a separate runtime signal would catch them first. The honest framing for an interview: **you are not deleting risk, you are relocating it in time.** Every move must name the control that catches what you let through. ## The failure modes of tiering - **The unwatched nightly.** A scheduled suite with no alert route fails for six weeks and everyone learns to ignore the red. If a tier-3 job has no owner and no notification, deleting it is more honest than running it. - **A trunk nobody can revert.** Teams move checks off the proposal gate, then discover a revert requires a data migration to be undone. The demotion was never affordable. - **Flaky tests promoted rather than fixed.** Moving a flaky suite to nightly to stop it blocking merges converts a visible problem into an invisible one. Quarantine and fix, or delete it. - **Tiering as a substitute for a slow suite.** If a single job takes 25 minutes because it was never optimised, moving it does not make it correct — it just means you pay the same time later. ## Guardrails that make tiering safe - A revert that is genuinely one action, and a deploy that can be rolled back without a human deciding anything clever. - Scheduled failures routed to a real on-call or team channel with an owner, not to an inbox. - A visible trunk-health signal, so "is the trunk green?" is answerable in seconds. - A rule that any check demoted for cost is reviewed again after a quarter, because the cheapest job today is the one that never catches anything and the world changes. ## Saying it in an interview A strong answer names the axis (who is waiting), gives the tiers, states the price of each demotion, and — crucially — refuses the demotion when the revert path is weak. A weak answer just proposes running less. The last note worth making is that this is not the only lever: making the suite faster and reducing what a change forces you to verify are the other two, and tiering is what you do when those are already exhausted or when the work is intrinsically slow.
- What has to be true about your revert path before you move a check off the proposal gate?Reverting must be routine and fast: one action to undo the change, a deploy that rolls back without clever human decisions, and no irreversible side effect — a one-way data migration, a published artifact consumers already pulled — in the window. If undoing takes a coordinated effort, a broken trunk is an outage rather than an inconvenience, and the check has to stay on the gate.
- A twelve-minute job on every proposal has not caught a real defect in six months. Do you move it or delete it?Ask what class of failure it is supposed to catch and whether that class can still occur. If it can — a rare but severe regression — move it to a scheduled tier with a real owner. If the class is gone, or the job's assertions no longer test anything meaningful, delete it. Keeping an unwatched job that catches nothing costs money and teaches people to ignore red.
- How do you stop the scheduled tier from becoming a graveyard of ignored red builds?Give every scheduled pipeline a named owner and a notification route into a channel that is actually read, and treat a failure there with the same triage discipline as a broken trunk. Track how long it stays red. If nobody will own it, that is a decision: delete the job rather than run something whose result changes nobody's behaviour.
saying these in an interview costs you the question
- Proposes running less without measuring what each job catches
- Treats tiering as removing risk rather than delaying detection
- Moves flaky suites to nightly instead of fixing or quarantining them
- Demotes checks without a fast, reliable revert path
- Adds scheduled pipelines with no owner and no alert route