What is a feedback budget for a test suite, and what happens when the suite outgrows it?
answer
- Ask what the result changes
- How long will someone actually wait
- Runtime is chosen, not discovered
- Judge the slow tail, not the mean
- Tier by risk before buying workers
basics
~20 sA feedback budget is the longest a check suite may run before its verdict stops changing what a developer does next. Past that limit people switch tasks, batch changes and stop running the suite locally.
solid answer
~50 sSuite runtime is not a fact you discover, it is a constraint you design to. Set the budget from how the result is used: single-digit minutes for the tier that runs on every push, because that is roughly how long someone will actually wait; tens of minutes for a shared-branch tier; hours only for a scheduled tier nobody is waiting on. Judge the suite against the budget at a high percentile rather than the mean, because the slow tail is what people plan around. When a suite outgrows its budget the damage is behavioural first: engineers stop running it locally, changes get batched, and a red run now covers many commits so attribution costs more than the fix. The legitimate remedies are tiering by risk, moving a check to the cheapest level that can hold it, parallelising, and fixing the slowest minority of cases - never deleting assertions to make the clock look better.
code
pseudocode · 18 linesBUDGET = { "commit": 8 * 60, "shared_branch": 25 * 60, "scheduled": null }
function report(tier, run_durations_seconds):
sorted_runs = sort(run_durations_seconds)
median = percentile(sorted_runs, 50)
p90 = percentile(sorted_runs, 90)
budget = BUDGET[tier]
if budget == null:
return { tier: tier, median: median, p90: p90, verdict: "unbudgeted" }
return {
tier: tier,
median: median,
p90: p90,
headroom: budget - p90,
verdict: (p90 <= budget) ? "inside budget" : "over budget at p90"
}go deeper
Be ready to say what a feedback budget is in one sentence and give a plausible number for a per-push tier. Knowing that a long suite makes people stop running it locally is the point being tested.
Explain the tiering mechanics: which checks belong on every push, which on the shared branch, which on a schedule, and why you judge runtime at a high percentile rather than at the mean.
Show that you have shortened a real suite. Talk about where the time actually went, why parallelism stopped paying, and how you separated the legitimate fixes from the ones that just make the number green.
Own the trade explicitly: a shorter budget means later detection somewhere else, and you should be able to say which risks you accepted moving out of the fast tier and what compensates for them.
### The budget is a decision, not a measurement A **feedback budget** is a ceiling you *choose* for how long a body of automated checks may take before its verdict arrives. It is derived from what the verdict is for, not from what the suite happens to cost today. The question that fixes the number is behavioural: *what will the person who triggered this run do while it runs?* If they will sit and watch, the budget is a couple of minutes. If they will start the next task, the budget is whatever keeps the context of the change recoverable when the result lands. If nobody is waiting at all — an overnight or pre-release tier — there is effectively no budget, and that tier is where genuinely slow checks belong. Most teams end up with two or three tiers rather than one number: - a **commit tier** that runs on every push and is measured in single-digit minutes; - an **integration tier** on the shared branch, tens of minutes, where slower checks live; - a **scheduled tier** that may run for hours because its result is read at a planning cadence. Each tier gets its own budget, and a check is placed in a tier by the risk it retires against the time it costs — not by which folder it lives in. ### Measure the tail, not the mean Runtime is a distribution. A suite whose mean is 12 minutes but whose slowest one run in twenty takes 31 minutes is experienced by the team as a 31-minute suite, because the slow runs are the ones people remember and plan around. Judge the suite against its budget at a high percentile — the ninetieth or ninety-fifth — and report that number alongside the median. Variance is itself a finding: a wide spread usually means contention for shared infrastructure, uneven sharding across workers, or a small number of very slow cases dominating one shard. ### What actually goes wrong when the budget is blown Take a ride-hailing dispatcher service whose check suite grew, over about eight months, from roughly nine minutes to **27 minutes**, with a ninetieth-percentile run of 34 minutes. Nothing about the suite was broken. What changed was behaviour: 1. Engineers stopped running the suite before pushing, because 27 minutes is longer than the pause between finishing a change and wanting to hand it over. 2. Because the local run disappeared, failures were first seen on the shared branch instead of on a workstation, so every failure now interrupted somebody else as well. 3. Changes got batched. When a red run covered eleven commits instead of one, attributing the failure took longer than fixing it, and the temptation to re-run rather than investigate rose. 4. The slow tier quietly became the only tier, so a check that was cheap and important sat behind 27 minutes of checks that were expensive and rarely interesting. The concrete damage in that team was a **permission escalation**: a change to a role check let a support-desk role reassign rides belonging to another operator. A case guarding exactly that rule existed and was correct — it simply lived in the tier nobody ran before pushing, and the failure was attributed three days later out of a batch of 22 commits. ### The legitimate ways back inside the budget Ranked roughly by how much they buy for how little they cost: - **Tier by risk.** Move checks that rarely find anything, or that only matter before a release, out of the per-push tier. This is the largest single lever and costs nothing but judgement. - **Push the check down a level.** A rule that can be verified against a small in-process unit does not need a full deployed environment to verify it. The cheapest level that can hold the assertion is where it belongs. - **Parallelise.** Split independent cases across workers. This shortens wall-clock time only up to the length of the slowest shard, and it stops helping entirely once fixture setup, container startup or a shared dependency becomes the bottleneck — so measure before buying more workers. - **Fix the slow minority.** Runtime is usually skewed: a handful of cases with real sleeps, heavy fixtures or unnecessary environment startup often account for a large share of the total. Report the slowest cases every week and treat the list as a work queue. - **Cache and reuse deliberately.** Reusing a prepared environment or prebuilt artefact is legitimate when the reuse cannot leak state between cases; when it can, you have traded runtime for nondeterminism, which is a bad trade. What is *not* a legitimate fix is weakening the oracle: deleting assertions, skipping cases, loosening a comparison so it stops failing, or reporting only the first failure so the run exits early. Each of those makes the number go green while removing exactly the signal the budget existed to deliver quickly. ### Saying it in an interview The sentence that lands is: *runtime is a constraint we design to, the same way we design to a latency budget, and the cost of blowing it is paid in behaviour before it is paid in minutes.*
- How do you pick the budget number itself, rather than copying someone else's ten minutes?Derive it from observed behaviour, not from a blog post. Watch what people do while a run is in flight: if they wait, the current runtime is already inside the budget; if they reliably switch tasks, the budget is shorter than what you have. Then check the cost side - how much rework a late verdict causes when a red run covers several changes. Write the number down per tier and revisit it when team size or change rate shifts.
- The suite averages twelve minutes but one run in twenty takes half an hour. Does that matter?Yes, and it is often the more important number. People plan around the worst case they have recently experienced, so a fat tail costs behaviour even when the mean is comfortable. A wide spread also points at a specific cause worth fixing: contention for shared infrastructure, uneven sharding so one worker carries the slow cases, or a small group of very slow cases. Report median and a high percentile together and treat the gap as a finding.
- What do you do when the suite genuinely cannot be made to fit the budget?Split it rather than shrink it. Keep a fast tier that runs on every push and holds the checks most likely to catch a mistake in the change being made, and move the rest to a shared-branch or scheduled tier with an honest, separate budget. Be explicit that the slower tier now detects later, and compensate with a policy for how fast its failures get triaged. That is a deliberate trade, unlike silently deleting cases.
It is the same logic as a page-load budget: nobody argues about whether the page is correct, they argue about whether anyone is still there when it finishes.
saying these in an interview costs you the question
- Treats suite runtime as a fixed property rather than a budgeted constraint
- Deletes or skips assertions to bring the clock inside the budget
- Quotes only the mean runtime and ignores the slow tail
- Assumes adding parallel workers always shortens wall-clock time
- Believes nobody minds waiting as long as the suite is thorough
- Puts every check in one tier because splitting feels like extra work