skip to content

How do you buffer a test estimate for fix-and-re-test cycles and days lost to blocked builds?

level: seniorimportance: should knowfreq 51%

answer

  1. The first pass is not the whole job
  2. Defects arrive, get fixed, force re-runs
  3. One shared defect can invalidate a slice
  4. Availability divides, churn adds
  5. Name the buffer and its trigger

basics

~20 s

Model the drivers rather than adding a flat percentage: expected defect arrival times re-test cost per defect, divided by the measured fraction of days the build and environment are usable. Keep the buffer named, visible and tied to stated assumptions.

solid answer

~50 s

A first-pass estimate answers only 'how long to run everything once', and the cycle is rarely dominated by the first pass. Two drivers add the rest. Defects arrive, get fixed and force re-runs - the confirmation of the fix plus whatever regression slice the fix invalidated - and each round costs calendar as well as effort because it waits on fix turnaround. Separately, a share of planned days is lost to unusable builds and environment outages. Size both from measured history: defects per unit of scope, days per re-test round, and the fraction of planned days actually usable last cycle. Then keep the buffer as a **named line item** with its assumption attached, rather than padding each task quietly. A visible buffer with a condition on it survives challenge; hidden padding gets negotiated away without anyone knowing what they removed.

code

pseudocode · 11 lines
pseudocode
first_pass_days   = 36.05
defect_rate       = 0.9     // defects per first-pass day, measured
retest_per_defect = 0.35    // days: confirmation + invalidated slice
usable_fraction   = 3.8 / 5 // measured usable days per planned day

defects_expected = first_pass_days * defect_rate
retest_days      = defects_expected * retest_per_defect
total_days       = (first_pass_days + retest_days) / usable_fraction

print(retest_days)   // 11.36
print(total_days)    // 62.4

go deeper

for a junior

Be ready to say that a failing check is not the end of the work: the fix has to be confirmed and nearby coverage re-run, and some planned days are lost to builds that will not deploy.

for a middle

Explain the mechanics of both drivers - defect arrival times re-test cost per defect, and the measured fraction of planned days that are usable - and why availability divides the estimate rather than adding to it.

for a senior

Demonstrate that you have run this for real: where your defect and availability numbers came from, how you handled a fix that invalidated a large slice of already-passing coverage, and when you re-issued the estimate mid-cycle.

for a principal

Own the contingency policy across teams: how allowances are named and reviewed, what evidence releases them, and how to stop a visible buffer being treated as slack to be reclaimed at the first schedule squeeze.

## The first pass is not the job Sizing the first pass - execute every planned check once - is the easy half. What consumes a cycle is what happens after the first failure, and estimates that stop at the first pass are the most common way a test schedule goes wrong. Two distinct mechanisms add the rest, and they need separate treatment because they respond to different levers. ## Driver one: fix-and-re-test churn A defect costs far more than the minutes spent reporting it. Each one carries investigation and isolation before it can be filed, a wait on fix turnaround that consumes calendar rather than effort, confirmation that the fix works, and re-execution of whatever coverage the fix invalidated. That last part is where estimates break. Consider a grant-application review queue whose 340-case regression pack was mostly green when a defect surfaced in the ordering assumption: applications were being ordered by score where the queue's rule was submission timestamp within a funding round. The fix changed the ordering everywhere, which meant 61 already-passing cases across pagination, reviewer assignment and the summary counts were no longer trustworthy and had to run again. Three fix rounds later, the pack had been partly re-run three times. None of that appeared in a first-pass estimate. The sizing inputs are measurable: **defects per unit of scope** from comparable past cycles, **median fix turnaround** (that queue's previous cycle measured 1.7 days), and **re-test cost per defect** as a fraction of first-pass effort - the confirmation plus the invalidated slice. Multiply, do not guess. And treat the tail honestly: the distribution is skewed, because a single defect in a shared assumption can invalidate a large fraction of the pack while twenty cosmetic ones invalidate nothing. ## Driver two: availability The second driver is that planned days are not usable days. Builds fail to deploy, an environment goes down, a dependency's stub is stale, the test data has not been refreshed. Track the fraction of planned days on which the team could actually test - if the last cycle gave 3.8 usable days out of every 5, that 0.76 is a measured property of the environment, not pessimism. Note the arithmetic difference between the two drivers. Re-test churn *adds work*; availability *shrinks the working week*. Availability therefore divides rather than adds, and that distinction matters when the number is challenged: ``` first_pass_days = 36.05 defects_expected = 32.4 // from measured arrival rate retest_days = 32.4 * 0.35 // confirm + invalidated slice per defect usable_fraction = 3.8 / 5 // measured last cycle total = (36.05 + 11.35) / 0.76 // = 62.4 elapsed tester-days ``` ## Named buffer, not hidden padding The instinct is to pad each task by a comfortable margin. Resist it, for three reasons. Padding compounds invisibly, so nobody - including you - can say how much of the estimate is contingency. It is unfalsifiable, so it cannot be defended when challenged; a manager who cuts 20% has no idea what they removed. And it destroys the feedback loop, because the actuals come back against inflated tasks and your rates never improve. Instead make the buffer a line item with a **stated trigger**: "re-test allowance: 11.4 days, assuming defect arrival matches the last two cycles" and "availability allowance: assumes at least 3.8 usable days per five-day week". Now the buffer is a claim about the world. If the environment turns out to be up every day, the allowance is released and you say so. If the build is blocked for a week, the assumption has failed publicly and the estimate changes with a reason attached rather than a quiet overrun. ## Quote a range, carry the assumptions The output shape follows from all of this: a range, with the assumptions that would move it named beside it. The optimistic end is the world where defect arrival is low and the environment holds; the pessimistic end is the world where a shared assumption breaks late and re-invalidates a large slice. Listing the assumptions is what makes the range actionable - each one is something somebody can go and change. Better environment availability, faster fix turnaround, or earlier delivery of the riskiest area all narrow it, and that is a far more productive conversation than arguing about the number itself. ## Re-estimating in flight A buffer is not a one-time calculation. Once the cycle starts, real defect arrival and real availability replace the assumed ones, and the estimate should be re-issued when they diverge materially. Two things are worth stating up front so this is expected rather than alarming: which measurements you will watch, and what magnitude of divergence triggers a re-issue. An estimate that is never revised in the face of contradicting evidence has stopped being an estimate and become a promise.

  • Why is a flat 20% contingency worse than an equally sized named buffer?
    Because it cannot be argued with or released. A flat percentage says nothing about what it is covering, so it is cut arbitrarily under pressure and nobody knows what risk they just accepted. A named allowance - re-test churn at the last two cycles' defect rate, availability at 3.8 usable days in five - is a falsifiable claim: it can be checked against reality, released when the assumption over-delivers, and defended when it does not.
  • How does one defect in a shared assumption differ from twenty isolated defects, for estimation purposes?
    Isolated defects invalidate only their own cases, so their re-test cost is roughly linear and predictable. A defect in a shared assumption - an ordering rule, a currency conversion, a permission model - invalidates every case that relied on it, so a single fix can force a large re-run and reset progress the team thought it had banked. That skew is why re-test allowance should carry a pessimistic end rather than a single average.
  • The environment turns out to be available every day. What do you do with the allowance?
    Release it explicitly and say so, ideally in the same report where you would have flagged an overrun. That is what earns the buffer credibility next time: it visibly behaves as a conditional allowance rather than a permanent tax. Handing back time you did not need is the strongest argument you have for being given it again.

saying these in an interview costs you the question

  • Adding a flat percentage with no stated driver
  • Padding each task quietly instead of one named allowance
  • Estimating first-pass execution only, ignoring re-test rounds
  • Assuming every planned day is a usable testing day
  • Never re-issuing the estimate when arrival or availability diverges
  • Treating re-test cost as linear in defect count

context