skip to content

When do you estimate test effort from historical throughput rather than expert three-point ranges?

level: middleimportance: should knowfreq 46%

answer

  1. Two honest sources for one number
  2. Measured rates need a comparability assumption
  3. Three numbers instead of one
  4. Weighted mean plus a spread
  5. Estimate silently before anyone speaks

basics

~20 s

Use historical throughput when the coming work resembles work you have already measured on the same product and team. Use expert three-point ranges and consensus rounds when the work is new, the process changed, or no comparable history exists.

solid answer

~50 s

Metrics-based estimation multiplies measured rates - days per scope item at each depth, defect arrival per cycle, days lost to unusable builds - by the scope in front of you. It is only as good as its comparability assumption: same product area, same team, same level of automation, same environment. When that assumption fails, or when the work has no precedent, you fall back on expert judgement, and the discipline there is to ask each estimator for three numbers - optimistic, most likely, pessimistic - rather than one, and to combine them into a weighted mean plus a spread. Consensus methods add a second discipline: estimators produce their numbers independently before anyone speaks, then discuss only the outliers and re-estimate. In practice you mix the two, using history for the recurring parts of a cycle and expert ranges for the genuinely novel slice.

code

pseudocode · 9 lines
pseudocode
optimistic   = 2.0
most_likely  = 3.0
pessimistic  = 12.0

expected = (optimistic + 4 * most_likely + pessimistic) / 6
spread   = (pessimistic - optimistic) / 6

print(expected)   // 4.333
print(spread)     // 1.667

go deeper

for a junior

Be able to name the two sources of an estimate - measured history and expert judgement - and say that three-point ranges ask for optimistic, most likely and pessimistic rather than a single figure.

for a middle

Explain the mechanics: which rates are worth keeping from past cycles, the comparability assumption underneath them, the weighted mean and spread of a three-point estimate, and why consensus rounds collect numbers independently first.

for a senior

Show judgement about which source applies to which slice of a real cycle, how you widened a range when the history was borrowed, and how you fed actuals back to correct a standing bias.

for a principal

Own the calibration loop across teams: what gets recorded, how correction factors are derived and reviewed, and how you keep estimates from becoming commitments that nobody dares revise.

## Two sources for the same number Every test estimate answers the question "where did this number come from", and there are only two honest answers: *we measured something like it* or *people who have done it judged it*. Interviewers ask about the split because candidates who only know one of them either extrapolate history into work that has no precedent, or guess when perfectly good data was sitting in the last three cycles. ## Metrics-based estimation Metrics-based estimation takes rates from completed cycles and applies them to the scope in front of you. The rates worth keeping are: - **Throughput per depth band** - days per scope item for deep, standard and shallow coverage. Over four cycles on a grant-application review queue, a team might see the shallow band run at 22.5 items per tester-day and the deep band at 0.4, and that ratio is far more informative than either number alone. - **Defect arrival per unit of scope**, which drives the re-test tail. - **Fix turnaround**, because a defect that waits three days to be fixed occupies calendar, not just effort. - **Availability** - the fraction of planned days on which the build and environment were actually usable. The strength of these rates is that they already contain the friction nobody remembers to estimate: interruptions, meetings, the half-day lost to a bad deployment. The weakness is the comparability assumption sitting underneath them. History from a different product, a different team, a pre-automation state of the same suite, or a cycle run under different environment conditions is not your history. Say which cycles the rates came from and why the coming work resembles them; that sentence is the estimate's most important assumption. ## Expert-based estimation When there is no comparable history - a new integration, a first cycle on a rewritten area, a technique the team has not used - you are estimating with judgement, and the technique is about extracting judgement honestly rather than pretending it is data. **Three-point estimates.** Ask for optimistic, most likely and pessimistic values instead of one number. A single number invites the estimator to collapse a genuinely uncertain distribution into a point they will then defend. Three numbers expose the uncertainty: a task quoted 2 / 3 / 12 days is telling you something very different from one quoted 2 / 3 / 4, even though both have the same most-likely value. The commonly used weighting from programme-evaluation scheduling practice is (optimistic + 4 x most likely + pessimistic) / 6, with (pessimistic - optimistic) / 6 as a rough spread. Treat that weighting as a convention that behaves reasonably, not as a law - it assumes a particular shape of distribution. **Consensus rounds.** The wideband-Delphi family of methods fixes the social failure of group estimation: whoever speaks first anchors everyone else, and the most senior voice anchors hardest. The procedure is to brief the group on the scope, have each estimator produce a number **independently and silently**, reveal all of them at once, discuss only what the highest and lowest estimators know that the others do not, then re-estimate. The convergence matters less than what the discussion surfaces - the outlier usually knows about a data dependency or an ordering assumption that nobody else had priced in. ## Mixing them, which is what actually happens A real cycle is mostly recurring work with a novel slice in it. Estimate the recurring part from history and the novel part from expert ranges, and label which is which. This has a practical benefit: when the estimate is challenged, you can point at the part that is measured and defend it differently from the part that is judged. It also tells you where to spend your uncertainty reduction - a half-day spike on the novel slice narrows the range far more than another hour arguing about the recurring part. ## Calibration and the honest caveats Estimates are only worth what their feedback loop is worth. Record the estimate, record the actual, and look at the ratio per cycle; a team that is consistently 1.4x over on deep-band work has just discovered a correction factor worth more than any technique. Note also that the evidence base for specific estimation techniques is contested in the literature - claims that one method is reliably more accurate than another rarely survive the difference between organisations. What is not contested is the direction of the bias: unaided estimates tend to be optimistic, and the tail risks - blocked environments, defect churn, an assumption discovered late - fall almost entirely on the pessimistic side. That asymmetry is why three-point ranges and measured availability earn their place regardless of which school you started from.

  • Why do consensus methods ask for silent, independent estimates before any discussion?
    To stop anchoring. The first number spoken - especially by the most senior person present - pulls every later number towards it, and the group then mistakes that convergence for agreement. Collecting estimates independently preserves genuine disagreement, and the disagreement is the valuable part: the outlier normally knows about a dependency or an assumption the others had not priced. Discussion happens after the reveal, aimed at the extremes.
  • Your throughput data comes from a team that has since lost two of its four testers. Can you still use it?
    Only with the change declared. Per-tester-day rates are more portable than per-cycle totals, but they still embed team knowledge, and replacements ramp slowly. Use the old rates as a starting point, widen the range, and state the assumption explicitly - that new testers reach a stated fraction of the old rate by a stated point. Then measure early and re-issue rather than defending the original number.

saying these in an interview costs you the question

  • Applying rates from a different product or team without saying so
  • Collapsing a three-point range to the most likely value immediately
  • Letting the most senior person estimate first in a group
  • Treating the weighted three-point formula as a proven law
  • Never comparing past estimates against actuals

context