skip to content

Test Effort Estimation

Sizing a testing effort from risk and scope rather than case counts, with buffer for defect churn and re-test rounds. Interviewers ask because the date is usually fixed before the estimate exists.

on this pageshow

questions

4

Why is a test case count a weak basis for sizing a test effort?

level: juniorimportance: must knowfreq 58%

answer

  1. Not every case costs the same
  2. The expensive tail dominates the schedule
  3. Splitting cases inflates the count
  4. Size from scope and depth per band
  5. Convert using throughput from past cycles

basics

~20 s

Because cases vary hugely in cost: a one-step check and a case needing days of seeded data both count as one. Size instead from scope and the depth each risk band needs, then convert with measured throughput.

solid answer

~40 s

A case count treats every case as the same size, and real cases are nothing like the same size. One case may assert a single rejection message; another may need a seeded dataset, several roles and a multi-day setup before it can run at all. The count is also gameable - splitting one case into five inflates the estimate without adding an hour of work - and it says nothing about the activities that dominate a cycle: analysis, data setup, defect investigation and re-test rounds. A defensible estimate starts from scope (what changed, which interfaces and configurations are in play), assigns depth per risk band, and converts that to effort using throughput measured on comparable past work. Case counts are still useful as a coverage inventory; they are just not a unit of effort.

code

pseudocode · 11 lines
pseudocode
bands = [
  { name: "deep",     items:  6, days_per_item: 2.5  },
  { name: "standard", items: 19, days_per_item: 0.75 },
  { name: "shallow",  items: 34, days_per_item: 0.2  }
]

first_pass_days = 0
for b in bands:
    first_pass_days = first_pass_days + b.items * b.days_per_item

print(first_pass_days)   // 36.05

go deeper

for a junior

Be ready to say plainly that cases differ in cost and give one cheap and one expensive example. Knowing that setup and data often outweigh execution is enough at this level.

for a middle

Explain the mechanics: what a count hides, why it can be inflated by splitting cases, and how you convert scope and depth into days using throughput measured on comparable work.

for a senior

Show you have estimated a real cycle. Talk about where your rates came from, how you handled heterogeneous work, and what you did when the pack contained a handful of cases that dominated the schedule.

for a principal

Own the consequence for the organisation: if teams are measured or funded on case counts, they will write more cases. Be ready to describe a sizing basis that cannot be gamed and still survives a budget conversation.

## The unit problem Any estimate is a count multiplied by a rate. That only works when the thing being counted has a roughly stable cost. A test case does not. Case cost is heavily skewed: most cases are cheap, a few are enormous, and the expensive tail is where the schedule actually goes. Take a grant-application review queue. One case asserts that a rejected application shows the correct rejection reason - about four minutes to run, less to re-run. Another exercises the queue's ordering rule across twelve reviewer accounts and three overlapping funding rounds, and cannot start until someone has seeded applications with staggered submission timestamps, scores and reviewer assignments; call it two and a half days including the data build. Both are one line in the pack. Multiply a count of 340 cases by any average and you have averaged those two together, which produces a number that is wrong for every case in the pack. ## What the count hides Four costs disappear inside a case count: - **Setup and data.** Constructing and refreshing test data is often larger than execution. Cases sharing one dataset are cheap after the first; a case needing its own fixture is not. - **Oracle difficulty.** How hard it is to decide whether the observed behaviour is correct. A case with an exact expected value is quick. A case whose oracle is "the ordering is defensible given these scores" needs a person to reason. - **Environment and configuration.** The same case run against four configurations is four executions, one line. - **Everything that is not execution.** Test analysis and design, defect investigation and reporting, confirmation and regression re-runs, and the meetings a cycle drags along. ## The count is gameable Case counts are an authored artefact, not a measurement of the system. The same coverage can be written as 340 cases or 90 cases with richer steps. If the estimate is a function of the count, then the estimate is a function of a writing style, and anyone who wants a bigger number can produce one by splitting cases. That is a strong signal the unit is wrong: a good sizing basis should be hard to inflate without doing more work. ## What to size from instead Size from **scope** and **depth**, then convert with **measured throughput**: 1. **Enumerate scope items**, not cases: the changed features, the interfaces touched, the configurations and data variants that must be covered, and the non-functional concerns in play. 2. **Attach the depth each item gets.** The ranking that decides which item deserves deep coverage is a risk exercise in its own right; estimation consumes its output. What estimation adds is a cost per depth level - a deep item costs multiples of a shallow one. 3. **Convert with history.** Days per item at each depth, taken from comparable past cycles on the same product and team, is far better calibrated than minutes per case. If no such history exists, that absence is itself an assumption to declare. 4. **Add the non-execution activities explicitly**, so they are visible and cannot be quietly negotiated away. ``` deep : 6 items x 2.5 days = 15.0 standard : 19 items x 0.75 days = 14.25 shallow : 34 items x 0.2 days = 6.8 first-pass execution = 36.05 days ``` ## Where case counts are still right This is not an argument that case counts are useless. They are a good **inventory**: which requirements have cases against them, what exists to be maintained, what the regression pack contains. They are also acceptable as a *progress* unit inside one homogeneous slice - counting executed against planned within a set of similar cases tells you something real. The failure is specifically using them as the **sizing unit across heterogeneous work**. ## Say the honest thing about rates There is no credible industry constant for minutes per test case, and any figure quoted as one should be challenged: it depends on the product, the level being run, the automation state and the data. Your own recent cycles are the only rate worth using, and even they need the assumption stated - that the coming work resembles the work they came from. If the estimate is quoted as a range with its assumptions attached, the case count can appear in it as a description of the pack without ever being the thing that generated the number.

  • If case counts are a poor sizing unit, what is the count still good for?
    As an inventory and a traceability artefact: which requirements have cases against them, what the regression pack holds, what has to be maintained. It also works as a progress unit inside one homogeneous slice, where executed-against-planned means something because the cases really are similar in cost. It stops being trustworthy the moment the slice mixes a four-minute assertion with a multi-day data build.
  • You have no historical throughput for this product. What do you do?
    Say so, and make the absence an explicit assumption in the estimate. Then borrow the nearest comparable rate you do have, mark it as borrowed, and quote a wider range because of it. Where possible, timebox a small slice first, measure the real rate on it, and re-issue the estimate - a measured half-day beats a confident guess. Never present a borrowed rate as if it were this team's history.
  • How does a large automated pack change the sizing?
    It moves cost from execution to maintenance and triage. Running an automated pack is cheap and roughly fixed regardless of size, so the count matters even less; what matters is authoring cost for new coverage, the time to investigate failures, and the churn when the interface changes. Size automated work as build-and-maintain effort, and size the manual and exploratory work separately - the two have completely different cost curves.

Estimating a move by counting boxes: a box of pillows and a box of books are both one box, and the piano is not a box at all.

saying these in an interview costs you the question

  • Multiplying total cases by an average minutes-per-case figure
  • Quoting an industry-standard minutes-per-case constant as fact
  • Treating a bigger case count as proof of more coverage
  • Estimating execution only, ignoring analysis, data and re-test
  • Assuming every case is independent of setup and data cost

context

open as a page

When do you estimate test effort from historical throughput rather than expert three-point ranges?

level: middleimportance: should knowfreq 46%

basics

~20 s

Use historical throughput when the coming work resembles work you have already measured on the same product and team. Use expert three-point ranges and consensus rounds when the work is new, the process changed, or no comparable history exists.

open as a page

How do you buffer a test estimate for fix-and-re-test cycles and days lost to blocked builds?

level: seniorimportance: should knowfreq 51%

basics

~20 s

Model the drivers rather than adding a flat percentage: expected defect arrival times re-test cost per defect, divided by the measured fraction of days the build and environment are usable. Keep the buffer named, visible and tied to stated assumptions.

open as a page

What does test point analysis size a test effort from, and why is it rarely used?

level: middleimportance: nice to knowfreq 14%

basics

~20 s

Test point analysis is a formula-based method that sizes testing from a counted functional size of the system, adjusted for quality characteristics, risk and environment factors. It is rare because almost nobody maintains a functional size count to feed it.

open as a page