skip to content

What does test point analysis size a test effort from, and why is it rarely used?

level: middleimportance: nice to knowfreq 14%

answer

  1. A third family beyond history and judgement
  2. A formula, not a judgement round
  3. Input is a counted functional size
  4. Weighted for quality characteristics and risk
  5. Rare because the input is not maintained

basics

~20 s

Test point analysis is a formula-based method that sizes testing from a counted functional size of the system, adjusted for quality characteristics, risk and environment factors. It is rare because almost nobody maintains a functional size count to feed it.

solid answer

~50 s

Test point analysis sits in a third family alongside measured throughput and expert judgement: parametric sizing, where a formula converts a counted property of the system into test effort. Its input is a functional size - a function-point-style count of what the system does - which is then weighted for how thoroughly each function must be tested, the quality characteristics in scope, and the risk attached, with further factors for the environment and team productivity and a fixed proportional overhead for planning and control. The attraction is that it is repeatable and independent of who is estimating. The reasons it is rare are practical: it needs a functional size count almost no team maintains, published variants disagree on the weights, and the productivity factor has to be calibrated locally anyway - at which point measuring your own throughput directly is cheaper and answers the same question.

go deeper

for a junior

You will not be marked down for not knowing this. It is enough to recognise that formula-based sizing methods exist alongside measured history and expert judgement.

for a middle

Be able to say what such a method takes as input - a counted functional size rather than code volume - and name one honest reason it is rarely used in practice.

for a senior

Show you can place it: repeatable and early, but dependent on an input few teams maintain and on a locally calibrated productivity factor that makes direct measurement the cheaper path.

for a principal

Own the governance angle: when an organisation wants a defensible, comparable number across teams, say what a parametric method really buys and what it costs to keep its inputs alive.

## A third family of estimate Most test estimating in practice draws on two sources: rates measured on comparable past work, and the judgement of people who have done it. A third family exists - **parametric sizing**, where a formula converts a counted property of the system into an effort figure. Test point analysis is the best-known member on the testing side. ## What it counts The input is a **functional size**: a count of what the system does, in the function-point tradition, rather than how large its code is. Function points count things like the inputs, outputs, enquiries and data stores a system offers, weighted by complexity, so that the size is a property of the delivered functionality and is independent of language and implementation. Test point analysis takes that size and adjusts it along several axes: - **How much testing each function warrants** - the importance of the function, its usage intensity, its interfacing with other functions, and its structural complexity. - **Which quality characteristics are in scope** - testing for performance or security is not the same amount of work as testing functionality alone, so the characteristics selected scale the size. - **Environment and productivity factors** - the maturity of the test process, availability of tooling and test data, the team's experience, all of which convert an abstract size into hours for *this* organisation. - **A proportional allowance for planning and control**, added on top of execution-shaped work, which typically scales with team size and management demands. The published variants differ in the exact factors and weights, so treat the shape as the point rather than any particular number, and be sceptical of anyone quoting a definitive multiplier. ## Why it appeals Three genuine attractions. It is **repeatable**: the same system and the same weights produce the same estimate regardless of who runs it, which removes the personality effects that dog expert judgement. It is **auditable**: every part of the number traces to a counted input and a stated factor, so a challenge lands on a specific factor rather than on someone's confidence. And it is **early**: a functional size can be counted from a specification before any code exists, which is exactly when the date is usually being fixed. ## Why it is rarely used The reasons are practical rather than theoretical. **The input does not exist.** Parametric test sizing depends on a functional size count, and maintaining one is a discipline in its own right, historically done by trained counters against reasonably complete specifications. Teams working from incremental, changing scope generally have nothing to count. **The calibration is local anyway.** The productivity factor that converts abstract test points into hours has to be derived from your own completed work. Once you are measuring your own completed work, you are one step from simply using measured throughput, which needs no counting apparatus and answers the same question with fewer assumptions. **The weights are contested.** Different published versions of the method assign different values, and there is no strong, widely-accepted evidence that any particular weighting predicts effort better than a well-calibrated local rate. A formula's precision can flatter an estimate whose underlying uncertainty is unchanged - the number reads as objective because it came out of arithmetic, not because it is better founded. **It fits a process shape that is less common now.** The method assumes a specification substantial enough to count and a testing effort planned as a phase. Where scope arrives incrementally and testing runs continuously, both assumptions weaken. ## What to say in an interview Knowing this method is a differentiator, not a gate, and nobody's offer turns on it. The useful answer names the family it belongs to, states its input honestly, and then says what you would actually do: for a team with history, measured throughput per depth band; for genuinely novel work, expert three-point ranges with a consensus round. Parametric sizing is worth knowing about mainly as a reminder that a formula-derived figure carries exactly as much uncertainty as its inputs, however precise the output looks.

  • What is the appeal of a parametric estimate over an expert one, in a single sentence?
    It is repeatable and auditable: the same inputs give the same number regardless of who produces it, and a challenge lands on a specific counted input or stated factor rather than on someone's confidence. That also makes it available early, before code exists, which is often exactly when a date is being set.
  • Why can a formula-derived estimate be more dangerous than an openly judged one?
    Because the arithmetic makes it look objective. The output carries all the uncertainty of its inputs and its weights, but it arrives as a single precise figure with no visible range, so it invites more confidence than it has earned. An expert three-point range wears its uncertainty on its face; a parametric point estimate hides it behind a decimal place.

Estimating a paint job from the architect's floor-area figure rather than from how long your crew took on the last three houses - precise, portable, and useless if nobody has the floor plan.

saying these in an interview costs you the question

  • Believing a formula removes uncertainty from an estimate
  • Confusing functional size with lines of code
  • Quoting a definitive weighting as if universally agreed
  • Using a parametric method with no local calibration
  • Presenting a parametric figure as a single precise date

context