skip to content

How do you report a test cycle so it conveys confidence rather than a pass-rate percentage?

level: principalimportance: should knowfreq 39%

answer

  1. Ratio over whatever happened to run
  2. Every case weighs the same to it
  3. Three categories, not two
  4. Confidence needs its basis attached
  5. A target stops being a measure

basics

~20 s

Replace the ratio with statements a reader can act on: what was exercised and what was not, what the cycle learned about each area, where the remaining uncertainty sits, and how much of that is measurement rather than product. Keep the percentage as detail, never as the headline.

solid answer

~60 s

A pass-rate percentage is a ratio over whatever happened to run, weights every case equally regardless of consequence, and moves when cases are added, split or retired — so it can improve while the product gets worse. Report instead in claims: which areas were exercised and to what depth, which were not exercised at all and why, what the cycle found and what it therefore knows, and where the residual uncertainty sits. Distinguish sharply between *tested and behaved*, *tested and broken*, and *not tested* — that third category is the one percentages erase. Qualify confidence with its basis, not with adjectives: confidence in an area rests on how much of it ran, on how recently the code under it changed, and on whether the environment resembled production. Keep the numbers in the detail for people who want them, and lead with the sentences. And watch for the dysfunction: once a pass rate becomes a target, it stops measuring, because the cheapest way to raise it is to stop running the cases that fail.

go deeper

for a junior

Know that a percentage of passing cases is not the same as a statement about the product, and that a report should always say how much of the plan never ran. Practise writing one sentence about what was not covered.

for a middle

Explain the structural weaknesses — the denominator is whatever ran, every case weighs the same, the unit is unstable — and be able to replace the headline with two or three concrete claims about what was exercised and what was not.

for a senior

Demonstrate that you write reports in claims with evidence, split product uncertainty from measurement uncertainty, and concentrate the reader's attention where the change landed. Be ready to defend a report that contains no percentage in its opening lines.

for a principal

Own the convention across teams: which figures are permitted as headlines, that untested scope is always named explicitly, and that no pass rate is used as a target. Expect to explain the measurement dysfunction to an executive who wants one comparable number per team.

### Why the percentage fails as a headline A pass rate is passed divided by executed. Four properties make it a poor summary and all four are structural, not fixable by more decimal places. **The denominator is whatever ran.** Everything blocked or never reached is outside the ratio entirely. A cycle that executed 263 of 418 planned cases can report 95.4 percent while a third of the plan is untouched — and blocks cluster, so the untouched part is not a random sample of the product. **Every case weighs the same.** One failing case on the money-handling path and one failing case on a preferences screen move the number identically. The metric has no concept of consequence, which is the only thing the reader cares about. **The unit is unstable.** Split one case into four and the ratio changes with no change in the product. Retire twelve stale cases and it changes again. Any measure whose unit the reporting team controls invites drift, honest or otherwise. **It answers a question nobody asked.** Stakeholders want to know what might break, for whom, and how they would find out. A ratio answers none of those. ### What to report instead Write claims, and attach evidence to each. 1. **What was exercised, and to what depth.** "The checkout path was exercised end to end across three payment methods and two currencies" is a claim someone can challenge. A percentage is not. 2. **What was not exercised, and why.** The most valuable paragraph in any cycle report. Untested is not a confession, it is a finding — provided it is stated rather than left to be inferred from a number. 3. **What the cycle learned.** Not just counts of defects but the shape of them: where they clustered, whether they were the kind expected from the change, whether an area produced surprises disproportionate to its size. 4. **Where uncertainty remains, and of what kind.** Separate *product* uncertainty (we exercised it and it behaves, but only under one configuration) from *measurement* uncertainty (the environment differed from production, or results came from a suite whose own reliability is questionable). They call for different responses. 5. **Confidence with its basis attached.** "Higher confidence in catalogue browse than in receipt generation, because browse ran fully on a production-like data set and receipt generation ran on one locale only." Confidence that names its evidence can be argued with; confidence expressed as high, medium or low cannot. ### A worked report line An online bookstore checkout cycle: 263 of 418 planned cases executed, 251 passing. The percentage headline would be 95.4 percent, and it would be true and useless. The useful version: cart, address and catalogue browse were exercised fully and behave. Payment was exercised in one locale only; a locale-dependent format defect in the order total blocked 31 cases behind it for three days, so the tax, currency-formatting and receipt paths carry no observation this cycle. The performance case written against the expected 1,200-request-per-minute peak did not run, so nothing is known about behaviour at load. Uncertainty is concentrated in exactly the area the release changed, and it is measurement-side as much as product-side, because the block was in the test data rather than in the product. That paragraph tells a reader what they can rely on, what they cannot, and where a decision is now required. It also contains no percentage in its first sentence. ### The dysfunction to design against A measure that becomes a target stops being a good measure — the observation usually credited to Goodhart. Make a pass rate the number a team is judged on and the cheapest routes to improving it are all corrosive: quarantine failing cases out of the counted set, split passing cases to dilute failures, retire cases that keep going red, or execute the easy areas first so the ratio looks healthy on reporting day. None of these require anyone to lie, which is exactly why they happen. The defence is to make the untested and blocked columns as prominent as the pass column, and to report trends in what was *covered* alongside what *passed*. ### Keep the numbers, demote them None of this argues for a report without figures. Counts of executed, blocked and not-run against the plan, the age of the oldest block, and measured throughput all belong in the detail — they are what makes the claims checkable. The argument is about which of them leads. A reader who takes only the first three sentences should come away knowing what is uncertain, not knowing a ratio. ### Owning the convention At a lead level the deliverable is not one good report; it is a reporting convention the organisation can rely on, in which "we did not test it" is never expressible in the same phrase as "we tested it and it was fine". Fix that one rule and most of the rest follows.

  • A stakeholder insists on a single number for the cycle. What do you give them?
    A coverage-of-plan figure paired with the count of areas carrying no observation, rather than a pass rate — it at least points at what is unknown. Then say in one sentence which area holds the most uncertainty. If a single number is unavoidable, choose one that gets worse when scope goes untested, because that is the failure mode you want the organisation to feel.
  • How do you separate product uncertainty from measurement uncertainty in a report?
    Ask what would change the statement. If a better environment, more representative data or a more reliable suite would change it, the uncertainty is measurement-side and the remedy is investment in the test system. If only more information about the product would change it, it is product-side and the remedy is more or deeper testing. Conflating them sends effort to the wrong place.
  • What tells you a pass rate has become a target rather than a measure?
    The ratio improves while coverage of the plan does not, cases get retired or split around reporting time, consistently failing cases are quietly moved out of the counted set, and easy areas are scheduled first. Watching the executed-versus-planned trend beside the ratio makes those movements visible.
  • Should a cycle report ever say confidence is high?
    Only with the basis attached — what ran, on what data, in what environment, and how recently the underlying code changed. Bare adjectives cannot be argued with, so they transfer risk to the reader without giving them anything to challenge. A sentence naming the evidence lets a stakeholder disagree, which is the point of reporting at all.

A weather report that says "92 percent of the instruments returned a reading" tells you nothing about whether to take a coat. What you want is which readings came in, which stations were offline, and what that leaves unknown.

saying these in an interview costs you the question

  • Leads with a pass-rate percentage and stops there
  • Reports confidence as high or medium with no evidence attached
  • Lets untested scope be inferred rather than stated
  • Treats every case as equally consequential to the reader
  • Compares pass rates across cycles whose case sets differ
  • Sets a pass-rate target and expects the number to stay meaningful

context