For a feature whose answers come from a generative step, why does a test suite pass rate misstate readiness, and what replaces it?
answer
- A percentage of what, exactly
- Same cases, a different number
- Two surfaces averaged into one figure
- Residual error rate per user action
- Sample size, judging rule, conditions
basics
~20 sA pass rate describes one judged sample of varying output, not the feature itself: re-judge the same cases and the number moves. Report a measured residual error rate per user action, with its sample size, judging rule and split by consequence.
solid answer
~50 sA pass rate is a property of the case set and the judging rule, not of the feature. Because the generative step inside the product can answer the same input differently on two executions of the suite, the same cases yield a different percentage each time, and the figure carries no spread. It also blends two surfaces: the deterministic wrapper around the feature — routing, permissions, formatting, the fallback path — which should be at 100%, and the varying answers, which never will be. The honest replacement is a **residual error rate stated per user action**, over a stated number of judged attempts, with the rule that decided right from wrong written down and the errors split by what each costs. State the conditions it was measured under, because it is only valid under those.
code
yaml · 13 linesrelease_statement:
user_action: summarise an uploaded document
judged_attempts: 240
wrong_answers: 14
residual_error_rate: 5.8% (repeats ranged 4.2% to 7.5%)
judging_rule: reject if any stated fact is absent from the source
measured_on: input mix sampled from last month's uploads
deterministic_surface: pass (permissions, routing, fallback, formatting)
by_consequence:
discarded_by_user: 4.6%
corrected_before_sending: 1.0%
reached_a_customer: 0.2%
irreversible_action: 0% (gated behind confirmation)go deeper
Be ready to say what a pass rate is and why a feature that can answer the same question differently makes that percentage unstable. Knowing the number moves without the code changing is enough at this level.
Explain the mechanics: sample size and spread, the rule that decides each verdict, and the fact that one percentage averages the fixed software around the feature with the varying answers themselves.
Show that you would report a residual error rate per user action with its sample and its conditions, split by what each error costs, and that you would refuse to re-judge until the figure improves.
Own what the organisation may conclude from a number: which figures are allowed into a release document, what must always travel with them, and how you stop a rate measured under one judging rule being compared with another.
## What a pass rate actually measures A pass rate is a ratio: cases judged to have passed, over cases executed in one run of the suite. It is a property of three things — the set of cases somebody chose, the rule that decided each verdict, and the execution that happened to occur. On a feature whose logic is fixed, those three are stable enough that the ratio behaves like a property of the product: the same suite over the same build returns the same number, and a change in the number means a change in the build. A feature whose answers come from a generative step breaks every part of that. The same input can produce a different answer on two executions of the suite, so the ratio moves while the build stands still. Each verdict is itself a judgement about output that has more than one acceptable form, so the rule that judges it is a component of the measurement rather than an obvious fact. And the case set is a tiny sample of an effectively unbounded input space, chosen by people who could only think of the inputs they thought of. ## The three things the percentage hides 1. **Variance.** A percentage with no sample size and no spread invites the reader to treat 94% and 96% as different, when re-judging the same cases might produce either. Whoever reads the recommendation cannot tell a real improvement from a re-roll. 2. **Two surfaces averaged together.** Around the generative step sits ordinary deterministic software: routing, permissions, input validation, volume limiting, formatting, the fallback path, the audit record. That surface should be at 100%, and any failure in it is a plain defect. Averaging it with the varying answers produces a number that describes neither — 95% might mean a perfect wrapper and 90% correct answers, or a broken permission check and 99% correct answers. Those are opposite situations and the aggregate cannot separate them. 3. **Cost.** Counting failures treats a suggestion the user discards and a wrong figure copied into a customer's invoice as one unit each. ## What the recommendation carries instead The replacement is not a better percentage. It is a statement of **residual risk** — what will still go wrong after release, in units the person deciding can act on. | Instead of | State | |---|---| | "95% of cases passed" | "roughly 1 in 20 attempts at this action returns an answer a reviewer would reject" | | a bare ratio | the number of judged attempts behind it, and the spread across repeats | | an unstated judging rule | the rule that decided right from wrong, and who applied it | | one aggregate | a split by what a wrong answer costs | | a snapshot | the conditions it holds under: the input mix, the configuration, the material it draws on | | silence about the wrapper | the deterministic surface reported separately, as pass or fail | Two of those rows deserve emphasis. **Per user action** matters because nobody experiences cases; people experience attempts at a task, and a rate per attempt is the unit that translates into "how often will someone be let down". **Conditions** matter because the figure was produced under an input mix you chose; if real traffic skews towards inputs you under-sampled, the released feature will not reproduce your number and nobody will understand why. ## Reporting it so it survives contact with a decision-maker - Give the rate as a range, not a point, and say how many attempts produced it. If you judged 200 attempts and 12 were wrong, say so, and the reader can see that 6% could reasonably have been 4% or 9%. - Report the deterministic wrapper separately and as a binary. A failing permission check is not "part of the 5%" — it blocks. - Say what you did not measure. Inputs you know exist but did not sample are part of the residual risk even though they contribute nothing to the number. - Keep the judging rule attached to the number permanently. A rate measured under a strict rule and one measured under a lenient rule are not comparable, and six weeks later nobody remembers which was used. - Resist re-judging until the figure looks better. Selecting the best of several executions is the same error as reporting the best of several samples, and it is invisible in the finished document. The honest one-line version is not "the suite passed". It is: under this input mix and this judging rule, about this share of attempts returns something wrong; here is what those wrong answers cost, here is what happens when one occurs, and here is the signal that will tell us we were optimistic.
- Two executions of the same suite give 94% and 97%. What do you report?Both, with the number of attempts each one judged. Report a range rather than a point, and say that the difference sits inside the spread you observed rather than reflecting a change in the build. If that spread is wide enough to affect the decision, judge more attempts before reporting instead of choosing an execution.
- Why report the deterministic software around the generative step separately?Because it is held to a different standard. Routing, permissions, validation, formatting and the fallback path are ordinary software and must be correct every time, so their result is binary. Folding them into an aggregate percentage lets a genuine blocking defect hide inside an error rate the reader has already agreed to tolerate.
Reporting a pass rate here is like reporting one hand of cards as the strength of the deck: deal again and the number changes, though nothing about the deck did.
saying these in an interview costs you the question
- Reports a single pass percentage with no sample size
- Treats a better figure on re-judging as a real improvement
- Averages wrapper failures together with wrong answers
- Claims the feature is correct because the suite went green
- Omits the rule that decided right from wrong