A generative product feature's test case met its criteria in four of five repeats - what does a single pass mark hide?
answer
- a boolean where a rate belongs
- which repeat missed, and on what
- read it against this case's own history
- passed under a rule is not passed cleanly
- same count, same rule, same inputs
basics
~20 sA single pass mark collapses a rate into a certainty. It hides which repeat missed and on which criterion, how badly it missed, whether four-in-five is normal for this case, and whether the misses cluster on one input.
solid answer
~40 sReport the outcome as what it is - four of five repeats met the criteria under a stated rule - rather than as *pass*. A single mark hides four things a reader needs: which repeat missed and which criterion it failed, whether the miss was marginal or total, whether four-in-five is this case's usual level or a fall from five-in-five last month, and whether the misses concentrate on one input. It also makes builds incomparable, because two results both marked pass can hide different misses. So the outcome line should carry the repeat count, the number that met the criteria, the rule applied, and the criterion each miss failed - and the status should stay distinguishable from a clean sweep, so nobody reads a rate as a guarantee.
code
json · 11 lines{
"caseId": "summarise-long-thread",
"rule": "at least 4 of 5 repeats meet the criteria",
"repeats": 5,
"met": 4,
"status": "passed-under-rate-rule",
"misses": [
{ "repeat": 3, "failedCriterion": "stated-deadline-present", "severity": "omission" }
],
"recentBuilds": ["5 of 5", "5 of 5", "4 of 5", "4 of 5"]
}go deeper
Know that the result for a feature that varies is a count of repeats rather than a yes or a no, and that the count is worth writing down somewhere a person will read it.
Explain what a single mark destroys: which repeat missed, which criterion it failed, how badly, and whether that level is normal for this case.
Design the outcome line so a reader can act on it - repeats, number met, rule applied, criterion missed - and keep it comparable with the same case's earlier builds.
Own what the team publishes for varying cases and what a tolerated miss must look like on the record, so that it is a decision someone made rather than a green mark nobody read.
## A boolean where a rate belongs A case that runs five repeats against a product feature whose answer comes from a generative step, and meets its criteria in four of them, has produced a measurement: *four in five*. Reporting it as `pass` throws the measurement away and substitutes a certainty the run never established. The reader of the results sees the same green mark beside this case as beside a case that asserted an exact value and matched it - and those two results license entirely different conclusions. This is not pedantry. A rate that is not written down cannot be compared, and a rate that is never compared cannot be seen to fall. Most of the value of repeating the case is destroyed at the instant the repeats are collapsed into one mark. ## What the single mark hides - **Which repeat missed, and on which criterion.** A miss on "must mention the deadline the customer stated" and a miss on "must not invent an order number" are different severities with different owners. - **How badly it missed.** An answer that omitted one required element and an answer that was confidently wrong both count as exactly one miss under the rule. - **Whether this is normal for the case.** Four-in-five is unremarkable if the case has run at four-in-five since it was written, and alarming if it ran at five-in-five for a month. - **Whether the misses cluster.** One input missing repeatedly while its neighbours stay clean is a different finding from misses scattered evenly across inputs. - **That a miss happened at all.** This is the one that bites later, when somebody reads the results while looking into a user's complaint about exactly this behaviour and concludes the suite was clean. ## What the outcome should carry instead The unit of reporting for a case like this is a **line, not a mark**. Four fields make it readable, and all four are already in the run's hands: 1. the repeat count and the number of attempts that met the criteria, written as a figure such as `4 of 5`; 2. the rule that was applied, printed next to the figure rather than looked up somewhere else; 3. the criterion each miss failed; 4. a status distinguishable from a clean sweep, so that *passed under a rate rule* and *passed every repeat* are not the same word. | Reported as | What a reader can do with it | | --- | --- | | `pass` | nothing beyond "no action" | | `pass - 4 of 5` | see that a miss happened, and ask about it | | `pass under a 4-of-5 rule, 4 of 5, miss on required-field` | judge the miss, set it beside previous builds, decide | The fourth field is a naming decision rather than a technical one. If the results view offers only pass and fail, add a distinct word for *met the rate rule with observed misses*. Teams that do not add one, in practice, stop reading their varying cases at all. ## Comparability is the point Two figures are comparable only when they were produced the same way. `4 of 5` this week means something beside `5 of 5` last week only if both used the same repeat count, the same rule, the same inputs and the same criteria. Change the repeat count and the two are different quantities. Change the criteria and the rate is now measuring a different promise. So the reported line should carry enough of its own definition to be read next to its own history: - Keep the per-repeat outcomes for at least as long as the trend you want to be able to see. - Record the rule alongside the result, so that a later change to the rule appears as a change rather than as an unexplained step in the figures. - When the criteria change, treat the earlier figures as a separate series - the number moved because the question moved, not because the feature did. There is a cheaper version of all of this that is still worth having when the results view cannot be changed: put the figure in the case's own name or its failure text, so that `4 of 5` reaches the reader even when the surrounding tooling only knows how to draw a mark. What separates a strong answer here is the refusal to let the reporting layer round a probabilistic result into a deterministic one. The suite already knows the rate; the only reason it disappears is that nobody chose a place to put it.
- Why keep a status distinct from an ordinary pass for a case that met a rate rule?Because the two license different conclusions. A clean sweep says nothing was observed to miss; a rate-rule pass says a miss was observed and tolerated. Someone scanning results needs to see that difference without opening the run, because the two outcomes justify different next steps - and because a tolerated miss should be a decision on record rather than an unread green mark.
- What makes two builds' results for the same varying case comparable?The same repeat count, the same rule, the same inputs, the same criteria, and per-repeat outcomes kept for both. A rate measured over three repeats and one measured over twenty are different quantities, and a figure with no earlier figure beside it cannot be read as stable or falling. When the criteria change, treat the earlier numbers as a separate series.
- The results view can only draw pass or fail. What do you do?Put the figure where the view will carry it anyway - in the case's reported name or in the text attached to the result - so that `4 of 5` reaches the reader even though the mark itself is binary. It is a worse answer than a proper outcome line, but it preserves the one number that makes the result readable and comparable.
saying these in an interview costs you the question
- Reports a case governed by a rate rule as a plain green mark
- Keeps only the final pass or fail, discarding per-repeat outcomes
- Reads this build's figure with no earlier figure beside it
- Treats a marginal miss and a confidently wrong answer as the same outcome
- Lets the reported mark imply that every repeat succeeded