When a test cycle's pass and fail totals are computed, what is the difference between counting each item's latest attempt and counting its worst attempt?
answer
- same attempts, two reductions
- fail-then-pass: which outcome wins
- one rate climbs, the other only falls
- publish the re-run count beside it
- a rate without its rule is ambiguous
basics
~10 sLatest-attempt counting scores each item by its newest result, so a fail-then-pass item counts as a pass. Worst-attempt counting scores it by its poorest result, so the same item counts as a fail.
solid answer
~40 sBoth rules read the same stored attempts and produce different totals. **Latest-attempt** counting takes each item's newest recorded outcome, which is what almost every cycle board and progress tile does by default; a cycle where half the items failed and were re-run after fixes can therefore read as a clean pass. **Worst-attempt** counting takes the poorest outcome across an item's attempts, so any failure anywhere in the cycle keeps the item red. Latest-attempt answers "does the software work now?"; worst-attempt answers "how much trouble did this cycle actually find?". Neither is wrong, but a pass rate quoted without saying which rule produced it is ambiguous, and the gap between the two numbers is itself the useful signal — it is exactly the volume of re-run traffic the headline figure hides.
go deeper
Know that an item can hold several attempts and that the totals must pick one of them, and be able to say which outcome a fail-then-pass item contributes under each rule.
Explain the reduction itself: latest-attempt lets a rate climb as fixes land, worst-attempt can only fall, and both read exactly the same stored attempt rows.
Demonstrate the operational consequence — an exit threshold or a cross-cycle trend built on an unstated rule is unenforceable, and repeated automated attempts inflate a latest-attempt rate on unchanged code.
Own the reporting convention: mandate that any published rate names its rule and travels with a re-run count, so two cycles quoted at the same number are genuinely comparable.
## Two rules over the same rows Once a re-run appends an attempt instead of overwriting the previous one, an execution holds a small ordered list: fail, fail, pass. Any total the tool prints has to reduce that list to one outcome per item, and there are two reductions in common use. - **Latest attempt** — take the newest recorded outcome. `fail, fail, pass` counts as **pass**. - **Worst attempt** — take the poorest outcome anywhere in the list, on the tool's own severity ordering. `fail, fail, pass` counts as **fail**. A third variant appears occasionally: **first attempt**, which freezes each item at the outcome it recorded the first time it was run. It is rarer, and mostly shows up in exports built to describe what a build did on arrival. ## What each rule is actually asking | | Latest-attempt | Worst-attempt | |---|---|---| | Question answered | does it work now? | what did this cycle find? | | Fail-then-pass item | pass | fail | | Behaviour over the cycle | rate climbs as fixes land | rate is monotonic downward | | Typical home | the board, progress tiles, the default pass rate | audit views, quality retrospectives | | Blind spot | hides re-run traffic entirely | never recovers from an early failure | The asymmetry matters. Under latest-attempt the cycle *heals*: every fix pushes the number up, and a cycle that ends at a hundred percent looks identical whether nothing ever failed or everything failed twice. Under worst-attempt the cycle only ever gets worse, which makes it useless as a progress indicator but honest as a record of trouble found. ## Where the choice bites - **Exit criteria.** A pass-rate threshold written without naming a rule is unenforceable — the same cycle passes and fails depending on which reduction the reader applies. - **Comparing cycles.** A latest-attempt rate from one cycle and a worst-attempt rate from another are simply different measurements. Trending them together produces a line that means nothing. - **Automated results feeding a cycle.** Repeated attempts land as repeated rows. If the totals count the latest, a suite that re-runs its failures reports better than the same suite without re-runs, on identical code. - **Hand-off between teams.** "We're at ninety-four percent" travels well and loses the rule that produced it within one retelling. ## Reading a number someone else produced 1. Ask what happens to an item that failed and then passed — this single question separates the rules faster than any documentation. 2. Check whether the total can go **down** as the cycle proceeds. A number that only ever rises is latest-attempt. 3. Look for an attempt count per item anywhere on the view. If there is none, the reduction is unstated by construction. 4. Reproduce one item by hand: read its attempt list and see which outcome the tile agreed with. ## A convention that survives contact The practical answer is not to pick a winner but to publish both halves of the picture. Keep the board on latest-attempt, because a work queue needs to show what still needs doing. Then add one derived number beside the pass rate: **how many items needed more than one attempt**. That figure costs nothing — it comes from the same attempt rows — and it restores everything latest-attempt hides, without asking anyone to reinterpret a familiar tile. For an export that will be read outside the team, go further and emit per item: the first outcome, the final outcome, and the attempt count. Any reader can then compute either rule for themselves, and nobody has to trust an unstated convention. An export that carries one derived percentage and no attempt column is the shape that causes the argument later. ## The failure mode to name out loud The dangerous case is not a wrong number, it is a *comparable-looking* number. Two cycles, both quoted at ninety-six percent, one clean and one that got there after forty re-runs, are presented to a decision-maker as equivalent. Nothing in the tool is broken — the attempts are all stored, the arithmetic is right — but the reduction threw away the distinction before the reader ever saw it. Naming which rule is in force, and publishing the re-run count next to the rate, is the whole fix.
- Your cycle counts by latest attempt. What single extra number keeps the re-run traffic honest?Publish the count of items that needed more than one attempt, next to the pass rate. It is derived from the same attempt rows, so it costs nothing, and it separates a cycle that passed cleanly from one that passed on the third try — without anyone having to argue about which counting rule is correct.
- Which rule should an export used as release evidence apply?State the rule instead of picking a favourite: emit each item's first outcome, final outcome and attempt count. A reader can then compute either total. An export quoting one derived percentage with no attempt column forces every downstream reader to trust a convention nobody wrote down.
Think of a re-sat exam. The transcript quotes the latest mark, so the student reads as a pass; the registrar's file still shows the first sitting. Latest-attempt counting is the transcript, worst-attempt counting is the file — same events, two defensible summaries.
saying these in an interview costs you the question
- Quotes a pass rate without saying which attempt counts
- Assumes every tool counts the worst attempt
- Trends latest-attempt and worst-attempt rates on one chart
- Thinks re-running items cannot change the totals
- Treats two equal pass rates as equally clean cycles