A machine-drafted end-to-end case arrives with a green run; why is that weak evidence, and what do you read first?
answer
- A case that cannot fail is green
- Green means executable, not capable of detecting
- State the claim before reading steps
- Ask for one correct failure, not ten passes
basics
~10 sGreen proves only that the case executes here today; a case that can never fail is green too. Read the claim first, then the assertions, the setup, and whether the pack already proves it.
solid answer
~40 sA passing run establishes that the case is valid and executable against this target right now. It cannot distinguish a working oracle from one that is incapable of failing — and incapable of failing is exactly the defect drafted cases carry: assertions that restate a step, preconditions borrowed from an existing world, a claim the pack already proves. All three are green forever. So read in an order that disqualifies cheaply: state the fact the case protects in one sentence; check whether any assertion would survive the behaviour being removed; trace every precondition to the line that created it; search the pack for the claim. Then get the evidence that actually licenses a merge — run it against a build with the behaviour absent and watch it fail at the assertion it should.
code
text · 11 linesReview of a drafted case — evidence checklist
claim : a declined authorisation leaves the order unpaid
and releases the reserved stock
assertions : 2 of 5 survive removal of the behaviour -> trim 3
setup : account + order created in the case -> ok
stock level assumed, never created -> fix
already proven? : searched "unpaid" + "stock release" -> no match
correct failure : run against build with rule disabled
-> failed at the stock assertion -> ok
verdict : merge after the two edits abovego deeper
Remember that passing is not the same as working. A check that cannot fail passes too, so before trusting a new case, ask what would have to break in the product for it to go red.
Explain the mechanics of the review. Give an order — claim, assertions, setup, duplication, then the body — and say why each earlier step disqualifies more drafts per minute of reading than the one after it.
Demonstrate judgement about evidence. Rank what a green run, repeated green runs and one observed correct failure each establish, describe how you obtain that failure cheaply, and explain what a pack of never-failing cases does to a team's trust.
Own the economics. Drafting throughput can outrun reading capacity, so decide what evidence a drafted case must carry before it is even offered for review, and who is accountable when the bar quietly becomes a rubber stamp.
## What a green run proves, exactly A green run of a newly drafted end-to-end case proves three things and no more: the case is syntactically valid, it can execute against this target right now, and nothing in it threw. It says nothing about the only property that matters — whether the case can **fail for the right reason**. That gap is not academic. The characteristic defect of machine-drafted cases is precisely the kind that runs green forever: an assertion that restates a step the case just performed, a precondition borrowed from a world that happens to exist, a case that duplicates a claim already proven. Every one of those is green on the first run and green on the hundredth. Green is the state a broken oracle and a working oracle share, which makes it the weakest evidence available and, unfortunately, the cheapest to produce — so it is the evidence reviewers reach for. ## Make it fail on purpose The strongest cheap signal is the inverse of the run everybody does: **show the case failing for the reason it claims**. In practice that is one of three moves, in order of preference: - **Remove the behaviour.** Point the case at a build with the feature disabled or the rule switched off. If the case still passes, it is not an oracle for that feature, whatever its name says. - **Perturb the expectation.** Change the expected value to something wrong and confirm the case fails, and fails at the assertion you expect rather than somewhere earlier. - **Break the precondition.** Remove the setup and confirm the case reports a setup failure rather than a behavioural one — which also validates the guard. A drafted case that has been shown to fail correctly once is worth far more than one that has been seen to pass ten times, and the demonstration costs a single run. ## A reading order for a drafted case The order matters because attention is finite and the expensive checks must come after the cheap disqualifying ones: 1. **The claim.** Read the name and the final assertions and state, in one sentence, what fact this case protects. If you cannot, stop and ask; nothing downstream can be judged without it. 2. **The assertions.** Would any of them survive the behaviour being removed? Restating assertions are the single most common defect and they are visible in seconds. 3. **The setup.** Is every fact it depends on created by the case, or explicitly guaranteed? Is a wrong precondition caught at the guard? 4. **The pack.** Does something already prove this claim? A perfectly-written duplicate is still a rejection. 5. **The body.** Only now read the steps in detail, and only for the parts the claim depends on. 6. **The failure demonstration.** Ask for evidence of it failing correctly, or produce that evidence yourself. Steps 1 to 4 disqualify most bad drafts and cost minutes. Step 6 is what actually licenses the merge. ## Evidence, ranked | Evidence | What it establishes | Cost | | --- | --- | --- | | One green run | The case executes here, today | Minutes, automatic | | Ten green runs | It executes repeatably here | Multiples of the above | | A stated claim | The case has something to protect | Seconds of reading | | Assertions surviving the removal test | The case can detect | Seconds of reading | | An observed correct failure | The case does detect, and reports it usefully | One run | | A pack search for the claim | It is not paid-for twice | A few minutes | The table is the argument. The cheapest evidence sits at the top and proves the least; the evidence that licenses a merge sits in the middle rows and costs almost nothing more. ## Where the attention goes wrong Two habits waste most of a reviewer's budget. The first is **reading top to bottom**: opening the draft at its setup lines and working down, so the reviewer spends their sharpest attention on plumbing before they know what the case even claims. The second is **treating the run as the review**: seeing green, skimming, approving. Both are comfortable, both scale badly, and both get worse as drafting throughput rises, because the volume of drafts grows while the number of hours available to read them does not. ## Closing the review The merge comment for a drafted case is short and it names the evidence: the claim in one sentence, confirmation that the assertions fail when the behaviour is absent, and the case it was checked against for duplication. That comment is also the record the next person needs when the case starts failing six months later and someone has to decide whether the case or the product is wrong.
- How would you demonstrate that a drafted case can fail for the reason it claims?Run it against a build where the behaviour is switched off, and confirm it fails at the assertion that carries the claim rather than earlier. If that build is not available, perturb the expected value and check the failure lands where you expect. Either way it is one run, and it is the only evidence that separates a working oracle from a case that cannot fail at all.
- Why read the final assertions before the steps of a drafted case?The assertions carry the claim, and the claim is what every other judgement depends on: whether the setup is relevant, whether the pack already covers it, whether the steps are the right ones. Reading top to bottom spends the sharpest attention on plumbing before the reviewer knows what the case is for, and most bad drafts are disqualified by their assertions alone.
- Ten green runs instead of one — does that change your confidence in a drafted case?It changes a different confidence. Repeated passes speak to stability on this target, which is worth knowing, but they add nothing about detection: a case that cannot fail is stable by construction. Stability evidence and detection evidence are separate, and only the second one licenses the merge.
A smoke alarm that has never been tested is silent for two reasons that look identical from the hallway: nothing is burning, or the battery is dead.
saying these in an interview costs you the question
- Approves a drafted case because the run came back green
- Reads the case top to bottom starting from setup lines
- Thinks repeated passes prove the case can detect a break
- Never asks what single fact the case protects
- Skips the duplicate check because the case is well written