How do you catch a machine-drafted end-to-end case that duplicates existing coverage under a new name?
answer
- Duplicated intent, not duplicated text
- Reduce each case to a short signature
- Search the claim, never the name
- Would deleting it lose any detection?
basics
~20 sReduce the draft to a signature — precondition, trigger, and the fact its assertions prove — then search the pack for that fact rather than the name. Matching signatures mean a duplicate however different the wording is.
solid answer
~40 sBulk drafting produces duplicates of **intent**, not of text: new name, new wording, sometimes a different route, but the same fact proven at the end. Catch it by reducing the draft to three things — the condition it starts in, the action under test, and the claim its final assertions establish — and searching the pack for the claim, since names are the part authors improvise and the least searchable. Then apply the deletion test: if this case did not exist, would any regression go uncaught? Same claim with different input values is a data row, not a case. Same claim reached by a genuinely different actor or channel usually is a real case, but it must be named for that difference, or the next reviewer will face the same puzzle.
go deeper
Recall that two cases can prove exactly the same thing while looking completely different. Before adding a case, check whether the pack already proves the fact you are about to assert, and ask a reviewer if you are unsure.
Explain the mechanics of finding one. Reduce a case to precondition, trigger and claim; search on the claim because names are unreliable; and apply the deletion test to decide whether the draft adds detection or only runtime.
Show the judgement calls. Distinguish a duplicate from a genuine variant by whether the code paths differ, choose between folding inputs into an existing case and keeping a separate one, and explain what a pack of near-twins costs a team over a year.
Own the conditions that make this findable. Decide how cases are grouped and named so a duplicate surfaces in one search, and how review capacity is protected when drafting throughput can outpace anyone's ability to read the pack.
## Duplication of intent, not of text A duplicate drafted case rarely looks like a duplicate. Its name is new, its wording is new, its steps may take a different route through the product, and its data is freshly invented. What repeats is the **fact it proves**. Two cases are duplicates when deleting either one loses no detection: any regression that would break the first would also break the second. This is the form duplication takes when cases are drafted in volume, because a drafting model has no memory of the pack. It is asked for a case about declined payments and it writes a good one — the fifth good one, differing from the other four only in surface detail. Nothing in the diff says so. The reviewer is the only place this can be caught, because after the merge the case simply looks like one more green line. ## The signature of a case The practical trick is to reduce every case to a short signature and compare signatures rather than text. A workable signature has three parts: 1. **The precondition** — what world the case starts in, stated as a condition rather than as steps. 2. **The trigger** — the one action whose effect is under test. 3. **The claim** — the fact the final assertions establish. Two cases whose signatures match are the same case, however different the middles look. A case whose middle is elaborate and whose signature is identical to an existing one is *worse* than a plain duplicate: it costs more runtime for the same detection. Signature reduction is also how you find the draft's real claim when the name lies. A drafted title like *"verify checkout works for a declined card"* is not a claim. *"A declined authorisation leaves the order unpaid and the stock unreserved"* is. ## Searching for it before the merge 1. Reduce the draft to its signature. 2. Search the pack for the **claim**, not the name: the status, the invariant, the computed value. Names are the least searchable part of any pack, because they are the part authors improvise. 3. Search for the **trigger** as well; a duplicate often shares the trigger and reaches the claim by another route. 4. Read the two or three closest cases end to end. Diff their signatures, not their text. 5. Apply the deletion test: if this draft did not exist, would any regression go uncaught? ## Duplicate, variant, or genuinely new | The draft… | Verdict | Action | | --- | --- | --- | | Same precondition, trigger and claim | Duplicate | Reject; the pack already proves it | | Same claim, different input values | Variant | Fold the values into the existing case as extra rows | | Same claim, genuinely different route or actor | Distinct | Keep, and name it for the difference | | Same trigger, different claim | Distinct | Keep; it protects something else | | Broader claim that subsumes an existing one | Replacement | Merge, then retire the narrower case deliberately | The middle row is the one reviewers get wrong in both directions. A second set of input values is almost never worth a second case: it is worth a row. A second **actor** — a different role, a different entitlement, a different entry channel — usually is a real case, because the code paths differ, and it should be named so the difference is visible without opening it. ## What to do with the draft Rejecting is fine, but a bare rejection teaches nothing and the next drafting run produces the same case again. Three better endings: - **Fold and close.** Move the draft's inputs into the existing case as additional data, and say which case absorbed it. - **Keep and rename.** If it is a real variant, rename both cases so the axis of difference is in the names. Duplicates hide behind vague names more than behind anything else. - **Replace deliberately.** If the draft is genuinely better than what exists, merge it and retire the older case in the same change, rather than leaving both. ## Why the pack's organisation decides whether you can see this at all A duplicate is findable in proportion to how well the pack is indexed by behaviour. If cases are grouped by the behaviour they protect and named for the fact they prove, a reviewer's search finds the neighbour in seconds. If they are grouped by screen, by author, or by the sprint that produced them, the same search finds nothing and every reviewer independently concludes the draft is new. That is why this review step is worth practising even though nobody's offer ever turned on it: the cost of missing it is not one wasted case, it is a slow drift in which the pack grows and the set of facts it actually protects does not.
- A drafted case proves an existing fact but through a different user role. Duplicate or not?Usually not a duplicate: a different role exercises different authorisation and often different code paths, so a regression can break one and not the other. Keep it, but put the role in the name so the axis of difference is visible without opening the case. If the role changes nothing about how the outcome is produced, it is a variant and belongs as an extra data row on the existing case.
- Why is searching for a duplicate by case name so unreliable?Names are the improvised part of a pack. Two authors describing the same fact will phrase it differently, and a drafting model will phrase it differently again, so a name search finds nothing and the reviewer concludes the draft is new. Searching for the asserted fact — the status, the invariant, the computed value — matches on the part that cannot be paraphrased away.
- What makes duplicates almost impossible to find in some packs and easy in others?How the pack is indexed. Grouped by the behaviour protected and named for the fact proven, the neighbouring case turns up in one search. Grouped by screen, by author, or by the release that produced it, there is no path from a draft to its twin, and every reviewer independently rediscovers nothing. Organisation, not diligence, is what makes this review step affordable.
saying these in an interview costs you the question
- Compares case names or wording instead of asserted facts
- Treats a new set of input values as a new case
- Assumes a different route always means different detection
- Keeps both cases because deleting a test feels unsafe
- Rejects the draft without saying which case already covers it