skip to content

In a suite grown by bulk drafting, how do you find cases that duplicate an existing case's intent under a new name?

level: seniorimportance: nice to knowfreq 26%

answer

  1. Names are the weakest evidence
  2. Compare behaviour, not titles
  3. Executed units and checked outcomes
  4. Cluster cases with matching signatures
  5. Never the sole failure in a run

basics

~20 s

Compare what cases do rather than what they are called: the units each case executes, the outcomes it checks, and the requirement it claims to cover. Cases whose signatures coincide share one intent whatever their names say.

solid answer

~50 s

Names stop being evidence as soon as drafting is cheap, because a drafting pass renames freely while reproducing the same journey. The reliable signals are **behavioural**: capture, per case, the set of units its run executes and the set of fields and outcomes it checks, then cluster cases whose two sets coincide - an exact or near-exact match is duplicate intent regardless of the titles. Two cheap corroborations help. A tag linking each case to the requirement it covers makes one requirement carried by nine cases visible, because you count distinct requirements instead of cases. And the run history shows independent detection: a case that has never been the only failing case in any run has never demonstrated that it catches anything alone. It needs per-case execution data, so run it after a large batch rather than per change.

code

pseudocode · 14 lines
pseudocode
signature(case) = (frozen(executed_units(case)),
                   frozen(checked_outcomes(case)),
                   intent_tag(case))

for group in group_by(all_cases, key = signature) where size(group) > 1:
    report(group, reason = "same intent, different names")

# near-duplicates: identical journey, one case claims a subset of the other
for a, b in pairs(all_cases):
    if executed_units(a) == executed_units(b) and
       checked_outcomes(a) <= checked_outcomes(b):
        report(a, reason = "subsumed by " + b.id)

ratio = distinct(intent_tag(c) for c in all_cases) / count(all_cases)

go deeper

for a junior

Know that two cases with different names can still check exactly the same thing, and that the way to tell is to read what each one does and checks rather than what it is called.

for a middle

Be able to describe a comparison you could actually run: which units each case executed and which outcomes it checked, then look for cases whose two sets match or where one is contained in the other.

for a senior

Show that you can source the data - per-case execution records, an inventory of what each case checks, requirement tags, run history - and that you know execution overlap alone yields false positives when the checks differ.

for a principal

Own the definition of duplication your organisation works to, and the cadence at which evidence is gathered. An analysis that runs twice a year and is trusted is worth more than one that runs nightly and is ignored.

## Why names stopped carrying information Deduplicating a suite used to lean on names, and it worked for an accidental reason: writing a case was slow enough that people searched before they wrote, and a team's vocabulary was small enough that two cases about the same thing tended to be called similar things. Bulk drafting removes both conditions at once. Naming is the cheapest part of a draft to vary, and the drafting step consults nothing about what already exists. The result is the specific failure this leaf is about: **the same intent wearing a name that no search would match.** Once names are noise, duplication has to be detected from **behaviour**. The useful news is that behaviour leaves records. ## Three behavioural signatures - **Execution signature** - the set of units a case causes to run: files, functions, branches, interface operations. Two cases that walk the same journey produce the same set, whatever they are called. - **Assertion signature** - the set of fields, outcomes and transitions the case actually checks. This is the part that says what a case *claims*, and it is the part that decides redundancy. - **Intent tag** - an explicit link from the case to the requirement, rule or acceptance criterion it exists to cover. Cheap to add, and it converts "how many cases" into "how many distinct requirements", which is the count that stops rewarding volume. Redundancy is a statement about the first two **together**: two cases are duplicates when their execution sets coincide **and** their assertion sets are equal, or one is contained in the other. Execution overlap on its own is not enough, and treating it as enough is the most common way this analysis produces confident false positives. ## Comparing the signals | Signal | What it shows | Blind spot | Cost to obtain | | --- | --- | --- | --- | | Execution signature | a shared journey | says nothing about what is checked | needs per-case execution recording | | Assertion signature | a shared claim | misses the same check written differently | needs assertions to be inspectable | | Intent tag | shared purpose as declared | only as honest as the tagging | cheap, but must be enforced at authoring | | Sole-failure history | independent detection ever demonstrated | needs a long history and real failures | free, if run results are retained | No single row is sufficient, and that is the point. Each is wrong in a different direction, so agreement between two of them is worth far more than a strong reading from one. ## Corroboration from the run history The run history answers a question none of the static signatures can: has this case ever detected anything **by itself**? Retain results long enough and you can ask, per case, how often it was the only failing case in a run. A case that has never been the sole failure in a year has never demonstrated independent detection. That is evidence rather than proof - a case guarding a rare path may simply never have been needed - but set beside a matching execution and assertion signature it becomes a strong reading. Two readings come out of that history, and they point opposite ways: - A case that has **never** been the sole failure is a candidate: nothing it caught was ever caught by it alone. - A case that is **frequently** the sole failure is either guarding something nothing else guards, or is unreliable. Either way it is the opposite of redundant, and the analysis should leave it alone. ## A procedure that survives contact 1. Collect, for one full run, every case's execution set, assertion set and intent tag. 2. Group by exact signature match. Those groups are duplicates with high confidence and need no argument. 3. Look for **subsumption**: identical execution sets where one case's assertion set is contained in another's. These are the interesting ones and they need a human read. 4. Count distinct intent tags against case count. A ratio far below one says a batch multiplied cases without multiplying purposes. 5. Overlay the sole-failure history and rank the clusters by how little evidence of independent detection they contain. Run this **periodically** - after a large drafting batch, or on a slow cadence - rather than on every change. It needs a full run with execution recording switched on, which is expensive, and its output is a reading exercise rather than a gate. An analysis that runs twice a year and is trusted beats one that runs nightly and is skipped. ## What the analysis does not decide The output of all this is **evidence about which cases carry the same intent**, presented as clusters together with the signals that put them there. It is deliberately not a verdict. Two cases can be behaviourally identical and both worth having: - one exists to document a contract another team relies on; - one runs in a configuration the signature does not capture; - one is the only case a downstream group ever reads. What the analysis buys is that the conversation is about behaviour with data attached instead of about names, and that a suite which grew fivefold in a month can finally be described by the number of distinct things it claims rather than the number of files it contains.

  • Two cases share an execution signature but check different fields. Are they duplicates?
    No, they are complementary. One journey checked for two different outcomes is legitimate design, and only equal assertion sets, or one contained in the other, make a case redundant. Judging on the execution set alone is the most common way this analysis produces false positives, and it is the failure mode that discredits the whole exercise fastest.
  • Your runner cannot produce per-case execution data. What is the next-best signal?
    The intent tag and the run history. Counting distinct requirements covered rather than cases exposes nine cases carrying one requirement, and a case that has never failed alone across a year of runs has never shown independent detection. Both are weaker than an execution signature, both are nearly free, and together they still separate the obvious clusters.

Two people can describe the same walk in completely different words. You tell that it was the same walk from the footprints, not from the story.

saying these in an interview costs you the question

  • Deduplicates by comparing case titles and step descriptions
  • Calls two cases distinct because their inputs differ cosmetically
  • Treats an identical execution set alone as proof of duplication
  • Assumes drafted cases are unique because a machine varied them
  • Runs the analysis on every change and abandons it when it is slow