skip to content

Machine-drafted cases triple a regression suite's case count in a week. What has that growth not proven?

level: juniorimportance: must knowfreq 55%

answer

  1. Volume is not signal
  2. Reach is a union, not a sum
  3. Repeats change the count only
  4. Ask what became newly detectable

basics

~20 s

Case count measures how much was written, not how much can be detected. Drafted cases usually re-walk journeys existing cases already walk and check outcomes already checked, so the set of failures the suite can catch may not grow.

solid answer

~50 s

A suite's value is the **set of distinct failures it can detect**, and that set is a union, not a sum. Drafted cases come from the same requirement text and the same code the existing cases came from, so most of them repeat a journey that is already walked and check an outcome that is already checked; the count grows with drafting speed while the union barely moves. Structural coverage behaves the same way - it can only rise if something previously unexecuted now runs, and volume alone does not find unexecuted code. The honest reading of a batch is its **marginal contribution**: which behaviours became reachable that were not reachable before, and which wrong behaviour would now be caught that would previously have shipped. Growth in count without growth in either is inflation, and it is paid for on every run.

code

pseudocode · 11 lines
pseudocode
reach_before = union(executed_units(c) for c in existing_cases)
checked_before = union(asserted_outcomes(c) for c in existing_cases)

reach_after = reach_before | union(executed_units(c) for c in drafted_batch)
new_reach   = reach_after - reach_before        # empty => batch added no reach

for c in drafted_batch:
    if executed_units(c) <= reach_before and asserted_outcomes(c) <= checked_before:
        mark(c, "count only")

report(size(drafted_batch), size(new_reach), count_marked("count only"))

go deeper

for a junior

Be ready to say plainly that a case count counts artefacts rather than measuring what a suite can catch, and to give one concrete reason a brand-new case might add nothing at all.

for a middle

An interviewer expects the union effect explained: reach is the set of behaviours exercised and outcomes checked, so a batch adds nothing unless it enlarges that set. Show how you would compute the difference before and after.

for a senior

Demonstrate that you check a batch's marginal contribution before it merges, and that you can state its ongoing cost in run time and triage width rather than arguing from an intuition about size.

for a principal

Own the definition of progress your teams work to. Decide what a drafted batch must demonstrate - reach that did not exist, or a failure that would previously have shipped - and hold that line when volume is offered as evidence instead.

## What a case count counts A suite's **case count** is a count of artefacts: files, functions, rows in a plan. What a suite is *worth* is a different kind of object entirely - a **set**, namely the set of distinct wrong behaviours that would make at least one case fail. Call that set the suite's **reach**. Every case contributes two things to reach: the parts of the system its run causes to execute, and the outcomes it actually checks about that execution. The suite's reach is the **union** of those contributions across all cases, not their sum. A case whose executed parts and checked outcomes both already sit inside that union is adding a member to a set that already contains it. The count goes up by one. The reach goes up by nothing. That is the whole mechanism of a coverage illusion under bulk drafting, and it fits in one line: **counts add, reach unions.** ## Why drafted volume repeats itself Redundancy here is not carelessness; it is what the setup produces by default. - Drafts are made from the same sources the existing cases were made from - the same requirement text, the same interface description, the same code - so they rediscover the same journeys through the system. - The variation a drafting pass produces most cheaply is surface variation: different names, different step ordering, different sample values drawn from **inside one equivalence class**. - The drafting step usually has no picture of what the suite already reaches, so it cannot aim at the gap. Without that, heavy overlap with the existing suite is the expected outcome rather than a bad day. - Volume arrives faster than review capacity, so the moment where somebody once said "we already have that one" simply does not happen for most of a batch. - Where a description is ambiguous, a drafting pass tends to produce several cases for the *same* reading of it, rather than one case per reading - the opposite of what the ambiguity called for. ## Three numbers, and which of them can move | Number | Moves only when | Under a redundant batch | | --- | --- | --- | | Case count | any case at all is added | rises by the full size of the batch | | Structural coverage | code that no case executed now runs | flat, or drifts slightly on setup paths | | Reach | a new outcome is checked on behaviour the suite exercises | unchanged | The middle row is the one that fools people in both directions. A structural figure can only rise when some part of the system that nothing reached is now reached, so adding traffic to already-executed code cannot move it: a batch that leaves the figure flat has genuinely told you something. But a batch that *does* move it has told you only that something new **ran**. Reach needs the second ingredient as well - an outcome checked - before a wrong behaviour in that newly executed code could fail anything. ## Reading a batch honestly The honest read of a drafted batch is its **marginal contribution**, and getting it is mechanical: 1. Before the batch, record the union of executed units across the suite and the union of asserted outcomes. 2. Run with the batch and record both unions again, then subtract. 3. Mark every case whose executed set and asserted set both fall inside the old unions as **count only**. 4. For each case that remains, name the wrong behaviour it would catch, and check whether an existing case already catches it. The output is one small honest sentence. Not "200 cases added", but "six behaviours became reachable, four of which were already checked elsewhere". That sentence is what the batch actually did, and it is usually a great deal shorter than the batch. ## What the inflation costs Cases are cheap to produce and are never cheap to keep. - **Run time.** Every redundant case is paid for on every run, for as long as it exists, in wall-clock time and in machine capacity. - **Triage width.** One defect sitting under forty redundant cases produces forty failures to read, and forty failures look like a large problem rather than a single one. - **Reading cost.** Anyone changing the code beneath those cases must first work out what each of them claims, and forty near-identical claims take far longer to understand than one. - **False confidence.** The dangerous cost is the belief that an area is well covered because many cases mention it, because that belief stops the work that would actually have closed the gap. ## When volume genuinely pays Bulk drafting is not the problem; **unaimed** bulk drafting is. Point the same capacity at a gap that has already been identified - a list of branches nothing executes, an inventory of error paths nobody checks, the input classes a stated rule defines but no case supplies - and every draft lands outside the union by construction. Then count and reach rise together, which is the only condition under which the count was ever worth quoting.

  • How do you tell a redundant drafted batch from a deliberate set of equivalence-class variations?
    By what varies and why. A deliberate set walks one journey with inputs chosen from different classes - a boundary, an empty value, an over-long value - and every variation can fail on its own while the others pass. Redundant drafts vary surface detail such as names, ordering and phrasing while the inputs stay inside a single class, so no variation is capable of failing alone.
  • The batch cost nothing to write. Why not keep all of it and see whether it ever catches something?
    Because writing is the one cost that was cheap. Each case is paid for on every run in wall-clock time, in triage whenever it fails for an unrelated reason, and in the reading time of whoever must understand it before changing the code beneath it. A case with no path to a failure of its own charges that rent indefinitely and never repays it.

Photocopying a map does not enlarge the territory it shows. You end up carrying much more paper over exactly the same ground.

saying these in an interview costs you the question

  • Treats a larger case count as a stronger suite
  • Calls the batch a coverage win without checking what newly executed
  • Assumes drafted cases must be new because their names differ
  • Says volume costs nothing because nobody wrote it by hand
  • Cannot name a single wrong behaviour the new cases would catch