Your coverage number rose after a batch of generated unit tests. Why might the team be no safer?
answer
- It counts what ran, not what was decided
- Tests used to be expensive to write
- Cheap execution, unchanged cost of judgement
- Absence of coverage is still honest evidence
basics
~20 sThe number counts code that ran. Generation made running code cheap and left the hard half - deciding what the right answer is - exactly as expensive, so a figure that stood in for effort no longer does.
solid answer
~50 sA test has two halves: the part that runs the code, and the part that decides what the right answer is. Coverage has only ever counted the first. It was a usable stand-in for the second because writing a test cost enough that nobody wrote one without a reason to check something - the expensive half came attached. Generation removes that cost and leaves the stand-in behind, so the figure can now move at drafting speed while the set of wrong behaviours your suite would reject may not move at all. What the number still tells you honestly is where nothing runs: unexecuted code is genuinely untested, and that half of the signal survives. What it no longer licenses is the inference from a rising figure to a safer change. Read the new tests in the diff instead, and for each one name the wrong result it would reject.
code
pseudocode · 14 lines# three generated tests, all green; all of them execute the code under test
test "accrue returns a number for a mid-year joiner":
assert is_number(accrue(MAR_01, DEC_31)) # true of any return value
test "accrue is never negative":
assert accrue(DEC_15, DEC_31) >= 0 # also true when it wrongly returns 1.5
test "more service accrues no less":
assert accrue(MAR_01, DEC_31) >= accrue(MAR_01, JUN_30) # holds under the bug too
# the branch is executed three times over and the rule is checked zero times.
# the line that would have caught it, given the rule "completed months only":
assert accrue(DEC_15, DEC_31) == 0.0go deeper
Know that the figure records which code ran, and that a batch of generated tests can move it without adding anything that would reject a wrong answer.
Explain why the figure used to be a decent stand-in for care - tests were expensive - and which half of the signal survives generation and which does not.
Show what you read instead at the size of one change, and be able to say what the delta does and does not license in front of somebody quoting it.
Decide what a coverage threshold is for once producing execution is nearly free, and make sure the team is not gating on a number that is now trivially satisfiable.
## The proxy that generation broke A coverage figure has never been a measure of checking - what the instrumentation records and what it cannot see is a subject of its own, and a well-worn one. The interesting question here is different: **why did the figure work as a proxy for so long, and what changed?** It worked because of economics. Writing a test used to cost a developer real minutes: naming it, arranging the inputs, deciding the expected result, making it pass. Nobody paid that for a test they had no reason to write. So a line that was covered was, with decent odds, covered by somebody who had thought about what that line should do - and the figure inherited the credibility of the thought behind it. Generation changes one term in that equation and not the other. The half of a test that **executes** code is now nearly free. The half that **decides what the right answer is** costs exactly what it always did, because it is a question about the product rather than about the code. A batch of generated tests can therefore move the number a long way while adding little or nothing to the set of wrong behaviours the suite would reject. The proxy is not wrong about what it measures; it has simply stopped being attached to the thing you were reading it for. ## What still holds and what does not | the inference | still sound? | why | | --- | --- | --- | | code nothing executes is untested | yes | an absence is direct evidence, and generation does not change it | | the number fell, something is missing | yes | it still finds code that no test touches | | the number rose, we check more behaviour | **no** | execution rose; whether anything now rejects a wrong result is unaddressed | | the number is high, the risky paths are covered | **no** | it was never true, and cheap tests make it easier to believe | Keeping the first two is the point. The overcorrection - *coverage is meaningless now* - throws away a signal that is still honest in one direction, and it is as wrong as the inference it is reacting to. ## What to read instead, at the size a developer actually works at You are looking at a change with several new tests in it. The delta on the report is the least informative thing on the screen. 1. **Read the new tests, not the number.** In a diff-sized batch this is minutes of work. 2. **For each one, name a wrong result it would reject.** If you cannot, it contributed execution and nothing else. 3. **Ask where the batch landed.** Generated tests cluster where the code was easiest to call - pure functions with simple parameters - and thin out exactly where the risk is: the branch with three collaborators, the failure path, the awkward boundary. 4. **Look for the cases that are absent rather than the lines that are green.** The situation nobody implemented has no line to light up, so it cannot show as a gap in a figure that only counts lines someone wrote. ## The ecosystem the team works in changes how this bites In some ecosystems coverage instrumentation is part of the standard toolchain: one command, on by default in the build, printed on every change, and often wired into the merge conversation. In others it is a separate, awkward, sometimes unreliable add-on that a team may never get round to setting up. The visible-number ecosystems are where the broken proxy does its damage, because the figure is cheap, prominent and load-bearing - it is right there in the review, it moves, and a moving number feels like progress. The teams without a number are not better off; they are merely not being misled by that particular one. They lose the honest half of the signal too, and have to notice untested code by reading. The lesson is the same in both: **how much you trust a generated suite should not depend on whether your toolchain prints a figure.** It depends on whether the tests in it can disagree with the code, and that is visible only by reading them. ## What to say when somebody quotes the delta Answer in two parts, because both are true. The figure moved, and that is not nothing - the new tests do execute code that nothing executed before, so a crash on those paths would now be found. And the figure moving does not mean the suite would now reject a wrong result on those paths, because that depends on what each test asserts and where its expectation came from. If a decision is going to rest on the number, the cheap way to make it safe is to read the new assertions first and report *that*.
- Is there any reading of the number that generation has not weakened?Yes - the downward one. Code that nothing executes is untested, and that stays true however the tests were produced. Treat a fall as an alarm worth acting on and a rise as a fact about execution that licenses no conclusion about detection until somebody has read the assertions.
- How would you keep a batch of generated tests from inflating the figure unnoticed?Make the reviewable unit the tests rather than the delta: require that each new test names the wrong behaviour it would report, and read a sample of any large batch closely. If a threshold exists, be clear with the team that it is a floor against untouched code, not a claim about how much is checked.
saying these in an interview costs you the question
- The number went up, so the change is better tested than it was.
- Coverage is meaningless now that tests can be generated.
- A high figure means the risky paths have been exercised.
- Generated tests raise coverage, so they pay for themselves.
- If the threshold is met, nobody needs to read the new tests.