skip to content

What are the limits of code coverage as a quality signal, and how should an org set coverage policy without encouraging gaming?

level: principalimportance: nice to knowfreq 35%

answer

  1. Execution, not verification — assertion-free tests still 'cover'
  2. Can't see missing code or assertion quality
  3. Goodhart: a measure-as-target gets gamed (assertion-free/exclusion creep)
  4. Gate on new/changed code, branch/instruction, realistic floor
  5. Mutation testing (PIT) measures the assertion quality coverage can't

basics

~20 s

Coverage shows which code ran during tests, not whether the tests check the right things. You can hit 100% with tests that assert nothing. So set realistic targets, focus on new/changed code, and pair coverage with techniques like mutation testing rather than chasing a single big percentage.

solid answer

~50 s

Coverage measures **execution**, not **verification** — a test can run a line and assert nothing, so high coverage can coexist with weak tests and real bugs. It also can't see *missing* code (a forgotten case is never measured) and rewards quantity over assertion quality. By Goodhart's law, once a percentage becomes a target it gets gamed — assertion-free tests, exclusion creep, trivial getter tests. A sound policy: gate on **new/changed-code** coverage rather than a whole-project average (stops regressions without demanding legacy retrofits); prefer **branch/instruction** counters; exclude only generated/untestable code; and treat the threshold as a floor, not a goal. To judge *quality* of tests, add **mutation testing** (e.g. PIT), which mutates the code and checks that tests fail — directly measuring whether assertions catch faults. Use coverage to find *untested* code, and review/mutation testing to ensure the tests that exist are meaningful.

go deeper

for a junior

Understands that 100% coverage doesn't guarantee correct or well-asserted tests.

for a middle

Can explain that coverage measures execution not verification, and that it can't see missing code; knows assertion-free tests still register coverage.

for a senior

Articulates Goodhart-style gaming risks and recommends new-code gates, branch counters, and constrained exclusions; introduces mutation testing for critical code.

for a principal

Sets org-wide coverage policy: risk-weighted targets, new-code gates, mutation testing for high-risk modules, exclusion governance, and framing coverage as a discovery/regression tool rather than a quality certificate — explicitly designing against metric gaming.

## Coverage measures execution, not correctness **Code coverage** answers one narrow question: *was this code executed by some test?* It does **not** answer *did a test verify the code behaves correctly?* These differ because a test can call a method and make **no assertions** — the code runs, coverage counts it, but nothing would fail if the result were wrong. So coverage is a **necessary-but-not-sufficient** signal: low coverage definitely means untested code, but high coverage does **not** prove the tests are good. ## What coverage structurally cannot see - **Missing logic.** Coverage can only measure code that *exists*. If you forgot to handle the empty-list case and wrote no code for it, there's nothing to be uncovered — the gap is invisible. Coverage never flags a missing branch you didn't write. - **Assertion quality.** Two test suites with identical 100% coverage can have wildly different bug-catching power depending on what they assert (or don't). - **Oracle problems / integration.** Coverage of a unit says nothing about whether components compose correctly, about timing, concurrency, or real I/O. ## Goodhart's law and gaming **Goodhart's law:** *"When a measure becomes a target, it ceases to be a good measure."* Make a coverage percentage a hard gate and humans optimize the number, not the goal: - **Assertion-free tests** that execute code purely to color it green. - **Exclusion creep** — excluding packages until the bar is met. - **Trivial tests** of getters/setters/DTOs that add coverage but no value. - **Threshold inflation** demands (e.g. "go to 95%") that push teams toward the above. A bare percentage thus *can* drive worse engineering if mismanaged. ## Designing policy that resists gaming 1. **Gate on new/changed code, not the project average.** A whole-project ratio is dominated by existing code, so regressions hide and legacy gaps demand expensive retrofits. A **diff/new-code** gate (e.g. Sonar's *new code* condition, or a CI step that scopes coverage to changed lines) enforces coverage exactly where risk is introduced — high leverage, low friction. 2. **Pick the right counter.** Prefer **branch** or **instruction** over line: branch exposes untested true/false paths; instruction is formatting-independent and harder to game. 3. **Set a realistic floor, not an aspirational ceiling.** A modest, *enforced* floor (e.g. 70-80% branch on new code) beats an unenforced 95% wish. The floor's job is to stop *regressions*, not to certify quality. 4. **Constrain exclusions.** Allow only generated/untestable code; review exclusion changes like any other code so they can't be used to dodge the gate. 5. **Measure test *quality* separately with mutation testing.** A tool like **PIT** introduces small faults ("mutants" — flip `>` to `>=`, return `null`, remove a call) and reruns the tests; a mutant **"killed"** (a test fails) means the tests actually verify that behavior, a **"survived"** mutant means coverage was empty calories. Mutation score directly targets the assertion-quality gap that coverage can't see. It's more expensive to run, so apply it to critical modules or on a schedule. 6. **Risk-weight effort.** Demand high coverage where failure is costly (payments, auth, parsing) and accept less on low-risk glue. Uniform targets misallocate effort. ## How to talk about it Frame coverage as a **discovery tool** ("here's code no test touches — investigate") and a **regression guardrail** ("new code must be tested"), never as a proof of quality. Pair it with code review (do the tests assert the right things?) and mutation testing (do the assertions actually catch faults?). The mature org outcome: coverage finds the gaps, review and mutation testing ensure the tests filling them are real — and no one is rewarded for a green number alone.

  • How does mutation testing address the weakness that coverage can't measure assertion quality?
    It deliberately introduces small faults (mutants) into the code and reruns the tests. If no test fails, the mutant 'survived' — proving the tests executed but didn't verify that behavior. Killing mutants requires real assertions, so the mutation score measures bug-catching power, not mere execution.
  • Why does gating only on whole-project coverage average often backfire?
    The average is dominated by legacy code, so regressions hide beneath it while demands to raise it force expensive, low-value retrofits and incentivize gaming. Scoping the gate to new/changed code stops regressions at the source with far less friction.

saying these in an interview costs you the question

  • Treating a coverage percentage as proof of test quality or correctness.
  • Mandating 100% coverage, which pushes teams toward assertion-free and trivial tests.
  • Ignoring that coverage cannot detect missing code or weak assertions.
  • Setting one uniform target regardless of module risk.

context