A team meets its 85% coverage target every sprint yet boundary defects keep shipping — how do you investigate?
answer
- A measure under pressure stops measuring
- Investigate the artefact, not the percentage
- Split the figure against where risk sits
- Counters are per line, not per input
- Move the pressure off the global number
basics
~20 sRead the cases rather than the number: look for missing, tautological or collaborator-only assertions, check whether the coverage sits on risky code or on trivia, and remember that a boundary defect lives inside a fully covered line by construction.
solid answer
~50 sOpen with the mechanism: once a percentage becomes a target, the cheapest way to raise it wins, and raising coverage is far cheaper than improving verification — Goodhart's law, not bad faith. So sample the cases in the modules where the defects landed and look for the shapes that lift a number without checking anything: no assertion, a not-null or no-exception check, an assertion on a collaborator call, a broad walk-through that traverses a lot and checks one weak thing, or exclusions that have crept outward. Then split the figure by module against risk, because 85% overall can be 97% on serialization and 58% on pricing. Then say the structural point: a counter is per line, not per input, so an off-by-one at a peak-fare boundary sits inside a covered line and no coverage policy would have caught it. Fix the incentive rather than the people.
code
pseudocode · 9 lines// covers the peak-window line, cannot see the boundary
test "peak fare":
assertNotNull(calculateFare(2, 5, "09:29"))
// same lines covered, now with a real oracle on both sides
test "last off-peak minute":
assertEquals(3.40, calculateFare(2, 5, "09:29"))
test "first peak minute":
assertEquals(4.75, calculateFare(2, 5, "09:30"))go deeper
Know the headline: a high coverage figure does not mean the checks are strong, and a defect in a returned value can sit inside a line the suite executes on every run.
Be able to list the concrete shapes that raise a percentage without verifying anything — missing assertions, not-null checks, assertions on collaborator calls, broad walk-throughs — and explain why boundary defects survive them.
Show the investigation: sample the risky modules, read assertions, split the figure by module against risk, ask whether a wrong value would turn a case red, and name Goodhart's law as the mechanism rather than blaming people.
Own the incentive design. Decide what is gated versus merely reported, keep the metric out of performance conversations, tier expectations by risk, and be able to defend to a stakeholder why the headline number is going to be replaced rather than raised.
### Start by naming the failure mode A team that hits its target every sprint while defects keep shipping is showing you Goodhart's law in operation: once a measure is made a target, the cheapest way to move it becomes the way it gets moved, and it stops measuring what it used to. Raising a coverage percentage is much cheaper than improving verification — you can do it without ever thinking about expected values — so pressure applied to the number flows to the cheap path. This is a property of the incentive, not a comment on anyone's character, and an investigation that opens by accusing the team will learn nothing. Also hold open a second hypothesis before you start: the number may be honest and simply pointed at the wrong code, or measuring something orthogonal to the defects being reported. Both are worth testing. ### Read the artefact, not the number Sample the suite and read the cases, weighted toward the modules where the defects actually landed. In the fare calculator, that means the pricing and peak-window code rather than the request-mapping glue. The shapes worth looking for: - **Cases with no assertion**, or with only a not-null or no-exception check. - **Tautological assertions** — comparing the result against a value computed by calling the same production logic, so the check cannot ever disagree. - **Collaborator-call assertions standing in for outcome assertions**: the case proves a call happened, never that the fare was right. - **Broad walk-through cases** that traverse a great deal of code and check one weak thing at the end. These are extremely efficient at raising a percentage. - **Report exclusions that have crept outward** — the awkward package quietly removed from measurement. - **Cases written against the implementation** so closely that they restate it and could not fail for a behavioural reason. Then ask the direct question for a sample of them: if this returned value were wrong, would this case go red? Deliberately introducing a fault and watching whether anything fails is the technique that answers it, and it converts an argument about the number into evidence. ### Check where the coverage is, and what a line can hide Split the figure by module and compare it against risk. A suite reporting 85% overall can be 97% on serialization helpers and 58% on the fare rules, and that composition is entirely consistent with a stream of pricing defects. Then face the structural point that matters for the defects described. A coverage counter is per code unit, not per input. The peak-window comparison is one line; a single departure time covers it completely. An off-by-one at the peak boundary — the 09:30 traveller billed at the off-peak rate because the comparison excludes the boundary minute — lives entirely inside a fully covered line. No percentage of any structural criterion will ever surface it, however high, because the defect is a wrong value produced by executed code. That is worth saying explicitly in the interview: this class of defect is invisible to this class of measurement by construction, so the remedy cannot be "cover more". ### What you change Investigation done, the remedies are mostly about moving the pressure off the number: 1. **Stop gating on the global percentage** and gate on the lines each change touches, plus a floor that can only rise. The global figure becomes an observation, not an objective. 2. **Risk-tier the expectation.** Pricing and fare rules carry a higher bar than presentation code; a single flat number simultaneously over-tests the trivia and under-tests the core. 3. **Move the quality conversation into review.** Reviewers read assertions: is there a concrete expected value, is each boundary exercised on both sides, would a wrong value fail this. 4. **Add boundary discipline where the defects are.** Equivalence classes and boundary values are what actually catch an off-by-one; pair each boundary with cases on either side of it. 5. **Add a direct measure of assertion strength on the risky modules**, accepting that it costs run time — on a 27-minute suite you scope it to the pricing package and to changed code rather than running it whole. 6. **Never attach the number to individual performance.** The moment it appears in a personal objective, the cheap path is the only rational path. ### The interview signal What separates a senior answer here is refusing to treat this as a tooling question. The number is behaving exactly as designed; the design was to answer a different question. Say what the measure can and cannot see, propose reading the artefact rather than arguing about the metric, fix the incentive rather than the individuals, and be honest that no coverage policy would have caught a boundary defect inside a covered line.
- What is Goodhart's law, and how does it apply to a coverage target?It says that when a measure becomes a target it ceases to be a good measure. Coverage is unusually exposed to it because the cheap way to move it — executing more code — is completely decoupled from the expensive thing it was standing in for, which is checking that behaviour is right. The stronger the pressure, and the more it touches individual performance, the more the cheap path dominates.
- Would you set different coverage expectations per module, and how would you choose the tiers?Yes. Tier by blast radius and change rate: fare and pricing rules carry a high expectation, glue and presentation code a low one, generated code none. A single flat number simultaneously over-tests trivia and under-tests the risky core, because the cheapest lines to cover are almost never the ones that hurt when they are wrong. Keep the tiers few and written down, or they become a negotiation per change.
- How do you distinguish deliberate gaming from honest but low-value cases?You mostly do not need to, and trying to is the wrong move. The artefact is the same either way — no assertion, or one that cannot fail — and the remedy is the same: read assertions in review, gate changed lines rather than a global figure, and keep the number off personal objectives. Framing it as intent turns a systems problem into a conduct conversation and guarantees you learn nothing.
It is like judging a restaurant by how many dishes the kitchen sent out. Push on that number and plates come out faster; nobody counted whether anyone tasted them.
saying these in an interview costs you the question
- Blames the team instead of the incentive
- Proposes raising the target from 85% to 95%
- Claims a higher percentage would have caught the boundary defect
- Treats the number as the deliverable in code review
- Ties the coverage figure to individual performance ratings
- Never opens the cases behind the reported figure