Why gate a merge on new-code (diff) coverage instead of the whole repository's global percentage?
answer
- Think about the size of the denominator
- One change against the whole repository
- Insensitive, and unfair on legacy code
- Measure only the lines this change touched
- Let the total rise as a consequence
basics
~20 sA global percentage is a ratio over the whole codebase, so one change is too small to move it and the gate cannot detect an untested change. Diff coverage measures only the changed lines, so the signal is proportional and the fix is local.
solid answer
~50 sThe global figure has an enormous legacy denominator, so it is insensitive: in a repository of roughly 78,000 measurable lines sitting at 91.4%, a change adding 240 untested lines moves the total by about three-tenths of a point and passes any realistic floor. It is also unfair and unactionable — an engineer fixing two lines in an under-covered repository is blocked by debt they cannot repay in that change, so the gate gets disabled or worked around. Diff coverage restricts the ratio to the added and modified lines, which makes an untested change read as near zero, points the report at code the author just wrote, and lets the global figure rise as a consequence rather than as a target. Pair it with a ratchet on the global floor so existing coverage cannot silently erode.
go deeper
Know that the coverage check on a change usually looks only at the lines that change added or modified, and that this is why an untested small change fails even when the overall project figure is high.
Explain the denominator argument with rough arithmetic, contrast actionability and fairness against a repository-wide floor, and describe a ratchet as the complementary guard against erosion.
Show judgement about the exception cases: refactors and file moves, generated code, the minimum-lines rule, an override path with a recorded reason, and why the threshold sits below 100% deliberately.
Own the policy as a system of incentives. Decide what is gated versus reported, what the override costs, and how the change gate interacts with risk-tiered expectations so the pressure lands on the code that actually carries risk.
### The problem a global percentage has A global coverage percentage is a ratio over the whole codebase. In any repository with history, the denominator is enormous and almost entirely legacy, so the number is dominated by code nobody is touching this quarter. That has two consequences, and both of them break the number as a merge gate. First, it is **insensitive**. Take the fare calculator repository: roughly 78,000 measurable lines, currently at 91.4% overall. A change that adds 240 lines of new zone-pricing logic with not a single case behind it drops the global figure by about three-tenths of a percentage point. If the gate is "do not fall below 90%", the change sails through. The signal from the change is drowned by the mass of everything else. Second, it is **unfair and unactionable**. If the repository sits at 61% and the policy says 80%, an engineer fixing a two-line defect is told the build fails for a debt they did not create and cannot repay in this change. The rational response is to disable the gate, request an exemption, or write cheap cases against whatever is easiest to cover — which is never the risky code. ### What diff coverage measures instead New-code or diff coverage restricts the ratio to the lines this change added or modified. The denominator becomes the change itself, so the signal scales with the change rather than with the repository: - **Proportional.** Two hundred and forty new uncovered lines read as 0% on the change, not as a rounding error. - **Actionable.** The report points at lines the author just wrote and can still remember. The fix is obvious and local. - **Fair on legacy.** A newcomer to an under-covered repository is asked to leave their own change tested, not to repay decades of debt. - **Convergent.** Because every change now arrives covered, the global figure rises as a **consequence** rather than as a target — which is exactly the relationship that avoids gaming. The usual policy shape is a floor on the changed lines (something like 78% or 80% on the diff, not 100%, so that a genuinely untestable line does not block the change) plus a minimum-lines rule so that one-line changes are not judged on a sample of one. ### Pairing it with a ratchet Diff coverage governs new work; a **ratchet** protects what exists. Record the current global figure as a floor and fail the build if a run comes in below it; when a run comes in above, raise the floor to the new value. The number can then only move upward, without anyone having to pick an aspirational target that the codebase is years away from and that everyone therefore ignores. Two practicalities. Give the floor a small tolerance — a percentage point or so — or ordinary churn produces failures that have nothing to do with the change. And provide a documented override, because legitimate work lowers the figure: deleting a heavily covered module, or absorbing a large third-party file, moves the ratio for reasons no case can fix. ### What diff coverage does not tell you It is a better gate, not a good measure of quality, and a strong answer says so. - It inherits the whole assertion problem. A change can hit 100% on its own lines with cases that check nothing. - It says nothing about the **effect** of the change on code it did not touch. A one-line change to a shared rounding helper can break the fare boundary in a module whose lines are untouched and therefore unmeasured. - Refactors read badly. Moving 600 lines between files marks them all as changed, so a pure move can show poor diff coverage with no behaviour added at all — a case for review-and-override, not for writing cases to satisfy the gate. - It is gameable in its own way: split the risky lines into a change small enough to fall under the minimum-lines rule, or pad a change with trivially covered lines to lift the ratio. ### The interview shape The answer an interviewer is listening for is the sensitivity argument first — the global number cannot detect a change because the denominator is the whole repository — then actionability and fairness, then the ratchet as the complement, then an honest limitation. A candidate who presents diff coverage as a fix for the assertion problem has swapped one misplaced trust for another. It fixes **where** you point the measurement, not **what** the measurement can see.
- What is a coverage ratchet, and what problem does it solve?Record the current global figure as a floor, fail a run that comes in below it, and raise the floor whenever a run comes in higher, so the number can only move upward. It removes the need to pick an aspirational target the codebase is years away from and that everyone therefore ignores, while blocking slow erosion. Give it a small tolerance for churn and a documented override, since deleting well-covered code legitimately lowers the ratio.
- How can a diff-coverage gate be gamed?Split the change so the risky lines land in a change small enough to fall under the minimum-lines rule, pad a change with trivially covered lines to lift the ratio, or widen the report's exclusions so the awkward file is not measured. And the ordinary route is untouched: hit the threshold with cases that execute the new lines while asserting nothing meaningful.
- A pure refactor moves 600 lines between files and the diff-coverage gate fails. What do you do?Treat it as a known false signal rather than a defect. A move or rename marks every line as changed while adding no behaviour, so the ratio is measuring noise. The right response is a reviewed override with the reason recorded, not writing cases whose only purpose is to satisfy the gate — those add run time and maintenance for no verification.
saying these in an interview costs you the question
- Thinks diff coverage fixes the weak-assertion problem
- Sets the diff threshold at 100% with no exceptions
- Cannot explain why the global number is insensitive
- Blocks a two-line fix on repository-wide legacy debt
- Ignores that a refactor inflates the changed-line count