skip to content

A pipeline blocks merges whenever total test coverage falls below 80 percent. What makes a threshold gate like that useful or useless, and what would you gate on instead?

level: middleimportance: should knowfreq 48%

answer

  1. a total average moves too slowly
  2. blocks the bystander, not the offender
  3. cover the lines this change touched
  4. ratchet the floor upward
  5. a gate that fails open is advisory

basics

~20 s

A whole-repository coverage threshold is dominated by legacy code, so one change barely moves it and the gate rarely fires on the change that deserves it. Gate on coverage of the lines the change touched, or ratchet the number so it can never fall.

solid answer

~50 s

Total coverage across a large repository is a slow-moving average: a 300-line untested change might drop it by a fraction of a point, so the gate stays green for exactly the change you wanted it to catch, and then fires months later on someone unrelated who happens to cross the line. Two better shapes: **diff coverage**, which requires the lines this change touched to be covered, so responsibility lands on the author; and a **ratchet**, which stores the current value as the new floor so the number can only improve. The same reasoning generalises to scanners — gate on *new* findings against a recorded baseline rather than on a total that is already in the thousands. Beyond the metric, a gate is only real if it is deterministic, fast, fails closed when the tool itself errors, and has a documented exception path so bypassing is not the routine.

code

bash · 12 lines
bash
#!/usr/bin/env bash
set -euo pipefail

# Without pipefail, the step's status is tee's (always 0),
# so a failing scan would still pass the gate.
security-scan --format json | tee scan-report.json

# Explicit alternative when you must not use pipefail:
# security-scan --format json > scan-report.json
# status=$?
# cat scan-report.json
# exit "$status"

go deeper

for a junior

Know that a quality gate is an automated check that can block a merge or a deploy, and that a failing gate is meant to stop the change rather than be retried until it passes.

for a middle

Explain why a repository-wide threshold misfires — inertia, blocked bystanders, gameable exclusions — and describe diff coverage, ratchets and scanner baselines as the alternatives.

for a senior

Demonstrate gate hygiene in production: determinism, latency on the blocking path, failing closed when the tool errors, and the fact that a routinely bypassed gate is data about the gate, not about the team.

for a principal

Own the portfolio of gates across many repositories: which checks block versus report, how thresholds differ by blast radius, who owns false positives, and how gate latency turns into batch size and deployment risk.

## Why the whole-repository number is the wrong signal Coverage across a large codebase is an average over tens of thousands of lines, most written years ago. Its inertia is the problem. Add 300 untested lines to a 100,000-line repository at 80.4 percent coverage and the number falls to roughly 80.2 — still green. The gate did not fire on the change that earned it. Later, an unrelated small change nudges it to 79.9, and that author is blocked and told to write tests for code they never touched. The gate has inverted: it is silent for the offender and loud for the bystander. There is a second failure. A threshold is trivially satisfiable in ways that produce no value — tests that execute code and assert nothing, generated code excluded or included depending on which way the number needs to move, or configuration edits to the exclusion list. Any single number that people are blocked on will eventually be optimised directly rather than through the behaviour it proxies. ## The two shapes that work **Diff coverage (patch coverage).** Require that the lines added or modified by this change are covered, at some percentage. This puts the requirement on the person who can act on it, at the moment they can act on it, and it is indifferent to how bad the legacy baseline is. It is also naturally proportionate — a one-line fix needs one line covered. **The ratchet.** Record the current value and require that the next run not fall below it, updating the floor upward whenever it improves. A ratchet never blocks anyone for inherited debt, and it makes the number monotonically non-decreasing without anyone negotiating a target. The same pattern applies to lint warning counts, bundle size, and type-checking strictness. A useful third move is to gate different code differently: strict diff coverage on payment or authentication paths, advisory reporting elsewhere. Uniform thresholds are administratively simple and technically indefensible. ## Scanners: baseline, do not total Enable a security or quality scanner on a mature repository and it reports thousands of pre-existing findings. Blocking on the total means nothing merges until a months-long cleanup finishes, so the gate is turned off within a week. Blocking on *findings new since the recorded baseline* keeps the gate on from day one, holds the line, and lets the backlog be burned down deliberately. Severity filtering is the companion control: block on high and critical, report the rest. ## Gate hygiene, independent of the metric Several properties decide whether any gate survives contact with a team. **Determinism.** A gate that fails randomly — flaky tests, a scanner with a network dependency — trains people to re-run until green. Once "just hit retry" is the culture, every gate has been weakened, including the ones that were telling the truth. **Latency.** A gate that adds forty minutes to every merge will be routed around. Push slow analysis off the blocking path: run it on the default branch or nightly, and block only on what is fast enough that nobody resents it. **Failing closed.** This is where gates quietly die. If the tool errors and the step is configured to continue on error, a broken scanner and a clean scan are indistinguishable. Shell pipelines hide this too — the exit status of a pipeline is the status of its *last* command, so piping a scanner into a formatter discards the scanner's verdict unless `set -o pipefail` is set or `PIPESTATUS` is inspected. **Blocking versus advisory, chosen deliberately.** Every gate should be one or the other on purpose. Advisory results that are labelled as required, and required results that everyone knows are ignorable, are both worse than an honest choice. **An owner and an exception path.** Someone must own each gate's threshold and its false positives, and there must be a documented, recorded way to proceed without it. If there is no sanctioned exception, people invent an unsanctioned one, and the difference between the two is whether anyone can count how often it happens. ## What to say in an interview The strong answer is not "coverage gates are bad". It is: pick a metric the author of the change can act on, baseline what you inherited rather than blocking on it, make the gate fast, deterministic and fail-closed, and be explicit about which gates block and which merely report.

  • You enable a scanner on a mature repository and it reports 4,000 findings. How do you make it a gate this week rather than next year?
    Record the current findings as a baseline and block only on findings not in it, so the gate is enforceable immediately while the backlog is unchanged. Filter by severity so high and critical block and the rest report. Then burn the baseline down deliberately, shrinking it on a schedule rather than blocking merges on it.
  • Why does a flaky blocking gate weaken gates you did not touch?
    Because the team's response to a red check becomes "re-run it" rather than "read it". Once retry-until-green is the habit, a genuine failure gets the same treatment as a flake. Flakiness in one gate is therefore an availability problem for the whole set, which is why quarantining flaky tests matters more than it looks.
  • What would make you choose an advisory report over a blocking gate for a given check?
    When the check is slow, noisy, or its verdict needs human judgement — style opinions, performance drift, a scanner with a high false-positive rate. Publish it as a report or a comment and watch the trend. Promote it to blocking only once it has been quiet and correct long enough that blocking would rarely be wrong.

saying these in an interview costs you the question

  • Treats total coverage percentage as a quality measure
  • Blocks merges on a scanner's pre-existing finding backlog
  • Configures gate steps to continue on error so releases are not held up
  • Assumes a piped command's failure fails the shell step
  • Adds gates without an owner or a recorded exception path

context