What is a static-analysis baseline, and why create one when adopting an analyser on an existing codebase?
answer
- The blocker is the backlog, not the rules
- Record what already fails, once
- Rules stay on for new code
- Deferred debt, not forgiven debt
- It may only shrink
basics
~20 sA static-analysis baseline is a recorded snapshot of the findings an analyser already reports on existing code. The gate ignores those and fails only on findings outside the snapshot, so an old codebase can adopt analysis without a mass cleanup first.
solid answer
~40 sSwitching an analyser on in a mature codebase usually produces thousands of findings at once, which is an adoption problem rather than a code problem. A baseline records those findings - rule identifier, location and a fingerprint per finding - and the gate then treats any finding that matches an entry as already known, failing the build only on findings that match nothing. The rules stay enabled everywhere, so new and modified code is held to the full standard from the first day while the historical backlog is deferred rather than forgiven. The baseline is checked in and reviewed like any other file, which makes the accepted debt visible and countable. Two disciplines keep it honest: it may only shrink, and it is never regenerated just to turn a red build green.
code
pseudocode · 14 linesbaseline = load_baseline("analysis-baseline")
findings = analyser.run(source_tree)
new_findings = []
for f in findings:
key = fingerprint(f.rule_id, f.file, f.code_element_hash)
if key not in baseline:
new_findings.append(f)
report(new_findings)
if len(new_findings) > 0:
fail_build("analysis: " + len(new_findings) + " finding(s) not in baseline")
else:
pass_build()go deeper
Be ready to say in one sentence what a baseline holds and what the gate does with it: known findings are excused, anything unmatched fails. Knowing that the rules stay switched on for new code is the half candidates most often miss.
Expect to explain the mechanics - what an entry stores, how a finding is matched to it, and why a list of findings behaves differently from a bare per-rule count when a violation is fixed and another added.
Show the judgement about what belongs in the baseline: triage by rule, fix the probable-defect findings during adoption, baseline the cosmetic remainder, and name the habits that keep the register honest under delivery pressure.
Own the policy. Decide what the baseline may contain, who reviews additions, whether it is allowed to grow at all, and how adoption is sequenced across many repositories so teams get a gate quickly instead of a cleanup project they will never schedule.
## The adoption problem Static analysis is easy to add to a new project and hard to add to an old one. The rules encode a standard the code was never written against, so the first run measures the accumulated distance between the code as it is and the standard as it is now. On a mature service - a warehouse stock ledger with years of history, say - that first run can emit several thousand findings. Three responses exist and two of them fail in practice: 1. **Fix everything before the analyser is merged.** A stop-the-world cleanup that nobody funds, producing an enormous diff across code whose test coverage is usually weakest exactly where the findings cluster. 2. **Run report-only forever.** The findings scroll past in the build log, nobody reads them, and fresh violations arrive at the same rate as the old ones are ignored. 3. **Record what already fails and hold only new work to the standard.** This is the baseline, and it is the only option that gets the gate switched on this week. ## What a baseline actually is A baseline is a file, stored with the code and versioned with it, listing the findings the analyser reported at a chosen moment. A typical entry carries the rule identifier, the location, and a fingerprint of the finding - a hash over the code element or its surrounding text rather than a bare line number, so that inserting an unrelated line above a finding does not make it look new. On every later run the analyser produces its findings, matches them against the baseline entries, and reports only the unmatched ones. The gate fails on those, and only those. The granularity varies between tools: some baselines list every individual finding, others store a per-rule or per-module count that the build must not exceed. The list form is stricter, because it can tell a fixed finding from a newly introduced one at the same site; the count form is cheaper to store and merge but cannot. ## What it is not - **Not the same as disabling a rule.** The rule stays enabled everywhere. A site not listed in the baseline is still checked, which is the entire point: new code gets the full standard immediately. - **Not the same as an inline suppression.** A suppression lives in the source at one site, is written deliberately, and should carry a reason. A baseline is external, bulk and blind - generated in one command, not argued for finding by finding. - **Not a fix.** Nothing in the code changed. The baseline is a debt register, and reading its size honestly is part of the discipline. ## Hygiene The baseline is only useful while it is trusted, and a handful of rules keep it trustworthy. It may only shrink: entries disappear as the underlying code is fixed, and the file is never regenerated to make a red build green - that single habit converts the whole mechanism into an amnesty. Its growth is reviewable, because it is a checked-in file: a change that adds baseline entries shows up in review as clearly as a change that adds code, and the right reviewer question is why the finding was not simply fixed. And not everything deserves baselining. A rule that describes style debt is a fair candidate; a rule that flags a probable defect - an unchecked cast, a value that can be absent, an ignored result - should be fixed or raised as a defect rather than filed in a register nobody rereads. ## A worked adoption On the stock-ledger service the first analyser run reported 4,137 findings across 26 rules. The team triaged by rule rather than by file: 14 findings from two correctness rules, all around unchecked reads of stored ledger records, were fixed inside the adoption change; the remaining 4,123 were baselined and the gate was switched on the following morning. Across the next four three-week release trains the baseline fell to roughly 3,400 through opportunistic fixes made while touching files for other reasons, and no build ever had to be unblocked by regenerating it. The numbers are illustrative, but the shape - fix the small dangerous slice, baseline the large cosmetic one, gate immediately - is the standard recipe. ## Failure modes to expect Renaming or moving a file can unmatch every entry it owned, so a few hundred old findings suddenly present as new; teams handle moves as a deliberate, separately reviewed baseline update. Two long-lived branches that both regenerate the baseline conflict noisily, which is an argument for regenerating it rarely and on purpose. Coarse fingerprints let a stale entry survive after the code it described was rewritten, silently excusing a fresh finding at the same site. And the social failure is the most common of all: once one inconvenient finding has been filed away without discussion, the baseline becomes the place difficult things go, and its size stops meaning anything.
- How is baselining different from just turning off the rules that fire the most?Turning a rule off removes it everywhere, including from code written tomorrow, so the standard is permanently lowered and nobody can see what was given up. A baseline keeps the rule enabled and excuses a named, countable set of existing sites; every new occurrence still fails the gate, and the register of excused sites is visible in the repository and shrinkable.
- What stops a baseline from becoming a permanent amnesty?Three habits. It may only shrink, so entries leave as code is fixed and none are added casually. It is never regenerated to unblock a red build - that is the failure that empties the mechanism of meaning. And additions to it are reviewed like code changes, so a reviewer gets the chance to ask why the finding was not simply fixed.
- Are there findings you would refuse to baseline at all?Yes. Findings from rules that describe probable defects rather than style - an unchecked cast, an ignored return, a resource that is never released - should be fixed in the adoption change or raised as defects with owners. Filing them in a debt register makes them invisible while leaving the bug in production code, which is the worst of both options.
It is like taking a dated photograph of a rented flat when you move in: the existing scuffs are recorded so you are not charged for them, but every new mark is yours - and the list only means anything if nobody quietly retakes the photo later.
saying these in an interview costs you the question
- Thinks a baseline fixes or repairs the findings it lists
- Regenerates the baseline whenever the build turns red
- Says baselining a rule means the rule is off everywhere
- Baselines probable-defect findings alongside cosmetic ones
- Keeps the baseline out of version control, so growth is invisible
- Treats shrinking the baseline as optional busywork