Developers ignore a lint run that emits thousands of warnings - how do you make the ruleset trustworthy again?
answer
- Attention is the scarce resource
- Group findings by rule, not file
- Sample findings and classify them
- Off or enforced, never permanently warned
- Narrow by path before deleting outright
basics
~20 sTreat findings as an attention budget: sample each rule to see how often it is right, narrow or delete the ones that cry wolf, and give every survivor one decision - off, or enforced. A permanent warning heap teaches people to ignore everything.
solid answer
~50 sA ruleset nobody reads is worse than no ruleset, because the learned habit of skipping the output also skips the one finding that mattered. I would start by sorting findings **by rule rather than by file** - noise is almost always concentrated in a handful of rules. Then measure instead of arguing: sample a few dozen findings from each loud rule, have two engineers classify each as 'we would change this' or 'we would not', and set an explicit bar, say right in at least 92 of 100 sampled findings, for a rule to stay enforced. Rules below it get narrowed by path if they are only wrong about generated or vendored code, and deleted otherwise. Finally, remove the middle severity as a resting place: every rule ends up off or enforced, because a rule that has warned for eighteen months has already been decided against, just never in writing.
code
pseudocode · 9 linesfindings = run_analysis(repository) // 4317 findings, 63 rules
by_rule = group(findings, key = rule)
for rule in top_n(by_rule, 4): // these 4 produced 3880
sample = random_choice(by_rule[rule], 40)
verdict = two_reviewers_classify(sample) // "would change" / "would not"
rate = verdict.would_change / 40
if rate >= 0.92: enforce(rule)
else if wrong_only_in(rule, "generated/"): narrow(rule, exclude="generated/")
else: disable(rule)go deeper
Understand the basic dynamic: if output is too big to read, people stop reading it, and then genuine findings are lost too. Be able to say that fewer, better rules beat more rules.
Explain the mechanics you would use: group findings by rule to find the concentrated noise, distinguish narrowing a rule's scope from disabling it, and describe why the middle severity level tends to accumulate rules nobody acts on.
Bring a method and numbers. Show how you would sample findings, classify them with two reviewers, set an explicit accuracy bar, and drive each rule to a written decision - including the ones you delete despite them being technically correct.
Own the cost model and the governance: rules have owners, a review cadence, and a budget they must pay for in prevented defects. Be ready to say honestly that the link between finding density and defect density is contested, and to defend a small enforced set to people who want everything on.
### The failure being described A ruleset that produces more findings than anyone can act on does not produce zero value — it produces negative value. People learn, correctly, that the output is not worth reading, and that learned reflex applies to the one finding a year that would have mattered. Warning fatigue is not laziness; it is a rational response to a channel with a bad signal-to-noise ratio. So the fix is never "tell people to read the warnings". The fix is to make the channel worth reading. ### Attention is the budget Treat findings as spending from a fixed, shared attention budget. Every enabled rule buys a claim on that budget and must pay for it in defects prevented. This reframes the ruleset from a catalogue of everything the tool can check into a *curated set* where each rule has an owner and a justification. A rule that has never caught anything anyone cared about is not neutral; it is consuming the budget that the leak-detection rule needs. ### The three-state model, and why the middle state rots Most rule engines offer three states per rule: off, report as a warning, report as an error. The third state interrupts; the second does not. In practice the middle state is where rules go to die: a rule that has warned continuously for eighteen months has been decided against by the organisation, just never in writing. So force the decision. Each rule ends up either **off** (nobody is going to act on it, so stop printing it) or **enforced** (we mean it, and an occurrence stops the change). "Warning" survives only as a temporary, dated state while a rule is being measured — never as a resting place. Note that the severity you assign a rule and the technical severity of the defect it finds are different things: a rule that finds a low-impact issue with 99% precision may be worth enforcing, while a rule that finds a scary issue one time in ten is not. ### Measure precision before you argue about it Ruleset arguments run on anecdote unless you bring numbers. A concrete audit that finishes in an afternoon: take a run over the appointment scheduler — 4,317 findings from 63 enabled rules — and first sort by rule, not by file. Four rules produced 3,880 of the findings; the remaining 59 produced 437 between them. Sample 40 findings at random from each of the top rules, have two engineers independently classify each as "a thing we would change" or "not a thing we would change", and compute the rate. Now the conversation has evidence, and you can set an explicit budget: a rule may be enforced only if it is right in at least 92 of 100 sampled findings. Two of the four noisy rules came in near 12% and were deleted the same day; one came in at 96% and was promoted; the fourth was right almost always in application code and almost never in generated code. ### Scope before you disable That last case is the important one, because "the rule is noisy" and "the rule is wrong" are different diagnoses. Rulesets support per-path overrides: a base configuration for the repository, with narrower configurations layered over specific directories. Generated sources, vendored third-party code, database migration scripts and test fixtures each have legitimately different conventions, and a rule that is right about production code and wrong about generated code should be *narrowed*, not deleted. Keep the override list short and readable, though — a configuration with 40 path exceptions is its own kind of noise, and each exception should say why it exists rather than merely that it does. ### Volume matters independently of precision A rule that is right 92% of the time but fires 900 times has produced 72 wrong findings, which is 72 arguments and 72 chances to edit correct code. Precision and volume multiply. When a high-precision rule is simply loud, the right move is usually to narrow what it looks at rather than to accept the flood. ### What to say about evidence Be honest about the parts that are contested. Whether the density of static-analysis findings predicts defect density in a codebase has been studied repeatedly with mixed results, and the effect depends heavily on which rules are enabled. So do not claim that a lower warning count means better software. The defensible claim is narrower and stronger: a small, high-precision, enforced set changes behaviour, and a large unread set does not. ### The end state Under twenty enforced rules, each with a named owner and a one-line reason. Everything else off. Findings on a typical change in the low single digits, so the number is small enough that a person reads it rather than skims it. New rules enter in a measured, dated trial and leave it in one direction or the other. And the ruleset gets an explicit review on a cadence, because a rule that earned its slot two years ago may be checking for a mistake the current design makes impossible. ### The anti-pattern to name The other common response to a warning heap — disabling analysis entirely, or leaving it running with the build ignoring it — is worse than the heap, because it removes the one channel that catches the class of defect tests structurally miss: the wrong behaviour on the branch nobody exercised.
- How do you measure a rule's true-positive rate without auditing every finding it produced?Sample. Take a few dozen findings at random from that rule - not the first few dozen, which cluster in one file - and have two engineers classify each independently against the question 'would we change this code?'. Disagreements are the interesting cases and usually reveal that the rule is really two rules. A sample of that size gives a rate precise enough to decide keep, narrow or delete, and it finishes in an afternoon.
- A rule is right 92% of the time but fires 900 times. Do you enforce it?Not as it stands. Precision and volume multiply: 900 findings at 92% is 72 wrong ones, which is 72 arguments and 72 chances to change correct code. The usual move is to narrow what the rule looks at - a directory, a construct, a specific context - until the volume is proportionate, and only then enforce it. Loud and accurate is still loud.
- Why not simply leave the noisy rules reporting as warnings and ask people to skim them?Because that is exactly the state that produced the problem. A permanent warning stream trains everyone that the channel is optional, and the training generalises to the rules that matter. The middle severity is only defensible as a temporary, dated state while a rule is being measured; as a resting place it is a decision against the rule that nobody has written down.
- How do you decide which rules deserve a severity that actually interrupts someone?High precision, a repair that is clear from the finding, and a consequence that is genuinely worth stopping for. Note that the rule's severity and the technical severity of the defect it finds are different things: a low-impact issue detected with near-perfect accuracy can be worth enforcing, while a frightening issue detected correctly one time in ten is not.
A smoke alarm that goes off every time someone makes toast does not make the kitchen safer. It teaches the household to take the battery out, and then the real fire is undetected.
saying these in an interview costs you the question
- Leaves every rule at warning so nothing is ever decided
- Disables all analysis instead of the few noisy rules
- Judges a rule by findings produced, not defects prevented
- Claims a high finding count proves a low-quality codebase
- Enables every rule a shared ruleset offers because it exists
- Argues from anecdote without sampling the actual findings