CodeQL is enabled across your org and produced thousands of alerts nobody triages. How do you make it trustworthy?
answer
- Volume is not the metric
- Two problems, not one
- Fix the input, not the list
- Dismissal reasons are evidence, not cleanup
basics
~20 sSeparate the historical backlog from new findings: gate only on alerts a pull request introduces, tune the configuration so recurring false positives stop firing, require a dismissal reason plus comment, and measure the false-positive rate per rule rather than the raw alert count.
solid answer
~50 sAlert count is the wrong metric; **trust** is the metric, and it is destroyed by a rule that is wrong three times in a row. First, split the problem in two. The backlog is a debt project with owners and deadlines. New findings on pull requests are the ratchet, and only that side gets a merge gate — that way scanning is adoptable on a legacy codebase. Second, tune at the configuration, not by mass-dismissing. Use `paths-ignore` for generated and vendored code, `query-filters` to exclude a rule id that is structurally wrong for your stack, and keep the default query suite until someone is actually triaging before widening to `security-extended`. Third, make dismissal honest: `false positive`, `won't fix`, `used in tests`, always with a comment. That comment is your evidence base — when one rule accounts for most "false positive" dismissals, you have a configuration bug, not a security backlog. Finally, assign ownership per repository. Alerts with no owner are guaranteed to rot.
code
yaml · 12 linesname: "CodeQL config"
paths-ignore:
- '**/generated/**'
- '**/build/**'
- src/test
query-filters:
- exclude:
id: java/spring-disabled-csrf-protection
- exclude:
tags contain: maintainabilitygo deeper
Know that alerts can be dismissed with a reason — false positive, won't fix, used in tests — and that dismissing without reading is how a scanner stops protecting anything.
Explain the tuning surface: paths-ignore for generated code, query-filters for a rule that does not fit the codebase, and why query suite choice drives noise.
Show you would gate only alerts a pull request introduces, run the backlog as owned work, and treat repeated false-positive dismissals of one rule as a configuration bug.
Own the trust argument and the metrics: precision per rule, time to triage, coverage honesty for compiled languages, and a rollout that observes before it enforces.
## The real failure mode A scanner that fires mostly false positives does not produce a slightly worse outcome than a good scanner — it produces a *worse than nothing* outcome, because it teaches the team a reflex: see red, dismiss, merge. Once that reflex exists, the true positive that eventually appears is dismissed with the same two clicks. Every decision below follows from that single point. ## Separate the backlog from the ratchet GitHub's own model gives you the split for free: pull request annotations and the code scanning results check concern alerts the diff introduces, while the full alert list lives in the Security tab. Use it. - **New alerts** are a review-time concern. This is where a merge gate can go, because the volume is small, it is attributable to a specific change, and the author has context loaded. - **The backlog** is a programme of work: triage by severity and reachability, assign owners per repository or service, set a schedule, and track burn-down. GitHub has been adding organisation-level tooling for exactly this backlog-driving problem, but the tool is not the hard part — the ownership is. Never gate the merge button on the backlog. That converts adoption into a hostage situation and guarantees the gate gets disabled. ## Tune the configuration, not the alert list Mass-dismissing is a treadmill: the finding returns next quarter in a new file. Fix the input. - **`paths-ignore`** in the CodeQL configuration file removes generated code, vendored third-party trees and fixture data from analysis. Findings in generated code are almost never actionable, and they are usually the bulk of a first-run backlog. - **`query-filters`** exclude a specific rule id whose assumptions do not hold in your codebase — a taint source your framework sanitises centrally, say. Excluding it in configuration is honest, reviewable and applies everywhere; dismissing its findings one by one is neither. - **Query suite choice.** Start at the default suite. `security-extended` trades precision for coverage and `security-and-quality` adds maintainability rules; both are appropriate *after* there is a functioning triage habit, not before. Widening the suite on an untriaged repository is the single most common self-inflicted noise wound. - **Build correctness.** For compiled languages, a partial build produces both missing coverage and odd findings. Verify the database actually covers the code before concluding a rule is noisy. ## Make dismissal carry information Code scanning offers three dismissal reasons, and they mean different things: - **False positive** — the tool is wrong. This is a *configuration* signal. - **Used in tests** — the code is real but the risk context does not apply. This is a *path filter* signal. - **Won't fix** — accepted risk. This is a *governance* signal, and the one that should require a named accepter and a comment. Require a comment on every dismissal. The comments are your evidence base: aggregate them and you can say "rule X accounts for 60% of our false-positive dismissals", which is an actionable configuration change rather than a feeling. Without comments, all you have is a number going down, which is indistinguishable from people clicking through. ## Measure the right things Replace "open alerts" with metrics that reflect trust: - **Precision per rule** — dismissals as false positive divided by alerts raised, per rule id. Rules above a threshold get excluded or fixed. - **Time to triage** for new alerts on pull requests. If it is days, the gate is not working; the author has moved on. - **Dismissal-to-fix ratio** by team. A team that only ever dismisses is either drowning or unowned. - **Coverage** — repositories with scanning enabled, and for compiled languages whether the analysis genuinely covers the build. Silent under-coverage flatters every other number. ## Sequence the rollout Enable analysis broadly but non-blocking. Watch for a few weeks, tune out the top noise rules, then require the check for new alerts on one important repository, then widen. Assisted remediation — GitHub's Copilot Autofix can propose fixes directly on code scanning alerts — helps the fix side of the ratio, but it does not rescue a badly tuned rule set; a confident suggested fix for a false positive is worse than the alert alone. ## What a principal answer sounds like Ownership model, configuration as the tuning surface, dismissal reasons as evidence, gate only the delta, and metrics tied to trust rather than volume. The tell of a weak answer is a plan whose first step is "triage all nine thousand alerts": that plan has been tried in every organisation that ever asked this question, and it has never once finished.
- Why not simply require every existing alert to be resolved before enabling the gate?Because on a real codebase that plan never finishes, and until it does nothing merges. Gate the delta instead: alerts a pull request introduces are few, attributable and reviewable while the author has context. The historical backlog is separate, owned work with a schedule.
- What is the difference between dismissing an alert and excluding its rule in configuration?A dismissal closes one finding at one location; the rule fires again in the next file. An exclusion in the CodeQL configuration is repository-wide, lives in a reviewed file, and states the reasoning once. Use dismissal for a genuinely one-off judgment, configuration for a rule that is wrong for the codebase.
- Does assisted fix generation solve the noise problem?No. Suggested fixes shorten remediation for genuine findings, which improves the fix-versus-dismiss ratio. They do nothing about precision — a plausible-looking suggested fix attached to a false positive is worse than the bare alert, because it invites someone to change working code.
saying these in an interview costs you the question
- Planning to triage the whole backlog first
- Gating merges on all pre-existing alerts
- Mass-dismissing findings without recording a reason
- Turning on security-extended to look thorough
- Reporting alert count as the security metric