skip to content

How do you set the precision-versus-recall operating point for an org-wide SAST program?

level: principalimportance: should knowfreq 44%

answer

  1. The scarce resource is attention
  2. Start from who reads it
  3. Two audiences, two tolerable error rates
  4. Credibility spends down and does not return
  5. Measure coverage, not just findings

basics

~20 s

Set it per audience, not per tool. Findings that interrupt a developer must be near-certain; a broad, noisier stream belongs in a reviewed queue with a funded owner. Then measure per-rule true-positive rates, triage age and the analyzer's own coverage.

solid answer

~50 s

The binding constraint is human attention, so start from who reads each finding, on what clock, with what context. That gives two tiers: a **small, measured, high-precision set** whose findings are expected to be genuine and are handled with the change, and a **broad, lower-precision set** routed to a review queue with a named owner and a triage budget. Rules earn the first tier from a sampled per-rule true-positive rate and lose it when that rate falls; demoting a rule is legitimate, because credibility is finite - once people believe the analyzer cries wolf, its effective recall is zero. Measure triage rate and age, escaped defects found some other way, and **analyzer health** - files skipped, rules that failed to load, timeouts - because a green run over a third of the tree manufactures confidence. Then write down the categories you chose not to catch, with their compensating controls.

go deeper

for a junior

Be ready to say that a noisy security check gets ignored, and that ignoring it is the real failure - the value of a finding depends on somebody actually reading and acting on it.

for a middle

Explain precision and recall as a dial with no universally right setting, and why the same rule set can be right for a reviewed queue and wrong for something that interrupts a change.

for a senior

Show you would measure before promoting a rule to anything blocking, close false-positive families with model work rather than suppressions, and audit the analyzer's own coverage instead of trusting a clean dashboard.

for a principal

Own the whole tradeoff: audiences and tiers, funded triage capacity against real headcount, rule promotion and demotion driven by sampled data, escaped-defect review as your only honest recall signal, and a written list of accepted gaps with their compensating controls.

### The decision, stated properly Every security analyzer has a dial. Turn it toward **recall** and it reports more of the real defects, along with far more that are not defects. Turn it toward **precision** and nearly every report is genuine, at the price of quietly missing whole families. There is no setting that is correct in the abstract, because the constraint is not the analyzer - it is **the finite attention of the people who read the output**. Choosing the operating point is a capacity decision, and it is the part of a security-analysis program that a lead actually owns. ### Start from the reader, not from the rules Ask who is meant to act on each finding, on what clock, and what they are allowed to do about it. A developer looking at their own change in review has minutes and no security context; they can act only on findings that are almost certainly real and almost certainly theirs. A security reviewer working a queue has hours and full context; they can profitably sift a noisy stream. Those are two different audiences with two different tolerable precisions, and a program that sends one stream to both fails whichever audience it did not design for. That gives the standard two-tier shape: a **small, high-precision blocking set** whose findings are expected to be true and are fixed as part of the change, and a **broad, lower-precision advisory set** that lands in a review queue with an owner, an SLA and a triage budget. The interesting design work is deciding which rules qualify for the first tier, and the answer must come from measurement rather than from the vendor's confidence label. ### Credibility is a finite, non-renewable resource The failure mode that kills these programs is not a missed defect; it is a team that has learned the tool cries wolf. Once developers form the belief that findings are noise, they stop reading them, and recall on the blocking tier effectively goes to zero no matter what the analyzer reports. This is why precision on the tier that interrupts people is worth more than raw coverage, and why moving a rule *out* of the blocking tier is a legitimate, often correct action rather than an admission of defeat. ### Measure the program, not the findings count Total findings is a vanity number - it moves with codebase size and rule-set churn. Useful measures: * **Per-rule true-positive rate**, from a sampled triage of a fixed number of findings per rule. Rules below an agreed floor leave the blocking tier automatically. * **Time to triage and time to fix**, split by tier. A queue whose age is growing is a program that has already exceeded its capacity. * **Escaped-defect review.** When a security defect is found by any other means - a review, a penetration test, a report from outside - ask whether the analyzer could have found it and why it did not. That is the only honest estimate of recall you will ever get, because the false negatives are by definition invisible to the tool. * **Analyzer health.** Rules that failed to compile, files skipped, functions that hit the timeout, the share of the tree actually analysed. A green run that analysed a third of the code is worse than a red one, because it manufactures confidence. That last point is not theoretical. On a ride-hailing dispatcher's estate, the incremental-analysis cache was keyed on file modification timestamps, and a clock-skew artefact between two build agents - about 90 seconds apart - meant freshly changed files sometimes looked older than the cached result and were skipped. For 19 days the scan reported a healthy 0 new findings on a module that was being rewritten. Nothing in the findings dashboard could have shown that; only a coverage measure could. ### Make the accepted false negatives explicit A mature program writes down what it has decided *not* to catch this way - authorisation and business-logic defects, anything reached through dynamic wiring, whole languages or repositories with no usable rule set - and names the compensating control for each: design review, a targeted manual assessment, runtime testing, a monitoring detection. An undocumented gap becomes, over time, a belief that the gap does not exist. ### Fund the triage, and place it where the knowledge is Precision is also an economic choice about somebody's week. An 11-person team cannot absorb a thousand-item advisory queue on top of delivery, and pretending otherwise produces a queue that is formally owned and actually abandoned. Either the volume is cut to the capacity, or capacity is added - a rotating triage duty, a security partner, a budget for model work. And the model work is the compounding investment: every false-positive family closed by declaring a team validator or narrowing a sink rule raises precision permanently, for every future scan, which is what lets you move the dial back toward recall without spending more attention. ### The framing to leave the interviewer with Set the operating point per audience; buy precision on the interrupting tier with rules you have measured; keep recall on a tier somebody is genuinely funded to read; measure the analyzer's own coverage as carefully as its findings; and write down what you have chosen not to catch.

  • What single metric would you refuse to run the program on, and why?
    Total open findings. It moves with codebase size, rule-set churn and duplicate reporting, so it can fall while risk rises or climb after an improvement. It also rewards suppression. Prefer per-rule true-positive rate from sampled triage, time to triage and fix by tier, escaped defects found by other means, and coverage of what was actually analysed. Those answer whether the program is producing decisions, which the raw count never does.
  • How would you ever estimate recall, given that false negatives are invisible to the tool?
    By working backwards from defects found some other way: a design review, a manual assessment, an external report, a production incident. For each, ask whether the analyzer could in principle have found it and, if so, why it did not - unmodelled source, run-time wiring, a truncated path, a rule that was never enabled. Seeded-defect exercises give a rougher second estimate. Both are samples, so report them as trend and gap analysis, never as a coverage percentage.
  • A dashboard shows zero new findings on a service for weeks. What would you check before believing it?
    Whether the analysis actually ran over that code. Check files analysed against files changed, rules that failed to load, functions that hit the timeout, and whether an incremental cache decided the work was already done. Caches keyed on file timestamps are a known trap when build agents disagree about the clock: freshly changed files look older than the cached result and are skipped, and the dashboard reports a healthy zero the whole time.

It is alarm design: an alarm that wakes the whole building must almost never be wrong, while a monitoring console someone is paid to watch can afford to show maybes - wire both to the same siren and people start sleeping through it.

saying these in an interview costs you the question

  • Turns on every rule and calls it thorough
  • Blocks changes on rules with unmeasured precision
  • Reports total findings as the program metric
  • Treats a queue nobody is funded to read as coverage
  • Never demotes a rule out of the blocking tier
  • Assumes a green run means the code was analysed

context