skip to content

What is a severity threshold in a pipeline security gate, and what does it actually decide?

level: juniorimportance: must knowfreq 72%

answer

  1. turns a report into pass or fail
  2. a line someone drew, not a fact
  3. severity form versus count form
  4. at or above a chosen level
  5. coarse proxy, valued for uniformity

basics

~20 s

A severity threshold is the rule that collapses a scan report into pass or fail: block when a finding sits at or above a chosen severity, such as critical. It decides what stops a change, not what is genuinely risky.

solid answer

~50 s

A scanner emits findings, each carrying a severity label such as critical, high, medium or low. A pipeline gate needs one boolean, so the policy draws a cut: fail if any finding is critical, or fail if there are more than N highs. That cut is a policy choice, not a property of the report — the same findings pass or fail depending on where someone drew the line. Two shapes exist. A *severity threshold* blocks at or above a level; a *count threshold* blocks above a number at a level. The severity form is the more honest one, because the count form invites a team to sit just under the number instead of fixing anything. The point to make in an interview is that the threshold decides what blocks, and only loosely correlates with what would actually hurt you in production.

go deeper

for a junior

Be able to state that the threshold is the rule turning a findings report into pass or fail, name the usual severity buckets, and say plainly that green means nothing crossed the line rather than nothing is wrong.

for a middle

Explain the difference between a severity cut and a count cut, and why the count form drifts and invites gaming. Show that the same report yields different verdicts purely from where the line was placed.

for a senior

Demonstrate that you have operated one of these. Talk about the blocking lane being narrow and defensible while the reporting lane is wide, and about what happens socially when the first build blocked is one nobody can fix.

for a principal

Own the argument that a threshold buys uniformity rather than accuracy, and that its real cost is the credibility spent every time it fires on something the team cannot act on.

## The two halves of a gate A pipeline security gate has a detective half and a preventive half. The scanner is detective: it runs, inspects an artifact, a manifest, a plan or a dependency set, and produces a **report** — a list of findings, each with a rule identifier, a location, and a severity label, usually drawn from a scale like critical / high / medium / low / informational. A report on its own changes nothing; it observes. The gate is the preventive half, and a pipeline step can only do one of two things: continue or stop. So something must collapse a list of findings into a single boolean. That something is the **threshold**, and it is written by a person. ## The two shapes of threshold **Severity threshold.** *Fail if any finding is at or above critical.* One dial: where the line sits. It is easy to state, easy to explain to a blocked developer, and easy to audit — anyone can read the report and reproduce the verdict. **Count threshold.** *Fail if there are more than ten high findings.* Two dials: the level and the number. This shape looks pragmatic and behaves badly. It aggregates unrelated problems into one integer, so the number tells you nothing about which ten. It rewards hovering: a team at nine highs has every incentive to stay at nine rather than drive toward zero, and closing one finding buys headroom to add another. And the number drifts on its own — when the scanner's rule set grows or an advisory feed adds entries, the count moves without anyone touching the codebase, so the gate's meaning changes while the policy file stays identical. ## The threshold is chosen, not discovered This is the sentence to be able to say out loud. Given one fixed report, a gate at *critical* passes and a gate at *high* fails. Nothing about the artifact changed between those two runs; the only difference is where a policy author put the line. A candidate who talks about the threshold as if it were a measurement of risk has missed what the interviewer is probing. ## Why the number is a proxy, not the risk Severity labels are assigned by whoever wrote the rule or published the advisory, generically, with no knowledge of how you deploy. Your gate inherits that rating wholesale. A finding rated high in a component that only runs at build time may be unreachable in production; a medium in an internet-facing path may matter far more. Deciding which findings genuinely deserve attention is triage, and that is a separate discipline — the gate's job is narrower and more mechanical: apply the stated line consistently to every change, so the verdict is predictable and arguable. Saying this plainly is the mature answer. A threshold is a coarse, cheap, uniform filter. Its value is that it is uniform, not that it is accurate. ## What green does and does not mean A green threshold gate means: *no finding in this report crossed the line we drew*. It does not mean the artifact is free of vulnerabilities. It does not mean anyone looked at the findings below the line. It does not mean the scanner covered everything — it saw only what it was pointed at, with the rules it had on that day. Overclaiming here is one of the most common weak answers in the domain. ## Why the first threshold is usually wrong The instinctive first policy is *zero criticals*, because it sounds strict and nobody argues with it in the meeting. Applied to an existing codebase that already carries criticals, it fails every build immediately, including builds whose changes had nothing to do with the finding. What follows is predictable: the gate is downgraded to a warning, disabled, or routed around. A threshold that has no answer for the backlog that predates it is a threshold that will be switched off. ## What a workable threshold looks like - **Narrow and defensible in the blocking lane.** A small set of conditions you are willing to defend to a team at 5pm on a Friday. - **Wide in the reporting lane.** Everything else is recorded and visible without stopping anyone. - **Scoped to what the change introduced** rather than to the absolute state of the repository, so the person who is blocked is the person who can fix it. - **Accompanied by a written answer** to 'this fired on me, now what' — who to ask, what the exception path is, how long a fix is expected to take. In an interview, land three things: the threshold turns a report into a verdict; it is a choice rather than a measurement; and its number is a proxy for risk that must be defended socially, not just configured.

  • Why is 'no more than ten high findings' a weaker rule than 'no critical findings'?
    A count threshold aggregates unrelated problems into one integer, so nobody learns which ten matter. It rewards sitting just under the number instead of driving toward zero, and closing one finding creates room to introduce another. Worse, the count drifts when the scanner's rule set or advisory feed grows, so the gate's behaviour changes without anyone editing the policy.
  • The severity labels are assigned by the scanner, not by you. What does that do to your threshold?
    Your gate inherits someone else's rating. A rule reclassification or a feed update can push a batch of findings up a level overnight and start blocking builds when nothing in the code changed. Treat a sudden mass shift as a tooling event, investigate it as such, and design the gate so a single upstream re-rating cannot halt every team at once.
  • Does the scan make the pipeline preventive, or does the gate?
    The gate does. Scanning is detective — it observes and reports what is present. Enforcement is what makes the control preventive: the pipeline stops the change from reaching the next stage. Without a gate you have a report nobody acts on; without a scan you have nothing to gate on. Both halves are needed, and only one of them blocks.

It is a speed limit, not a measure of how dangerous a particular driver is. Its usefulness comes from being the same number for everyone, not from being right about any one car.

saying these in an interview costs you the question

  • Treats the threshold as an objective measurement of risk
  • Says a green gate proves the artifact has no vulnerabilities
  • Picks zero criticals without checking the existing backlog
  • Assumes severity labels are stable and comparable across tools
  • Cannot state who is blocked or what they should do next

context