skip to content

Risk and Confidence

Every alert carries two separate scores, and a scan rule has two separate dials that feed them. Interviewers ask because mixing up the score and the dial is the classic scanner-tuning mistake.

on this pageshow

questions

4

On an OWASP ZAP alert, what do the risk value and the confidence value each tell you?

level: juniorimportance: must knowfreq 70%

answer

  1. two scores, not one
  2. one is severity, one is certainty
  3. four risk rungs, five confidence rungs
  4. core's Alert object carries both

basics

~10 s

Risk is how damaging the finding would be if it is real. Confidence is how sure the scan rule is that it is real. They are two separate fields on the alert, set independently.

solid answer

~30 s

Every finding ZAP records is an `Alert` in the core program, and it carries two scores. `risk` runs `RISK_INFO` through `RISK_HIGH` and answers "how bad if this is real". `confidence` runs `CONFIDENCE_FALSE_POSITIVE` through `CONFIDENCE_USER_CONFIRMED` and answers "how sure is the rule". The scales are different lengths — four rungs against five — and a fresh alert defaults to the bottom of the risk scale but the middle of the confidence scale. Reports splice them as `Medium (Low)`, risk first. There is no combined score on the alert: deciding what a high-risk low-confidence row is worth is your job, not the tool's.

go deeper

for a junior

Be able to say the two words apart without hesitating: severity if real, certainty it is real. Then name the ends of each scale, because the labels are what you will actually see in a report.

for a middle

Explain that both are plain integer fields on the core alert object, that the scales are different lengths, and that a fresh alert defaults to the bottom of one and the middle of the other.

for a senior

Show that you read the pair rather than the first word. Say out loud what your team does with a high-risk low-confidence row, and be honest that the tool publishes no ranking between that and a low-risk certain one.

for a principal

The tradeoff to own is what the organisation does with the axis the tool refuses to collapse. Decide whether certainty or consequence orders the queue, write it down, and accept that either choice leaks findings the other would have caught.

## Two fields on one object, answering two different questions Everything OWASP ZAP reports is an `Alert` — one object in the **core** program (`org.parosproxy.paros.core.scanner.Alert`), raised by an active scan rule, by a passive scan rule, or added by hand. Two of its fields are scores, and they are not two names for the same idea: - **`risk`** answers *how damaging would this be if it is real?* Its constants run `RISK_INFO`, `RISK_LOW`, `RISK_MEDIUM`, `RISK_HIGH`, and they print as **Informational**, **Low**, **Medium** and **High**. - **`confidence`** answers *how sure is the rule that it is real?* Its constants run `CONFIDENCE_FALSE_POSITIVE`, `CONFIDENCE_LOW`, `CONFIDENCE_MEDIUM`, `CONFIDENCE_HIGH`, `CONFIDENCE_USER_CONFIRMED`, and they print as **False Positive**, **Low**, **Medium**, **High** and **Confirmed**. Both are stored as plain integers and both are used directly as an index into the label array that prints them, so the numbering is the storage rather than a formatting detail. ## The two scales are not the same shape | | `risk` | `confidence` | |---|---|---| | the question it answers | how bad if real | how likely it is real | | rungs on the scale | four | five | | bottom rung prints as | `Informational` | `False Positive` | | top rung prints as | `High` | `Confirmed` | | default on a fresh `Alert` | `RISK_INFO` | `CONFIDENCE_MEDIUM` | | who normally supplies it | the scan rule | the scan rule, then a person | Three things fall out of that table. First, the scales have **different lengths**, so any mental model that pairs them rung for rung is wrong. Second, they have **different defaults**: a bare `Alert` starts at the bottom of the risk scale and in the *middle* of the confidence scale, which is the honest default for a machine that has not been told how sure it is. Third, the ends of the confidence scale are not degrees of confidence at all: the bottom one says the finding is not real and the top one says it has been confirmed. Measured across the shipped scan-rule add-ons, rules raise only at Low, Medium and High — neither end is ever set by a rule. ## How the pair reaches you An alert's own report fragment writes all of it out: a `<riskcode>` element carrying the risk integer, a `<confidence>` element carrying the confidence integer, and a `<riskdesc>` element that splices the two labels together as **`Medium (Low)`** — the risk band first, the confidence in parentheses. That composite string is where most readers first meet the pair, and it is why the second word is so often skimmed past: it looks like a qualifier on the severity when it is an independent measurement. The control API keeps them equally separate. The `alerts` view takes a `riskId` filter and a `confidenceId` filter as two different optional parameters, and there are **two** distinct actions for changing a stored finding — one that rewrites confidence and one that rewrites risk. Nothing in that surface offers a single blended score, because there is no such field to offer. ## Why the separation is the useful part The combination carries information neither axis carries alone: 1. **High risk, high confidence** — the rule is sure, and it matters. Act. 2. **High risk, low confidence** — the rule saw a weak signal for something serious. This is the one a human must look at, because the tool cannot settle it. 3. **Low risk, high confidence** — an observation that is certainly true and mostly cosmetic. A passive rule that sees a version-disclosing response header is the archetype: the header is either there or it is not, so the certainty is total, and the consequence is small. 4. **Informational, any confidence** — a note, not a defect. A gate that sorts purely on risk treats rows 2 and 1 identically and will waste a reviewer's day on weak signals; a gate that sorts purely on confidence will happily ship a serious issue because one rule hedged. ZAP deliberately refuses to collapse them for you: there is no combined score field on the alert, no weighting on it, and no defined ordering between a High-at-Low and a Low-at-High. That policy is yours to write, and writing it is the whole reason both numbers are on the object. ## The one place they stop being independent There is a single exception, and it is worth knowing before you trust a summary. The bottom confidence value, `CONFIDENCE_FALSE_POSITIVE`, is treated across the program as *"do not count this"* rather than as *"barely sure"*: the desktop swaps the alert's risk-coloured flag for a plain OK icon with a source comment saying there is no risk, the site tree leaves such alerts out when it works out a node's highest risk, and the default alert listing omits them unless you ask for them. At every other value the consumers read the two fields independently. For that one, confidence overrides risk.

  • What are the printed labels at each end of the two scales?
    Risk prints as `Informational`, `Low`, `Medium`, `High`. Confidence prints as `False Positive`, `Low`, `Medium`, `High`, `Confirmed`. Note that neither end of the confidence scale reads as a degree of certainty, and neither is ever set by a scan rule: rules raise only at Low, Medium and High. That is the hint that those two values behave differently from the middle three.
  • What does a ZAP report's `riskdesc` field such as `Medium (Low)` actually contain?
    The risk label, then the confidence label in parentheses. It is built by splicing the two label arrays, so it is a rendering of both axes rather than a third score. Readers routinely skim the parenthesised half and treat it as a qualifier on the severity, which is exactly backwards.
  • Does ZAP combine risk and confidence into one priority for you?
    No. There is no combined score field on the alert, no configured weighting, and no defined ordering between a high-risk low-confidence finding and a low-risk high-confidence one. Both numbers are published and the ranking policy is left to whoever consumes them.

A severe-weather notice carries two separate things: how bad the storm would be, and how likely it is to arrive. Neither number tells you the other, and no single figure replaces the pair.

saying these in an interview costs you the question

  • Treats confidence as another word for severity
  • Assumes a High finding is automatically a certain one
  • Thinks confidence is a per-rule setting you turn up
  • Expects the tool to merge the two into one priority score
  • Reads Informational as meaning nothing was found
open as a page

How can a single OWASP ZAP scan rule raise alerts at several different risk and confidence levels?

level: middleimportance: should knowfreq 50%

basics

~20 s

Both scores are set per finding, on the builder the rule uses to raise each alert. An active rule seeds risk from its own rule-level value and may override it; a passive rule seeds neither, so every raise site supplies both.

open as a page

Which OWASP ZAP confidence value stops an alert being counted, and where does that show up?

level: seniorimportance: should knowfreq 42%

basics

~20 s

The bottom of the confidence scale, the value named False Positive. It is a confidence value rather than a flag, and consumers special-case it: such alerts drop out of the highest-risk figure, the default alert listing and the risk counts.

open as a page

What does an OWASP ZAP alert's alertRef identify, and what does Alert.setAlertRef refuse?

level: middleimportance: nice to knowfreq 30%

basics

~20 s

The reference names the kind of alert, not the rule and not the occurrence, so one rule can have several. Setting one throws unless the string starts with the raising rule's plugin id — a check that is skipped while the reference is empty.

open as a page