skip to content

How does the OWASP Core Rule Set's anomaly score decide a refusal, and what room does a raised threshold give an attacker?

level: middleimportance: must knowfreq 60%

answer

  1. matches are evidence, not verdicts
  2. severity maps to points, points accumulate
  3. one verdict at the end of the phase
  4. critical 5, warning 3, threshold 5
  5. the number is discoverable by probing

basics

~20 s

Each matching Core Rule Set rule adds points by severity (critical 5, warning 3, notice 2); the request is refused only when the total reaches the inbound threshold, 5 by default. Every point added to it is attacker headroom.

solid answer

~50 s

The Core Rule Set does not refuse a request the moment one rule matches. In its default anomaly-scoring mode each match adds points according to the rule's severity - critical 5, error 4, warning 3, notice 2 - the points accumulate across every rule that matched, and one verdict is taken at the end of the request phase against the inbound anomaly threshold. That threshold defaults to 5, which is why a single critical match is enough while a single warning is not. Two consequences follow. A request that passed is not a request that matched nothing: sub-threshold matches are still recorded. And the threshold is discoverable by probing - send graded requests, watch where refusals begin - after which an attacker simply keeps each request beneath it. Every point you add to rescue one noisy route is that much more room, on every route sharing the setting.

code

text · 13 lines
text
Request A   POST /api/checkout
  match: protocol-enforcement rule   severity WARNING    +3
  match: paranoia-level-2 rule       severity CRITICAL   +5
  inbound anomaly score = 8   threshold = 5   -> refused

Request B   POST /api/checkout
  match: protocol-enforcement rule   severity WARNING    +3
  inbound anomaly score = 3   threshold = 5   -> served, match still recorded

Request C   POST /api/checkout
  match: paranoia-level-2 rule       severity NOTICE     +2
  inbound anomaly score = 2   threshold = 5   -> served, match still recorded
...

go deeper

for a junior

Know that matches add points rather than refusing outright, and that a total is compared to a threshold at the end. Being able to say the default threshold is 5 is a good sign.

for a middle

Explain the severity-to-score mapping and work an example: why one critical match is decisive at the default, why one warning is not, and why two warnings are. Be clear that a served request may still have matched rules.

for a senior

Show that you treat the threshold as a shared security parameter, not a tuning knob. Explain what raising it by five points hands to an attacker on every route, and how you would find out what a proposed change would actually have served.

for a principal

Be ready to argue where the threshold decision should live at all - a single number for a whole tier is an organisational choice with no owner, and you should have a position on who signs for changing it.

## Scoring, not a trip wire Early rule sets were pass/fail per rule: one match, one refusal. That design makes every single rule a potential outage, so every rule must be written tightly, and tight rules are exactly what an attacker reshapes their payload around. The Core Rule Set's anomaly-scoring mode exists to break that coupling. The mechanism is arithmetic: 1. Every rule carries a severity. Severity maps to a score: **critical 5, error 4, warning 3, notice 2**. 2. When a rule matches, it adds its score to a running inbound anomaly total for that request. It does not, by itself, refuse anything. 3. At the end of the request phase, the total is compared to the **inbound anomaly score threshold**, which ships at **5**. 4. If the total reaches the threshold, the request is refused. Otherwise it is served, and the individual matches are recorded anyway. So the default is calibrated such that one critical match is decisive on its own, while a single warning is not, and two warnings (3 + 3 = 6) are. A moderately suspicious request can be served; a request that is suspicious in several independent ways is not. That is the point: a rule can now be a piece of evidence instead of a verdict, which is what makes loose, higher-paranoia rules deployable at all. ## What this means for a request that passed A served request tells you the total stayed under the threshold. It does not tell you nothing matched. This trips people up constantly, because it inverts the usual reading: the record of a passed request can still contain several rule matches, and those records are the only trace that anything unusual was tried. The direction of the claim matters - a refusal proves the score reached the threshold, not that an attack occurred; a pass proves the score did not, not that the request was benign. ## The attacker's move, which is the whole point of this leaf The threshold is a number, the rule set is public, and the response to crossing it is observable. That is everything an attacker needs to calibrate. They send a series of requests of graded suspiciousness against your endpoint and watch where the refusals begin. The boundary between served and refused localises your threshold within a point or two, and from then on the attack is an optimisation problem: deliver the payload while keeping each individual request's score beneath the number. The ways under are ordinary. Split the work across several requests so that no single one accumulates enough. Choose an encoding or a placement that trips only the lower-severity rules. Move the payload into a part of the request that the high-severity rules do not target on this route. None of these defeat the rule set's understanding of the attack; they defeat the arithmetic. This is why the threshold is not a free tuning knob. When a route owner's flow is being refused and the quickest fix is to raise the tier's threshold from 5 to 10, you have not solved a false positive - you have doubled the working room of every attacker against every route that shares that setting, and you have done it in a way that produces no alert and no visible change. ## Reading a score trace Given a record showing which rules matched and what each contributed, the questions worth asking are: what was the total, what was the threshold at the moment of the decision, which single match dominated the total, and would removing that one match have changed the verdict? That last question is the honest test of whether a proposed exclusion is narrow enough to be safe: if pulling one rule out of the arithmetic takes a request from refused to served, that rule was carrying the decision on that route, and you should know exactly what else it was carrying. ## Two settings, two different arguments Paranoia level and anomaly threshold are frequently conflated and they answer different questions. The level decides **which rules are allowed to speak**; the threshold decides **how much accumulated suspicion is required to act**. Lowering the level removes evidence sources fleet-wide. Raising the threshold keeps every source but demands more of them before acting. They fail differently too: a level that is too high produces loud, visible breakage that gets reported; a threshold that is too high produces silence, which is precisely what makes it the more dangerous of the two to reach for under pressure.

  • How would an attacker discover the threshold you are enforcing?
    By probing. They send requests of graded suspiciousness and watch where refusals begin; the boundary between served and refused localises the number. Nothing about the process is exotic, the rule set and its severities are public, and the refusal is directly observable in the response. From there the attack becomes an exercise in keeping each request's score beneath the number they measured.
  • A request scored 4 against a threshold of 5 and was served. What is left behind?
    The matches themselves are still recorded, with the rules that fired and the score each contributed. So the record shows a request that was suspicious but not suspicious enough, which is exactly the shape a calibrating attacker leaves. Turning that record into a detection is a different discipline and a different team's job, but the raw evidence exists whether or not anyone reads it.
  • Why is raising the threshold more dangerous than lowering the paranoia level?
    Both weaken the control, but they fail differently. A level that is too low removes rules openly, and the gap is visible in the configuration. A raised threshold keeps every rule running and simply demands more accumulated evidence before acting, so the control still looks fully deployed while quietly serving requests it used to refuse - and on a shared tier the change applies to every route at once.

saying these in an interview costs you the question

  • Says one matching rule always refuses the request
  • Assumes a served request matched no rules at all
  • Treats the threshold as a per-route knob when it is shared
  • Calls a refusal proof that an attack occurred
  • Cannot separate severity from paranoia level

context