skip to content

Why should a policy rule read normalised finding fields rather than each scanner's native output?

level: middleimportance: should knowfreq 55%

answer

  1. three tools, three shapes
  2. structure, vocabulary, identity
  3. one record shape, one rule set
  4. the mapping table is policy too
  5. unmapped severity becomes unknown, not low

basics

~20 s

Because a rule bound to one tool's schema must be rewritten for every other tool and breaks when that schema changes. A normaliser maps each report into one record shape, so policy is written once against stable fields.

solid answer

~50 s

Scanner output is heterogeneous in three ways at once: structure, vocabulary and identity. One tool emits SARIF, where a result carries a `ruleId`, a `level` of error, warning, note or none, and a location; another emits its own JSON with a `check` name and a `sev` of HIGH. Read those shapes directly and you get one rule per tool, three severity vocabularies inside your policy, and a rewrite every time a vendor reshapes a field. So the platform puts a normaliser in front: it maps every report into a single record — what was scanned, which tool and version said so, a check identifier, a canonical severity, a location, and an envelope stating the scan completed and when. Rules read only those fields. The cost is real: the mapping is lossy and opinionated, and the normaliser becomes the one place where policy can go blind.

code

json · 17 lines
json
{
  "sarifResult": {
    "ruleId": "RULE-107",
    "level": "error",
    "locations": [
      { "physicalLocation": { "artifactLocation": { "uri": "src/db.py" } } }
    ]
  },
  "vendorResult": { "check": "hardcoded-credential", "sev": "HIGH", "file": "src/db.py" },
  "normalised": {
    "subject": "sha256:9f2c1a...",
    "tool": "tool-a", "toolVersion": "3.4.1",
    "checkId": "RULE-107", "severity": "high",
    "path": "src/db.py",
    "scanStatus": "completed", "scannedAt": "2026-08-14T09:12:03Z"
  }
}

go deeper

for a junior

Know that different scanners emit different report shapes and different severity words, and that a gate is written against one agreed record shape rather than each tool's own format.

for a middle

Explain the three axes that differ — structure, severity vocabulary, and what the finding points at — and list the fields a normalised record must carry, including the scan envelope.

for a senior

Show that you treat the normaliser as a trust boundary: schema validation on input, no invented defaults, a reviewed severity mapping, fixture tests, and an unknown bucket for values you have never seen.

for a principal

Argue about where the mapping is owned and how a new tool is onboarded without every team rewriting rules, and be honest that the normaliser concentrates risk in one component you must staff, test and monitor.

## What 'heterogeneous' actually means Reports differ on three axes, and each axis breaks a rule differently. **Structure.** Where the findings live in the document, how deeply they nest, whether one finding is one object or a row joined to a rule definition elsewhere in the file. A rule that walks `runs[].results[]` cannot walk `issues[]`. **Vocabulary.** SARIF's `level` takes the values error, warning, note and none — a statement about how the tool wants a result treated, not a risk ranking. Other tools emit Critical/High/Medium/Low, or a numeric score, or pass/fail. Three vocabularies, no free translation between them. **Identity.** What the finding points at differs by class of tool: a file and line, a package and version, an image layer, a cloud resource. A rule that assumes every finding has a path silently ignores every finding that has a package instead. ## The canonical record The normaliser's output is a flat record designed for rules to read, not for humans to browse: - **subject** — the immutable identity of what was scanned, normally an artifact digest; - **source** — tool name, tool version, and the report format it arrived in; - **checkId** — the tool's own rule identifier, kept verbatim so it can be traced back; - **severity** — a value on one canonical scale, produced by an explicit mapping; - **location** — path and line, or component identity, depending on the finding class; - **envelope** — scan status, timestamps, and what the scan covered; - **raw pointer** — a reference back to the original record, so a human can audit the mapping. Rules then read `severity`, `subject` and `checkId` and never learn which tool produced them. ## Why rules must not see raw output One rule set instead of one per tool. Tool substitution without touching policy. Fixture-driven tests, because the rule's input is a shape you control. Mapping decisions reviewed once, in one component, instead of smeared across dozens of rules owned by different teams. And a policy repository that does not change every time a vendor ships a release. ## The mapping is the hard part, and it is lossy Mapping HIGH to high is trivial. Mapping one tool's `error` to another's `medium` is a judgment call that someone has to make and write down. Two rules of thumb. First, keep the mapping table in code, versioned and reviewed like policy, because it *is* policy. Second, an unrecognised value maps to **unknown**, never to the lowest severity — quietly flooring a value you have never seen is precisely how real findings disappear, and it fails in the direction that looks calm. Surface the unknown so the table gets updated. What you then *decide* on the basis of a severity — where the line sits, which severities block — is a separate question about thresholds, not about normalisation. ## The normaliser is a trust boundary Two disciplines follow from that. **Validate the input**: check each report against the schema you expect for that tool, and refuse to emit confident records from a document that does not validate; emit an unknown-status record instead, so the gate's fail-closed path fires rather than its pass path. **Never invent data**: a field the tool did not provide must be absent from the record, not filled with a plausible default. A defaulted severity or a defaulted subject gives rules a confident wrong answer, which is worse than no answer, because no answer can be detected. ## What normalisation does not do It does not decide which tool is right when two disagree. Presenting both statements in a comparable shape is the whole job; adjudicating between them is a triage question with a different owner. Nor does it deduplicate risk judgments for you — although a stable per-finding key, built from subject plus check id plus location, is worth producing, because anything that compares one run against another needs to know that a finding is the same finding. ## The visible payoff When a new tool is onboarded, the work is a mapping and a set of fixtures, not a policy rewrite. When a tool is retired, no rule changes. When an auditor asks what the gate actually reads, you can point at one schema instead of narrating a dozen vendor formats.

  • What belongs in a normalised record beyond the finding itself?
    The envelope and the provenance: what was scanned by immutable identity, which tool and version said so, whether the scan completed and what it covered, when it ran, and a pointer back to the raw record. Without those a rule can read findings but cannot tell whether it is looking at a real, current, relevant scan.
  • What should the normaliser do with a severity value it does not recognise?
    Record it as unknown severity, keep the raw string alongside, and raise it so the mapping table gets updated. Mapping it to the lowest level hides a real finding, and dropping the record destroys evidence. Unknown is the only option that fails in a direction someone will notice.
  • Doesn't normalising just move the coupling from the rules into the normaliser?
    Yes, deliberately. The coupling still exists, but it now lives in one reviewed, fixture-tested component instead of being spread across dozens of rules owned by different teams. That concentration is the point, and it is also why the normaliser needs its own tests, its own schema validation and its own alerting.

saying these in an interview costs you the question

  • Writes a separate rule per scanner and calls it flexible
  • Assumes every tool's severity words mean the same thing
  • Defaults an unrecognised severity to the lowest level
  • Lets the normaliser invent values the tool never supplied
  • Thinks normalisation decides which tool is correct

context