skip to content

A scanner upgrade reshaped its report and your gates stopped blocking — how would you catch that?

level: seniorimportance: nice to knowfreq 33%

answer

  1. an empty set of denials looks clean
  2. nothing errors when a field vanishes
  3. your fixtures still encode the old schema
  4. validate at the boundary, not in the rule
  5. a bad artifact that must be blocked

basics

~20 s

Rules reading a field that no longer exists produce no denial, which looks exactly like a clean estate. Validate reports against an expected schema at ingest, test the normaliser against stored real samples, and push a deliberately bad artifact through the real gate.

solid answer

~50 s

This is the failure the platform team owns, and it is silent by construction: a rule whose body reads a field the report no longer carries produces no result, and an empty set of denials is indistinguishable from a clean build. Nothing errors, nobody files a ticket, and the first signal is usually an incident or an auditor. Three defences, in order of value. Validate at the boundary — the normaliser checks each report against the schema it expects for that tool and emits an unknown-status record rather than records it cannot trust, so the fail-closed path fires instead of the pass path. Keep golden samples — real output from each tool version as fixtures, with tests asserting the normalised records, so the upgrade breaks the platform's own build. And run a negative canary: a synthetic artifact that must be blocked, pushed through the real gate on a schedule. If it is admitted, page someone.

go deeper

for a junior

Know that when a report field disappears the rule usually produces nothing at all rather than an error, so the gate keeps passing builds and looks perfectly healthy.

for a middle

Explain where the report should be validated — at the normaliser boundary against an expected schema — and why an unrecognised document should become an unknown result rather than an empty one.

for a senior

Describe the layered defence: schema validation at ingest, fixtures refreshed from real tool output, a negative canary through the real gate, and counters that alarm when decisions stop happening.

for a principal

Treat a scanner version change as a policy change with a blast radius, and decide in advance how the platform reports a blind window it found itself rather than hoping the question never comes.

## Why this failure is silent Three shapes of drift, all quiet: - a field is **renamed or moved** — the document still parses, the path the rule reads is simply empty; - a **value vocabulary changes** — a severity string you never mapped arrives, and an unforgiving normaliser floors or drops it; - the findings move under a **new envelope** — the report is valid, larger, and your extraction returns nothing. In every case the rule body produces no result, no denial is emitted, and the gate reports the same thing it reports for a perfect artifact. Gates are built out of denials, so the absence of denials is the absence of information, and it is displayed as success. ## Layer one: validate at the boundary Do not let unvalidated documents reach rules. The normaliser holds an expected schema per tool, checks each report against it, and on a mismatch emits a record with an unknown status instead of a plausible-looking empty result. That converts an invisible pass into a visible unknown, which the gate already knows how to fail closed on. Version the normalised record contract separately from any tool's schema, so a vendor change is contained in the mapping layer. ## Layer two: fixtures from real output Store genuine sample reports from each tool version in the platform's own repository, and assert on the normalised records they produce. This is the test that fails the day the fixtures are refreshed against a new tool version — which means refreshing fixtures must be part of the upgrade, not a follow-up chore. The trap here is worth stating plainly: **the rules' own unit tests will stay green throughout the outage.** They run against fixtures you wrote, which still describe the old shape. Rule tests prove the rule handles the input you imagined. They say nothing about whether the tool still produces that input. Only fixtures refreshed from reality, or an end-to-end canary, close that gap. ## Layer three: the negative canary Push a synthetic artifact that *must* be blocked through the real gate on a schedule — the whole chain, from scanner invocation to report upload to normalisation to rule evaluation to the gate's wiring. If it is admitted, alert loudly. This is the only check that covers the parts nobody unit-tests: the upload path, the file naming, the job that was quietly disabled, the credential that expired. Pair it with a positive canary — a known-good artifact that must pass — so the opposite failure, a gate that blocks everything, is caught too. ## Layer four: alarm on missing decisions, not only on failures Most monitoring counts failures. This failure mode is the absence of failures, so instrument the decision flow itself: per tool and per day, count reports ingested, records emitted per report, evaluations run, denials produced, and unknowns. The signatures to alert on are shapes, not absolute numbers — records-per-report collapsing while ingestion stays flat; unknowns spiking; a rule that used to fire regularly going to a flat zero. A gate that has never once said no is not a compliant estate, it is an untested gate. ## Change control on the tool If the scanner version can float underneath you, a vendor release becomes an unannounced policy change with an unknown blast radius. Control when the version changes, and land the fixture refresh in the same change. That is not conservatism about upgrades; it is the difference between a schema change being a build failure and a schema change being a silent hole. ## Designing so silence is harder Prefer rules that assert presence before content. If a rule first asserts *this report exists, validates, is complete, and describes this artifact*, then a vanished field trips that assertion instead of evaporating. It is a small structural discipline that converts an entire class of quiet failure into a loud one. ## The morning you find it Do not declare the blind window clean. Nothing was blocked because nothing was evaluated, so those artifacts are unevaluated, not compliant. Re-run the gate against the stored evidence for everything admitted during the window, and treat re-scanning as the fallback only where evidence was never captured. Reporting the window yourself, with a number attached, is a much better position than having someone else find it.

  • Your rule unit tests were green throughout. Why did they not catch it?
    They run against fixtures you wrote, which still describe the old report shape. Rule tests prove the rule handles the input you imagined; they say nothing about whether the tool still produces that input. Refresh fixtures from real output as part of every tool upgrade, and add an end-to-end canary that exercises the real chain.
  • What telemetry would have surfaced this within a day?
    Per-tool counters for reports ingested, normalised records emitted, evaluations, denials and unknowns. A schema change shows up as records-per-report collapsing while ingestion stays flat, or as a rule that used to fire dropping to a flat zero. Alert on the shape of the change rather than on an absolute count.
  • What do you do about artifacts admitted during the blind window?
    Treat the window as unevaluated rather than clean. Re-run the gate against the stored evidence for everything admitted while it was blind, re-scan only where no evidence was captured, and then decide what needs remediating. Calling them fine because nothing was blocked repeats the original mistake.

saying these in an interview costs you the question

  • Assumes a rule reading a missing field would throw an error
  • Trusts rule unit tests written against hand-made fixtures
  • Lets the scanner version change without refreshing fixtures
  • Alerts only on failures, never on decisions that stopped happening
  • Calls the blind window clean because nothing was blocked

context