skip to content

A regenerated findings baseline let a real regression pass a green gate. How did that happen?

level: seniorimportance: nice to knowfreq 32%

answer

  1. the reference moved, not the rule
  2. regenerate overwrites, merge adds
  3. run on a branch, absorbs the branch
  4. never let the pipeline rewrite it
  5. watch entry count, review the diff

basics

~20 s

Regenerating a baseline overwrites it with whatever the current scan reports, so a finding the branch just introduced is recorded as pre-existing and stops counting as new. The gate then passes truthfully while a real regression ships.

solid answer

~50 s

A new-findings gate is only as trustworthy as the baseline it compares against. Regeneration is a full overwrite: it takes the current scan and declares all of it known. Run on a branch that introduced a critical, it absorbs that critical into the baseline, and the very next evaluation finds nothing new. The gate is not lying — it is answering the question it was asked, against a reference someone just moved. The usual triggers are a developer clearing an unrelated blockage, a helper step that regenerates on failure, or a bulk refactor where regenerating was genuinely easier than reconciling fingerprints. The defences are process rather than cleverness: never regenerate inside the pipeline, review baseline diffs as security changes, regenerate only from a trunk scan, and alarm on baseline growth independently of build outcomes so an expanding file is visible even when every build is green.

go deeper

for a junior

Understand that the baseline is an ordinary file in the repository, that regenerating it replaces its whole contents, and that whatever the scan sees at that moment becomes part of the accepted starting point.

for a middle

Explain the difference between merging specific entries and regenerating wholesale, and why regenerating on a branch absorbs that branch's own new findings into the reference.

for a senior

Show the detection and prevention playbook: baseline diffs reviewed as security changes, entry count tracked independently of build outcomes, regeneration derived from the mainline and never performed by the pipeline itself.

for a principal

Be ready to defend baselining after an incident like this. Argue that a watched, movable reference outperforms an absolute gate that gets switched off, and specify who reviews reference changes and who is accountable when the count grows.

## Why the mechanism is fragile A new-findings gate answers a **relative** question: is anything here that was not in the reference set? Its correctness therefore depends entirely on the reference set being trustworthy. The reference set is a file in the same repository, editable by the same people the gate constrains, and updating it is a routine, encouraged operation. That combination is the fragility. ## What regeneration actually does Updating a baseline comes in two flavours, and conflating them is the root of the incident. - **Merge**: add specific, named findings to the known set, leaving everything else alone. Reviewable, because the diff is small and each added entry can be argued about. - **Regenerate**: scan now, write the entire result as the new baseline. Fast, tidy, and it silently absorbs anything present at that moment — including whatever the branch just introduced. Regenerate on a branch that added a critical and the critical becomes 'known'. The subsequent evaluation reports no new findings, which is exactly right given the reference it was handed. The gate did not fail; the reference moved. ## How it happens without malice Almost never as an attempt to smuggle a vulnerability. The realistic paths: - A developer is blocked by unrelated fingerprint churn — a reformat, a file move — and regenerating is the documented, one-command fix. It clears the churn and their own new finding at the same time. - A convenience step regenerates the baseline automatically when the gate fails, which converts the gate into a no-op by construction. This is the worst version because it is systematic rather than occasional. - A large refactor produces hundreds of unmatched fingerprints for problems that all pre-date it, and reconciling them by hand is genuinely infeasible, so someone regenerates and a real regression rides along in the noise. ## Detecting it - **Diff the baseline, not just the code.** An added entry is a security change and deserves the same scrutiny as an edit to the policy itself. Route baseline changes to a reviewer who is not the author of the change that needed them. - **Watch the size.** Track how many entries the baseline holds over time. Growth is a policy event even when every build was green, and it is the only signal that survives when the gate itself has been rendered silent. - **Compare against a trunk-derived reference.** Periodically scan the mainline independently and compare its findings against the committed baseline. Entries that appear on a branch and never existed on trunk before that branch are exactly the pattern you are hunting. - **Report the total, always.** The reporting lane exists partly for this: if the absolute count jumps while the gate stays green, something moved the reference. ## Preventing it - **Never regenerate inside the pipeline.** The gate must not be able to rewrite its own reference. If a build step can update the baseline, the gate is decorative. - **Regenerate only from a scan of the mainline**, never from a feature branch, so a branch's own additions can never enter the reference. - **Prefer merge over regenerate** for day-to-day updates, and make regeneration a deliberate, announced operation with a review. - **Make entries expire.** An entry with an owner and a date is one that someone must revisit; a permanent entry is one nobody ever looks at again. - **Separate the refactor case.** When a legitimate large change churns fingerprints, handle it as its own reviewed change that touches only the baseline, so the reconciliation is visible and is not mixed with functional edits. ## The honest framing This is the standing cost of baselining. Grandfathering is what made the gate survivable on a real codebase, and the price is a reference file that the gated team can move. You do not eliminate that; you make moving it visible, reviewed and rare. If an interviewer pushes on whether the mechanism is therefore worthless, the answer is that an unbaselined absolute gate is not stricter in practice — it is simply switched off faster, and a gate that is off catches nothing at all. A movable reference that is watched beats a perfect rule that nobody runs. ## What to say in the incident review Name the sequence plainly: the baseline was regenerated on a branch that carried a new finding, the finding entered the reference, and the gate correctly reported no new findings against a reference that now contained it. Then propose the process controls, and resist the temptation to describe it as the gate having failed — it did what it was configured to do, which is a more uncomfortable and more useful conclusion.

  • Why is a step that auto-regenerates the baseline when the gate fails worse than an occasional manual regeneration?
    Because it removes the gate entirely by construction. Every failure is converted into an updated reference, so no finding can ever remain new and the build is green forever. A manual regeneration is an occasional mistake with a diff someone could have reviewed; the automatic version guarantees the outcome and leaves no moment where anyone would notice.
  • A refactor legitimately churns hundreds of fingerprints. How do you handle it without opening this hole?
    Split it. Land the refactor and the baseline reconciliation as separate, reviewed changes, and derive the new reference from a scan of the mainline after the refactor rather than from the branch. Compare entry counts before and after: the reconciliation should re-key existing problems, not increase how many are grandfathered.
  • Every build is green and you suspect the reference has been moved. What do you look at?
    The baseline's history and its size over time, and the absolute findings total from the reporting lane. A growing entry count or a jump in the total while the gate stayed green is the signature. Then scan the mainline independently and look for entries in the committed baseline that never appeared in a trunk scan before the branch that added them.

It is like resetting a burglar alarm's baseline photo while someone is standing in the room: nothing new is ever detected again, and the alarm is telling the truth.

saying these in an interview costs you the question

  • Calls it a gate failure rather than a moved reference
  • Lets a pipeline step regenerate the baseline on failure
  • Reviews only code diffs and never baseline diffs
  • Regenerates from a feature branch scan
  • Concludes baselining is worthless and returns to an absolute gate

context