skip to content

Regression by Re-Execution

An agent upgrade or a renamed field can silently retire a detection, and the only proof it still works is executing the technique again on a live host. Interviewers ask how you staff that.

on this pageshow

explore

questions

4

Why doesn't replaying saved log records prove a detection still works after a sensor upgrade?

level: juniorimportance: must knowfreq 58%

answer

  1. the test starts halfway down the path
  2. upstream stages are assumed, not exercised
  3. collection is the fragile part
  4. only a real action reaches every stage

basics

~20 s

Replaying saved records tests only the stages after collection: parsing and rule logic. It cannot show that the upgraded sensor still emits that event with the same fields, or that the forwarder still ships it. Only re-executing the real technique exercises the whole path.

solid answer

~50 s

A detection depends on a chain: the action happens, the source decides to record it, a forwarder ships the record, the platform parses it into fields, the rule evaluates those fields, an alert reaches a human. A replay injects stored records partway down that chain, so it proves the logic still matches records of that shape and nothing more. Upgrades break the stages a replay skips — an appliance firmware upgrade that resets its syslog destination, drops an audit category from the default set, or moves a field so it is no longer populated. None of that errors, and the rule still looks healthy. Re-executing the technique — performing the real action again in the estate — is the only thing that traverses every stage. If no alert appears, confirm the action actually succeeded before calling the detection broken.

go deeper

for a junior

Be ready to name the stages between an action happening and an analyst seeing an alert, and to say which of them a replayed record never touches.

for a middle

Explain the concrete upstream failures an upgrade causes — an audit category dropped from the defaults, a reset syslog destination, a field that stops being populated — and why none of them raise an error anywhere.

for a senior

Show the judgment of running both layers at different frequencies, and of verifying that a technique actually executed before you report a silent detection as broken.

for a principal

Own the framing that a validation result is only as strong as the stage it starts from, and be prepared to say what your organisation may and may not claim about coverage on the strength of replay alone.

## The chain a detection actually depends on A working detection is not one artefact, it is a chain of six stages, and every one of them can fail independently: 1. **The action happens** — someone authenticates to the backup server's management console, or an administrator adds an account on the hypervisor management appliance. 2. **The source decides to record it** — an audit policy, a logging level, or a vendor's default set of enabled audit categories determines whether an event is written at all. 3. **A forwarder ships it** — an agent, or in the appliance case a syslog destination configured on the box itself. 4. **The platform ingests and normalises it** — the raw message is parsed into fields the rule can name. 5. **The rule evaluates those fields** and matches, or doesn't. 6. **An alert is raised and routed** to somebody who will act on it. ## What a replay tests, and what it assumes A replay takes records captured during an earlier successful test and re-injects them, typically at stage 4. It therefore exercises stages 4 to 6 and *assumes* 1 to 3. That is a genuinely useful thing to assume when what changed is downstream — a query-language migration, a normalisation change, an edit to the rule. It is worthless when what changed is upstream, and upstream is where most breakage lives. ## Where upgrades actually break things Take a detection over on-premises virtualisation and backup infrastructure: an authenticated administrative login to the backup server's management interface from an account or subnet that has no business there. This is the step that precedes destroying restore points, and it is normally mapped to valid-account use (`T1078`) followed by inhibiting recovery (`T1490`). The telemetry is the appliance's own application audit log, forwarded by syslog. A firmware upgrade on that appliance can, entirely without error: - reset the syslog destination to unset, or to a port the collector does not listen on; - change which audit categories are enabled by default, so login events are simply no longer written; - reformat the message so the source-address field lands somewhere the parser does not look, leaving the rule's field empty rather than wrong. In all three cases the rule is untouched, a replayed capture still matches, and the source is silent. The estate looks green and is not. ## Why re-execution is the only proof Re-execution means performing the real action again in the live estate — a genuine authenticated login against the real management console, from the account and network position the detection describes — and watching for the alert. That single act traverses all six stages. It is the only test that can distinguish *the sensor still emits this* from *the rule still parses this*. ## Verify in both directions If an alert fires, confirm it is the alert you expected, on the right entity, with the fields an analyst needs. If no alert fires, do **not** immediately declare a regression. Confirm the technique actually executed and succeeded — the operator's own notes, a session visible in the console, the appliance's local log. Silence can also mean the action never happened, or that a preventive control blocked it, which is a different and better outcome. An absent alert on its own proves nothing about the detection. ## What re-execution costs, and why that matters A replay costs seconds and no permission. A re-execution costs a change window on infrastructure everything else depends on, an approval from the platform owner, careful handling of genuinely privileged credentials, an announcement (or a deliberate non-announcement) to the on-call analyst, and a rollback plan for anything you touched. That cost is exactly why re-execution is scheduled rather than continuous, and why deciding which techniques earn a recurring slot is a real question rather than a formality. ## Use both, at different frequencies The sane arrangement is layered. Replay is cheap and runs whenever the logic or the platform changes — it answers *does this logic still match this shape*. Re-execution runs on a cadence and whenever a source, agent or platform component changes — it answers *does this estate still see this behaviour*. Reporting a replay pass as if it were coverage is the mistake this whole practice exists to prevent.

  • If a replay can't prove collection works, is a replay harness worth keeping at all?
    Yes, for a narrower job. It costs seconds and no approvals, so it can run on every rule edit and after any query-language or normalisation change, catching logic and parsing breaks immediately. It just answers a smaller question than people credit it with. Keep it as the fast layer underneath scheduled re-execution, and never report its result as coverage.
  • Your re-run produced no alert. What do you check before declaring the detection broken?
    First that the technique really executed and succeeded — the operator's notes, the session visible in the target console, the appliance's local audit trail. Then walk the chain from the source outward: was the record written on the box, did it leave, did it arrive, did it parse into the fields the rule names. Silence can also mean a preventive control blocked the action, which is a pass, not a failure.
  • Which stages does a replay actually exercise?
    Parsing and normalisation from the injection point onward, rule evaluation, and alert routing. Everything before the injection point — whether the source recorded the event, and whether the forwarder shipped it — is assumed by construction and never tested.

Replaying stored records is like testing a fire alarm by pressing the panel's self-test button. Re-executing the technique is holding a smoke source under the detector: only one of them tells you the detector is still wired in.

saying these in an interview costs you the question

  • Treats a passing replay as proof the detection still works
  • Thinks replayed records exercise the sensor or the forwarder
  • Assumes no alert during a re-run means the technique was blocked
  • Calls a silent detection a new coverage gap without checking the source
  • Believes an untouched rule cannot stop working

context

open as a page

Which platform changes should trigger re-executing a detection test out of cadence?

level: middleimportance: should knowfreq 46%

basics

~20 s

Any change to something a detection depends on: the source's audit configuration or firmware, the agent or forwarder version, the network path to the collector, and the platform's ingest or normalisation schema. Re-run the specific tests that depend on the changed component, not the whole suite.

open as a page

With capacity to re-execute 20 of 200 passing attack techniques a quarter, how do you choose?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Rank by consequence of a single success, fragility of the telemetry path, and exposure to change since the last pass. Give a small permanent core the estate cannot survive missing, spend the rest on recently changed and single-source detections, and rotate the tail so nothing is untested indefinitely.

open as a page

A re-executed attack technique that alerted last quarter now produces nothing. What do you do?

level: seniorimportance: nice to knowfreq 34%

basics

~20 s

Confirm the technique really executed, then treat it as a regression rather than a new gap: find what changed, scope every other detection sharing that dependency, and account for the blind window between the change and this discovery. It is a change-management finding, not a coverage backlog item.

open as a page