Which platform changes should trigger re-executing a detection test out of cadence?
answer
- start from what each detection depends on
- source, transport, platform, identity
- the change calendar is the trigger
- the rule never errors when inputs vanish
- narrow run, not the whole programme
basics
~20 sAny change to something a detection depends on: the source's audit configuration or firmware, the agent or forwarder version, the network path to the collector, and the platform's ingest or normalisation schema. Re-run the specific tests that depend on the changed component, not the whole suite.
solid answer
~40 sWork from a dependency map: each detection records the source it needs, the audit category that source must have enabled, the fields it keys on, and the collector that carries them. Then any change to one of those components fires the tests that depend on it. In practice that means source-side changes (appliance or application upgrades, audit-policy edits, a vendor changing default audit verbosity), transport-side changes (agent major versions, forwarder or certificate changes, a collector move or a firewall change on the syslog path), and platform-side changes (ingest schema or normalisation changes). Wire this to the change calendar rather than to memory: the platform owner's change ticket should name the regression tests it puts at risk. Out-of-cadence runs are narrow — a handful of techniques tied to one component, not the quarterly programme.
code
text · 7 lines# before the upgrade
Aug 12 02:11:07 bkpsrv01 backupd: audit user=svc_backup action=console_login src=10.20.4.51 result=success
# after the upgrade — same event, no source address
Aug 14 02:11:07 bkpsrv01 backupd: audit user=svc_backup action=console_login result=success
...
# the rule keys on src outside the admin subnet; it now never matches, and nothing errorsgo deeper
Know that upgrading an agent, appliance or forwarder can stop a detection working even though nobody edited the rule, and that this is why changes get re-tested.
Be able to list triggers across source, transport and platform, and to explain how a dependency map turns a change ticket into a specific short list of tests.
Demonstrate wiring the trigger into the change calendar so it fires without anyone remembering, and scoping the run narrowly enough that the platform owner keeps agreeing to it.
Frame the missing link between change management and detection validation as an ownership problem, and be ready to say who is accountable when a change ships without re-validation.
## The premise: a detection has a bill of materials You cannot decide which changes matter until each detection states what it depends on. The minimum is: the **source** (which appliance, host agent, service or audit trail), the **event category** that source must have enabled, the **fields** the logic keys on, and the **collector or forwarder** that carries the record. That list turns a vague worry into an index you can query when a change ticket appears. ## Source-side changes These are the highest-yield triggers because the source is where the record either exists or does not. - **Firmware or version upgrades on an appliance** — a hypervisor management console or a backup server. Vendors change which audit categories are on by default, and they change message formats. - **Audit-policy or logging-level edits**, whether deliberate or a side effect of restoring a configuration from a template. - **Licence or tier changes**, which on some products silently remove audit or forwarding features. - **Migration of the source** to a new host, VLAN or site, which usually means a new path to the collector. ## Transport-side changes - **Agent major versions** — new field names, changed defaults, telemetry moved behind a policy toggle. - **Forwarder or collector upgrades**, certificate rotation on a TLS syslog path, or a firewall rule change on the port the appliance sends to. - **Credential changes** on the account a collector uses to pull an audit trail; a permission quietly lost is indistinguishable from an estate with nothing to report. ## Platform-side changes - **Ingest schema or normalisation changes** that rename or retype a field. - **Routing or index changes** that put records somewhere the rule does not look. ## Non-technical triggers worth honouring Re-execute after an incident in which the technique appeared, after you materially edit the detection itself, and after any change to who owns the source — an ownership handover is a good moment to prove the plumbing still works while both parties are paying attention. ## Why the rule looks fine throughout The thing that makes this discipline necessary is that none of these changes touch the rule. Its logic is unchanged, its stored test inputs still match, and its dashboard shows no error. The signal you have lost is *absence*, and absence is what a healthy quiet estate also looks like. That is why the trigger has to come from the change calendar rather than from the detection platform noticing something. ## What the trigger actually launches An out-of-cadence run is deliberately narrow: the specific techniques whose dependency list names the changed component, executed once, soon after the change window closes. It is not the quarterly programme brought forward. Keeping it narrow is what makes it affordable enough to actually happen — a handful of executions against one appliance is a conversation with one platform owner, not a scheduling exercise. ## Wiring it so it happens without heroics The practical mechanism is boring and effective: detection engineering subscribes to the change calendar for the components in the dependency index, and change tickets for those components carry a line naming the regression tests at risk. The platform owner is not being asked to understand your detections; they are being asked to tell you when they touched your inputs. When that link is missing, the failure mode is predictable — the breakage is found by the next scheduled run, months later, and the window between the change and the discovery is a period you cannot make any coverage claim about. ## A worked shape A backup appliance is upgraded on a Thursday night. Its audit events reach the platform by syslog, and one detection keys on an administrative console login from outside the admin subnet. After the upgrade the events still arrive, but the source-address field is no longer present in the message, so the rule's condition can never be true. Nothing errors; the appliance is healthy; the detection is dead. The only cheap way to have caught it that week was a change ticket that named the test, and a single re-executed login the following morning.
- The platform owner says they upgrade that appliance monthly. How do you keep this affordable?Scope the trigger to the one or two techniques whose dependency list names that appliance, and agree a standing slot in the morning after each change window rather than negotiating every time. If the technique itself is expensive or risky to execute, agree a cheaper proxy action that traverses the same source, category, field and collector, and re-execute the full technique on the quarterly cadence.
- How do you tell a lost field from a technique that simply didn't run?Check the source first: does the appliance's local audit log contain the event you expected. If it does and the platform has nothing, the loss is in transport or parsing. If the appliance has nothing either, the loss is at the source or the action never happened, and the operator's own record of the execution settles which.
- Which of these triggers do teams most often forget?Identity and permission changes on the account a collector uses to read an audit trail, and vendor changes to default audit verbosity that arrive with a routine upgrade. Both remove records without touching anything the security team owns, so nothing in the detection platform reflects them.
saying these in an interview costs you the question
- Waits for the quarterly cadence after a known source upgrade
- Re-runs the entire suite instead of the dependent tests
- Tracks only agent versions and ignores appliance firmware
- Assumes a lost field will surface as an ingest error
- Has no record of which source and fields each detection needs