skip to content

An incident lead asks the exact date your DNS rule stopped matching - how do you establish it?

level: seniorimportance: should knowfreq 36%

answer

  1. date the data, not the narrative
  2. step change in field population
  3. release note dates a deploy only
  4. retention bounds what you can show
  5. no later than, not exactly on

basics

~20 s

Date the data, not the story: plot the daily count of events populating the field the rule filters on, find the step change, and correlate it with the parser deploy record. Bound the claim to what retention supports.

solid answer

~50 s

Three dates get confused here and you should separate them out loud: when the rule last produced a real match, when the field it depends on changed shape, and when the rule last could have matched. Only the middle one is directly measurable, and the measurement is a per-day count of events populating that field over the retention window - a step change is unmistakable. Then corroborate: the parser pipeline's deploy timestamp, a before-and-after record pair showing the rename, the rule's own version history, the scheduled-search execution telemetry, and the last canary firing if one existed. A release note by itself dates a deploy, not your rule's behaviour. Finally, bound the claim honestly: if the break predates retention you can only say `no later than the fourteenth`, and you must say which half is observation and which is inference. Dating the mute does not establish that a working rule would have matched this tunnelling.

go deeper

for a junior

Know that a claim about when a detection stopped working has to rest on data you can show, and that a change ticket or release note tells you when something was deployed, not how the rule behaved.

for a middle

Be able to build the measurement: a per-day count of events populating the field the rule filters on, over the retention window, plus a before-and-after record pair. Explain why a search that ran and matched nothing is not the same as one that never ran.

for a senior

Show the whole reconstruction and its limits - field population, deploy record, execution telemetry, the replay trap when data has been reparsed - and state the conclusion as a bound with observation and inference kept apart.

for a principal

Own what this costs the organisation: what assurance is now owed to people who acted on the old coverage claim, and which cheap long-lived artefacts, like aggregate field-population metrics, you will fund so the next reconstruction is measurable rather than argued.

## Separate the three dates before you answer An incident lead who has been told for five months that DNS tunnelling was covered wants one number. Before you give one, split the question, because three different dates are being run together: - **The last true match.** When the rule last fired on something real. This is the easiest to look up and the most misleading: a rule that models a rare behaviour may legitimately not have fired for a year before it broke, so this date is an upper bound of the loosest kind. - **The date the input changed.** When the field the rule filters on was renamed, reshaped, moved or rerouted. This is the one you can actually measure. - **The date the rule could no longer match.** What the lead is asking for. In the rename case it equals the input-change date, but only once you have shown the rule depended on the old shape. ## The evidence, strongest first **1. Field population over time.** Count, per day, the events in the relevant log source that carry the field the rule filters on, populated and of the expected shape. A parser change produces a step: four million a day, then zero, with the new field name appearing at the same instant. This is the sharpest artefact available and it comes from data you already hold. Do the same for the value shape when the field survived - average and maximum queried-name length per day will show a truncation change just as cleanly as presence shows a rename. **2. A before-and-after record pair.** Two real events either side of the boundary, showing the old and new field names on the same resolver and the same client. This is what makes the step change legible to someone who is not a detection engineer. **3. The change record.** The parser pipeline's deploy timestamp and release note, and the rule's own version history showing it was not updated. Note the epistemics: a release note is evidence of a *deploy*, not of your rule's behaviour. It corroborates the step change; it cannot substitute for it. **4. Execution telemetry.** The scheduled search's run history - it ran, on time, scanning a healthy event count, matching zero. This closes the alternative explanation that the search simply stopped running. **5. The last canary firing.** If a canary existed, the last one it produced is a hard lower bound on the rule's health, dated to the minute. If none existed, say so plainly; that absence is part of the answer to why nobody noticed. ## Replay, and the trap inside it The obvious instinct is to re-run the rule over historical data. Be careful about what you would be measuring. If your platform stores events as parsed at ingest, historical data before the change carries the *old* shape and today's rule will match it - useful, and it confirms the rule's logic. If the store was reindexed or reparsed, everything now carries today's shape and a replay tells you about today, not about March. So establish which of those your platform does before quoting a replay result. The field-population series does not have this problem, which is why it leads. ## Bound the claim to what retention supports Suppose the parsed store keeps ninety days and the suspected break is five months old. You cannot show the step change; the oldest surviving day already carries the new field name. The honest statement is then a bound with its reasoning attached: *at the start of my retention window the field was already renamed, and the parser deploy record puts that change on the fourteenth of March; so the rule was mute no later than the fourteenth of March, and probably from that date, but the direct evidence for the exact day is gone.* Say which half is observation and which is inference. An incident lead can work with a bound; nobody can work with a confident number that turns out to be a release note read backwards. This is also the argument for keeping cheap aggregate metrics far longer than raw events. A daily count of events by log source and key field costs almost nothing to retain for years and would have closed this gap exactly. ## What the date does, and does not, license Two claims the lead will try to derive, and only one of them follows. It **does** bound the window in which this rule could not have contributed. That is a real, useful statement about your coverage during the intrusion, and it belongs in the case record with the queries used, so somebody else can re-derive it. It **does not** establish that a working rule would have caught the activity. `Would have evaluated the data` and `would have matched` are different claims, and the second depends on whether this tunnelling implementation's query shape crossed the rule's thresholds. If pre-change data survives, test it and report the result. If it does not, say so - resist trading a broken assurance for an unprovable one. It also does not establish when the intrusion began; the tunnelling has its own timeline, built from the resolver data itself, and it may predate the mute entirely, which would mean the rule ran against the traffic and did not match it. That finding is worse than the parser bug and it is the one you must not bury.

  • Retention is ninety days and the break is five months old. What can you still claim?
    That at the oldest surviving day the field already carried the new name, plus the parser deploy record - which supports no later than the fourteenth of March, and nothing sharper. Mark clearly which half is observation and which is inference. A retained daily aggregate of events by log source and key field, had you kept one, is the artefact that would have closed the gap.
  • Why not just replay the rule over the historical index and see when matches stop?
    Because the answer depends on whether your store holds events as parsed at ingest or as reparsed later. If it was reindexed, everything now carries today's shape and the replay tells you about today. Establish which your platform does before quoting the result; the field-population series has no such ambiguity, which is why it leads.
  • The lead asks whether you would have caught it in January. What do you say?
    That a working rule would have evaluated the data, which is a weaker claim than would have matched. Whether it matched depends on this tunnelling's query shape against the rule's thresholds; test it if pre-change data survives, and say plainly if it does not. Do not replace a broken assurance with an unprovable one.
  • What if the tunnelling started before the rule went mute?
    Then the rule ran against the traffic and did not match it, which is a logic failure rather than a pipeline failure, and it is the more serious finding. Build the activity's own timeline from the resolver data, compare it with the mute date, and report the overlap explicitly rather than letting the parser bug absorb the whole explanation.

saying these in an interview costs you the question

  • Offers the parser release note alone as the break date
  • Dates the break to the rule's last true-positive alert
  • Claims a day older than retention can support
  • Promises the intrusion would have been caught before the break
  • Replays today's logic on reparsed data and calls it history
  • Presents inference without labelling it as inference

context