skip to content

You ship an exclusion to a live detection rule — how do you find out later what it hid?

level: middleimportance: should knowfreq 47%

answer

  1. a suppressed match makes no noise
  2. the cost never announces itself
  3. where it is enforced decides what survives
  4. the records still reach the log store
  5. re-run the strict logic over the window

basics

~20 s

An exclusion suppresses matches silently. Enforce it in the rule's logic rather than at the sensor, so the records still reach the log store, and record its scope and start date so you can search that gap later.

solid answer

~50 s

A suppressed match is silent: no alert, no ticket, no complaint, so an exclusion's cost never announces itself the way noise does. Two decisions at the moment you write it fix that. First, enforce it in the detection rule's own condition rather than in the agent's collection filter or the ingest pipeline — a rule-level exception still lets the records reach the log store, so the gap stays searchable, while dropping the telemetry removes it permanently and for every other rule as well. Second, ship it as a change in version control carrying its scope, owner and start date, because the question you will actually be asked is "what were we not alerting on, on those hosts, since 4 March?". Then, when an intrusion lands in the excluded population, you re-run the un-excluded logic over the retained data for the exception's lifetime and find out what it swallowed.

go deeper

for a junior

Know that an exclusion applies to every future match and produces no visible sign when it suppresses one, so nothing will prompt you to look at it again.

for a middle

Explain the layers an exception can be enforced at — agent collection filter, ingest pipeline, rule condition, alert suppression — and say which of them destroy the underlying record and which leave it searchable.

for a senior

Show the procedure: bound the exception's window from its change history, re-run the un-excluded logic over retained telemetry when an intrusion lands there, and name retention as the limit on how far back that works.

for a principal

Own the standard: exceptions ship through version control with scope and dates, are enforced at the layer that preserves data, and have lifetimes no longer than the period over which they could be audited.

## The request A detection on cross-process memory reads keeps firing on a developer-laptop fleet, and the cause is settled: a named engineer profiles services they own, with a tool your team has no authority to remove. You have accepted that an exception is warranted. What arrives next is a request in ordinary English — "just exclude the profiler on the dev fleet" — and your job is to turn it into a change that stops the noise without removing your ability to find out, later, what it stopped. That last clause is the whole difficulty, because of one property of exclusions that is easy to state and easy to forget. ## A suppressed match is silent by construction Noise announces itself. Every false alert costs an analyst time, so a noisy rule generates pressure until somebody fixes it. An exclusion has the opposite feedback: when it suppresses something, nothing happens. No alert, no ticket, no complaint, no queue growth. Nothing in the system tells you that the exception is now covering an intruder rather than a profiler, and no moment arrives at which you are asked to reconsider. So the cost has to be made visible deliberately, at the time you write the change, through two decisions: **where** the exclusion is enforced, and **what you write down** about it. ## Where you enforce it decides what survives The same logical exception can be implemented at four different points, and they are not equivalent. | Enforced at | The alert | The record | | --- | --- | --- | | The sensor or agent's collection filter | never raised | never generated | | The forwarder or ingest pipeline | never raised | dropped before indexing | | The detection rule's own condition | never raised | indexed and searchable | | The alerting layer's suppression list | not delivered | indexed, and usually the match too | The bottom two leave you evidence; the top two do not. That is the difference between a rule you tuned and a hole you cannot examine — and it is why a request that sounds identical to everyone ("stop this alert firing") has radically different consequences depending on who implements it. Someone under pressure to reduce ingest cost reaches for the collection filter, because it saves money as well as noise. From a detection standpoint it is the most expensive option on the list: it removes visibility from every rule over that data, present and future, not only from the rule that was noisy, and it removes it for hunting too. Prefer the rule's own condition. It keeps the exception next to the logic it modifies, it leaves the underlying telemetry intact, and it is undone by editing one line rather than by re-deploying an agent policy and waiting for data to start arriving again. ## Write it so a stranger can bound the gap The second decision is documentary and it is cheap. Ship the exception as a change in version control, so it carries a diff, an author and a date, and record alongside it the scope (which condition, which assets), the owner, and the alert that justified it. The date is the part people leave out and the part that gets asked for. When something goes wrong on that fleet, the question is not philosophical — it is "what were we not alerting on, on those hosts, between 4 March and today?" An exception with a start date and an explicit scope answers that in one sentence. An exception typed by hand into a console, with no history behind it, cannot be dated at all, and the honest answer becomes "we do not know" — which is the answer that turns a bounded incident into an open-ended one. ## Finding out what it hid Together those two decisions give you a procedure rather than a hope. An intrusion is confirmed on a host in the excluded population. You take the rule's logic with the exception removed, and you run it over the retained telemetry for the exception's lifetime. Whatever it surfaces is what the exception was swallowing; a clean result narrows the intruder's likely path and is worth recording in writing either way. The limit is retention, and it is a hard one. If telemetry is kept for thirty days and the exception has been live for eight months, seven of those months cannot be examined by anyone, ever. That is not a reason to skip the check. It is a reason to state the un-searchable window as an explicit assumption in the incident's scope rather than glossing over it, and a reason to prefer exceptions whose lifetime is shorter than the window over which you could audit them. ## What none of this buys you Auditability is not safety. A dated, rule-level, well-scoped exception still removes coverage for as long as it lives, and reconstructing what it hid is something you do after the damage. What it buys is the ability to answer the question honestly and to bound the incident — which, when the alternative is a shrug, is a great deal.

  • Telemetry is retained for thirty days and the exclusion has been live eight months. What can you still establish?
    Only the last thirty days can be re-examined against the un-excluded logic; the earlier seven months are gone for everyone. Do the check anyway, then state the un-searchable window as an explicit assumption in the incident's scope, and lean on surfaces the exception never touched — identity, network, the process-creation records — for the period you cannot reconstruct.
  • A platform engineer proposes implementing the exclusion in the agent's collection filter to cut ingest cost. What do you say?
    The saving is real, but it is paid in evidence rather than in money. A collection filter removes those records from every rule and every hunt over that data, now and in future, not just from the noisy rule, and nothing can recover them afterwards. If cost genuinely forces it, scope the filter as narrowly as the noisy pattern and keep a second surface that still records the behaviour.
  • The exclusion has been in place a quarter with no alerts and no complaints. What does that tell you about its cost?
    Nothing at all — silence in that population is the exception's designed output, not evidence about it. The only way to price it is to run the rule's logic without the exception over the retained telemetry for that quarter and look at what comes back. Absence of alerts from an excluded population is a fact about the exclusion, never about the hosts.

Excluding at the rule is telling the reviewer to skip one van on the tape. Excluding at the sensor is never recording that van. Both stop the calls; only one leaves you something to re-examine.

saying these in an interview costs you the question

  • Assumes a suppressed match still surfaces somewhere as a counter
  • Drops the telemetry at the agent to make the rule cheaper
  • Reads a quiet excluded fleet as evidence the exclusion was safe
  • Cannot say when the exclusion started or what it covered
  • Expects to reconstruct a gap older than the retention window

context