Your DNS-tunnelling detection has raised no alerts in five months - what does that silence prove?
answer
- three causes, not one
- no output is not a measurement
- did the search run at all?
- quiet estate versus muted logic
- positive evidence, not absence
basics
~20 sNothing on its own. Zero alerts has at least three causes: there was no such activity, the rule is broken or muted by a pipeline change, or the data never reached the rule. Silence is a question, not an assurance.
solid answer
~50 sOn its own it proves nothing, because three very different situations produce exactly the same output: the estate really was quiet, the rule ran but its logic can no longer match anything, or the events it needed never arrived in the first place. A detection only speaks when it fires, so its own output can never tell you what it missed - the absence of an alert is indistinguishable from an absence of activity. To turn silence into a claim you need positive evidence: that the scheduled search actually executed, that its input was non-empty (a count of events matching just the log source and prefilter portion of the rule), and that a known-matching synthetic event still produces an alert end to end. Without those three, `no alerts` is a sentence about your telemetry, not about your adversary.
go deeper
Be ready to say plainly that zero alerts has several possible causes, and name them: no activity, muted logic, or data that never arrived. Do not offer the estate is clean as your first answer.
Explain how a rule can run to completion, report success and still match nothing, because an empty result set is a legal answer rather than an error. Say which two queries you would run to separate the cases.
Show how you convert silence into a checkable claim: execution telemetry, an input-volume floor on the rule's prefilter, and a synthetic event that must produce an alert. Say which one you would build first with limited time.
Own the framing that a detection's own output cannot measure what it misses. Be ready to state what liveness assurance you will fund, for which detections, and what you will tell a risk owner about everything you chose not to instrument.
## The claim being made, and why it is not supported When a detection produces no alerts for months, the sentence people hear is *nothing bad happened*. The sentence the data actually supports is *this rule emitted no output*. Those are only the same thing if you can also show that the rule ran, that it ran over the data it was written for, and that its logic still selects the thing it was written to select. None of that is implied by the absence of alerts, and in a real pipeline each of the three fails routinely. A worked case makes the gap concrete. A corporate network runs its own recursive resolvers, and every client query is logged: the name asked for, the client that asked, the query type. A detection watches those logs for the signature of data tunnelled inside DNS - a large number of distinct subdomains under one registered zone, unusually long labels, high-entropy names, an odd query-type mix - all pointing at a name server the attacker controls. The rule exists. It is deployed. It has fired nothing for five months. Then an outside party tells you that traffic from your network has been reaching that authoritative name server the whole time. The rule was not wrong about DNS tunnelling; the rule stopped seeing the field it filters on after a parser release, and its silence was read as safety. ## The three explanations, and how to tell them apart **1. The estate was quiet.** The rule ran, its input was populated, its logic evaluated real events and matched none. This is the only case where silence means what people wanted it to mean - and it is the case you must *earn* the right to claim, not the default. **2. The logic is mute.** The search executes and succeeds, but a predicate can no longer be true: a field was renamed or dropped by a parsing change, a value's shape changed (case, a trailing dot on a fully qualified name, truncation of a long name), the timestamp the search windows on moved, or the rule points at an index or log source that no longer carries this data. Nothing errors, because an empty result set is a legal answer to a query, not a fault. **3. The data never arrived.** The rule is fine and there is nothing to evaluate: the sensor stopped, the collector's queue is dropping, verbose query logging was turned off on the resolvers for performance. (Detecting that across a whole estate is a discipline in its own right; here the point is only that it is a third explanation, and it looks identical from the alert queue.) The order of investigation follows cost. Check the rule's execution telemetry first - did the scheduled job run, over how many events, with how many matches. `Ran, scanned zero events` and `ran, scanned four million, matched zero` are two completely different findings and both are cheap to get. Then check the population of the specific fields the rule filters on, as a count per day, which will show a step change if a parsing release moved them. Only when the job ran, the input was full and the fields were populated can you say the third thing: this rule saw its data and matched nothing. ## Why the rule cannot grade itself A detection's failure mode leaves no trace in the detection's output. A false positive is visible - somebody worked the alert and closed it as benign. A false negative produces nothing at all: no record, no row, no count. That asymmetry is why precision can be estimated from a queue and a miss rate cannot, and why a rule's own history is evidence about the times it worked and silent about every time it did not. The practical consequence is that assurance for a detection has to come from *outside* its output. Three signals cover most of it: - **Execution telemetry** - the scheduled search ran on time and completed. - **An input-volume floor** - the count of events matching only the rule's log source and prefilter stays above a level you expect, so you learn when the input dries up rather than when the alerts do. - **A known-matching event** - a synthetic canary injected on a schedule, travelling the same collection and parsing path as production data, which must produce an alert. With those in place, silence finally carries information: the plumbing is alive, the logic still matches what it is supposed to match, and nothing that looks like the modelled behaviour occurred. Without them, five months of quiet is exactly as strong as five months of not looking.
- Which of the three causes would you check first, and why that order?Cheapest and most decisive first. Start with the rule's execution telemetry - did the scheduled search run, and over how many events - because `scanned zero` and `scanned four million, matched zero` are different findings for a couple of minutes of work. Then look at the daily population of the fields the rule filters on. Only when both are healthy may you conclude that the rule saw its data and matched nothing.
- Somebody outside tells you the tunnelling has been running all along. Does that make your rule a false negative?Not automatically, and the distinction matters for the fix. A rule that ran, saw populated data and did not match is a logic miss. A rule whose predicate could not be true because a field moved never evaluated anything - that is a broken contract between the parsing pipeline and the detection content. Both look the same from the queue; only one is fixed by editing the rule.
- What positive evidence would you want standing behind every high-value detection?That the scheduled job executed, that its input was non-empty measured on the log source and prefilter portion alone, and that a synthetic event known to match still produces an alert through the normal path. With those three, silence becomes a statement about the environment; without them it is a statement about nothing.
saying these in an interview costs you the question
- Treats zero alerts as proof the estate is clean
- Says the rule works because it deployed without errors
- Assumes a green scheduled job means data was present
- Confuses no alerts with a low false-positive rate
- Claims a miss rate from the rule's own output