skip to content

Your ESXi hunt query returned zero rows — how do you establish the zero is real before writing it up?

level: seniorimportance: should knowfreq 38%

answer

  1. distrust the zero until the data is proved
  2. run a search that must return something
  3. events by host, by day
  4. a zero background rate is a data smell
  5. silent fields fail without erroring

basics

~20 s

Prove the data was there before concluding the behaviour was not. Check which hosts delivered events and when, check real retention, check the filtered fields are still populated, and run a broader search that must return known-benign activity.

solid answer

~50 s

I treat a zero as suspect until I have shown the search could have returned something. Four checks. First, a positive control: run a deliberately broad search over the same sources and window that must return known-benign activity — an administrator's routine shell session or a backup job's authentication. If that comes back empty too, the problem is my pipeline, not the estate. Second, the per-host delivery picture: which of the forty-eight ESXi hosts sent events, on which days, and when any of them stopped. Third, real retention: the source may hold fourteen days while the interface offers ninety. Fourth, the fields: if a parser change stopped extracting the field I filter on, the filter matches nothing and raises no error. Only once the search demonstrably had ground under it does the zero become a security result rather than a data result.

go deeper

for a junior

Know that an empty query result might mean the data was missing rather than the activity was absent, and that you should check the source was reporting before you report a clean hunt.

for a middle

Be able to name the concrete ways a zero is manufactured: hosts not forwarding, retention shorter than the search range, and a filtered field that stopped being extracted after a parser change.

for a senior

Demonstrate the positive control as a habit — a broad search that must return known-benign activity — and show that you read a zero background rate as a data smell rather than as a clean estate.

for a principal

Set the expectation that no negative is banked without evidence the search had data under it, and that hunts blocked on telemetry are recorded as blocked rather than as clean.

## A zero is a measurement, and measurements need controls The worst outcome of a hunt is not missing an intruder. It is writing up an empty result that was empty because nothing could ever have matched, and then having that record cited as evidence of coverage. A zero row count is produced identically by a clean estate and by a broken search, and nothing in the interface distinguishes them. So the discipline is: before the zero becomes a conclusion, demonstrate that the search had ground under it. ## The positive control The strongest single check is a search you know must return rows. Pick activity you can corroborate outside the log platform — a scheduled maintenance window, a named administrator's routine session, a backup service authenticating on a known cadence — and search for it over the same sources, the same hosts and the same window as the hunt. If that returns rows, the pipeline was carrying data and your narrow query's zero is meaningful. If it returns nothing, you have found a telemetry failure and the hunt has not started yet. A weaker but useful variant is to strip your own query down: drop the qualifying conditions and search for the noun alone — any remote shell enablement on any host at all, rather than one done outside a change window. That coarse version should return a background rate. **A background rate of exactly zero is a red flag about the data, not good news about the estate**, because in a real cluster administrators do these things routinely. ## The delivery picture, per host and per day Count events by host and by day across the window. This answers three questions at once: which hosts were in the search at all, whether any host stopped mid-window, and whether the volume looks like the cluster you know. On a virtualisation estate this matters more than almost anywhere else, because the hosts carry no endpoint agent and appear in your data only if each one was individually configured to forward its syslog. A host that sent nothing was not searched and cannot be reported as clean; it belongs in the write-up as a gap. The instinct to drop it from the denominator so the percentage looks better is exactly the wrong move. It is also worth knowing that host-local ESXi logs sit on a small scratch area and may not survive a reboot, so "we can go back and check that host later" is often false. ## Retention, honestly Search interfaces will let you select a range far longer than a given source retained. The result for the missing period is zero rows and no warning. Establish the earliest timestamp actually present for each source in the hunt, and record that as the covered window. If the sources hold fourteen days, the negative is fourteen days long no matter how the search was run. ## The fields you filtered on A filter on a field that is no longer extracted matches nothing and reports success. Parser and schema changes do this quietly: a message format changes on an upgrade, the command text ends up inside an unparsed blob, and every rule and query keyed on the extracted field goes silent without erroring. Sample a handful of raw records from inside the window and confirm the fields your logic depends on are populated in them. This is also why a query written against a normalised field is worth re-running against the raw text once, as a cross-check. ## The logic itself Finally, sensitivity. Describe the record that would have matched. If you cannot write that sentence, you cannot claim the query could have found anything. And be explicit about variants: a hypothesis drawn from one public write-up encodes one implementation, so an actor who achieved the same objective through the management API rather than a host shell produces no match and no contradiction. ## Where this stops This is verification of a one-off hunt's ground truth, not health monitoring of deployed detections — those are separate practices with separate mechanisms. Here the question is narrower and answerable in an afternoon: did my search have data under it, on which hosts, for how long? Answer that first, and the write-up becomes a security statement instead of an unexamined zero.

  • The broad control search also comes back empty. What do you conclude and what do you do next?
    I conclude the hunt has not happened yet — the zero is about the pipeline. Then I work backwards: are the hosts configured to forward, is the collector receiving, did the parser change, is the index the one I think it is. I record the hunt as blocked on telemetry rather than as a negative result, because banking it as a clean hunt would put a false coverage claim on the record.
  • Eleven of forty-eight hosts sent no events. Why not just report the negative for the thirty-seven that did?
    That is exactly what I report — but the eleven must be named alongside it, not dropped. Reporting a percentage over the reduced denominator makes the coverage look complete when a quarter of the cluster was never looked at, and those are precisely the hosts an intruder would prefer. The eleven become an engineering item with an owner, which is often the hunt's most valuable output.
  • How would a parser change produce a zero without anyone noticing?
    A message format shifts — an upgrade, a new agent version — and the field the query filters on stops being extracted. The query is still syntactically valid, runs successfully and returns no rows, because the value it is comparing against now lives inside unparsed text. Nothing errors. Sampling raw records from inside the window and confirming the field is populated catches it in minutes.

Before reporting that the metal detector found nothing on the beach, you sweep it over your own keys. If it stays silent there, you have learned about the detector, not the beach.

saying these in an interview costs you the question

  • Accepts a zero row count as a security result
  • Never checks which hosts actually delivered data
  • Confuses configured retention with data present
  • Reports coverage over the reduced denominator
  • Assumes a valid query means a working query

context