skip to content

What telemetry preflight checks come before the first threat-hunt query?

level: juniorimportance: must knowfreq 60%

answer

  1. start from the hypothesis, not the source
  2. coverage, fields, look-back window
  3. which hosts, which build images
  4. zero rows can mean no collection
  5. missing source becomes a filed requirement

basics

~20 s

Check which sources actually record the behaviour you are hunting, on which hosts and build images, which fields those records carry, and how far back they reach. A query over a source that never collected the record returns zero rows and proves nothing.

solid answer

~50 s

A preflight turns "we have Linux logs" into "this record type exists, on these hosts, with these fields, for this many days back". I start from the hypothesis — say persistence planted as a `systemd` timer or a user cron entry across our server fleet — and name the records that would betray it: `auditd` `execve` syscall records for whatever the timer executes, write watches on the unit and cron directories, resolver query logs if the payload calls out. Then I confirm each one exists where I intend to search: which build images have the audit rules loaded, whether the fields I need survived normalisation, how many days are searchable. If something is missing, the hunt is deferred or scoped down and the gap becomes a named collection requirement. What I am avoiding is a zero-row result being written up as "no evidence of compromise" when the estate never recorded the evidence at all.

go deeper

for a junior

Be ready to list what you check before querying: which record would show the behaviour, on which hosts, with which fields, and over how many days. Say plainly that zero rows from a source that never collected is not evidence of absence.

for a middle

Expect to explain how you verify coverage per build image rather than per fleet, and how a field can vanish in normalisation between the host and the search index.

for a senior

Show that the preflight changes the deliverable: a scoped hunt with a stated denominator, plus a filed collection requirement, rather than a quiet empty result presented as a clean estate.

for a principal

Own the argument that unhunted is not the same as clean in SOC reporting, and that collection gaps need named owners and review dates rather than being absorbed silently by the hunt team.

## What a preflight is A threat hunt is a search for adversary behaviour that no alert fired on. The search runs over records that some machine wrote down earlier — and if nothing wrote the relevant record, the search cannot fail in a visible way. It returns zero rows, which looks identical to a clean estate. The telemetry preflight is the short piece of work done **before** the first query that establishes whether the question is answerable at all, and over what. It has four parts, and they are worth naming separately because they fail separately. **1. Which source would carry this behaviour?** Start from the hypothesis, not from the sources you happen to like. A hypothesis about persistence installed as a `systemd` timer or a user cron entry on a Linux fleet is a hypothesis about (a) a file appearing under a unit or cron directory and (b) something being executed on a schedule. Those map to different records: file writes are visible to an audit watch rule on the directory or to a file-integrity tool; execution is visible to `auditd` `execve` syscall records, or to an EDR sensor's process telemetry if one is installed. If the payload reaches out, the resolver query log and proxy or flow records may show it. Writing this mapping down is the preflight's first output. **2. Does that source exist where I intend to search?** This is the part people skip. On Linux, `auditd` only records what a loaded rule matches — there is no default "log everything" state — so a fleet built from mixed images can have execution auditing on the newer builds and nothing at all on the older ones. Older images may also carry no EDR sensor. The honest preflight answer is a denominator: not "we have auditd", but "execve rules are loaded on 214 of 480 hosts, covering these two build images". **3. Do the records carry the fields the query needs?** A record type existing is not the same as the field existing. Process telemetry that gives you only an image name will not answer a question about what a timer executed, because the interesting part is in the arguments. A DNS resolver query log shows the *name a client asked for* — it does not show what the process ultimately received, and it is not a record of a connection. Check the fields, and check them after normalisation, because a mapping step can drop or rename what you planned to filter on. **4. How far back is searchable?** A hunt hypothesis usually implies a look-back window. If the resolver query log holds two weeks and the hypothesis concerns activity from last quarter, the query will run happily and answer a much smaller question than the one you asked. That constraint belongs in the write-up, not in the footnotes. ## Why this matters more than it sounds The direction of the claim is what is at stake. A hunt that finds nothing over a source that *was* collecting is weak evidence of absence — it says the behaviour, as you defined it, did not appear in those records. A hunt that finds nothing over a source that was never collecting is **no evidence of anything**, and reporting the two the same way is the characteristic failure of an inexperienced hunter. Leadership reads "we hunted for scheduled-task persistence and found nothing" as reassurance, and the reassurance is manufactured. ## What the preflight produces The preflight has deliverables even when it fails: - a **scoped hunt** over the portion of the estate that does report, with the scope stated in the result; - a **collection requirement**: a named, specific ask — which record type, on which hosts, retained how long, and which hunt it unblocks — handed to whoever owns those systems; - a note in the hunt backlog that this hypothesis is **deferred**, not answered, so nobody re-reads the empty result later as a clean bill of health. It is cheap. Confirming that the record exists on a representative host per build image, and that the field survives normalisation, takes an hour. Discovering the same thing on day four of a week-long hunt costs the week.

  • How is a zero-result hunt over a collecting source different from a zero-result hunt over a source that was never collecting?
    The first is weak evidence of absence: the behaviour, as defined by the query, did not appear in records that existed. The second is no evidence at all — you have measured your own blind spot. They must be written up differently, because a reader treats both as reassurance unless the scope is stated on the result.
  • You have half an hour before a hunt is due to start. What do you check first?
    The single record type the hypothesis depends on most, on one representative host per build image: does it exist, does it carry the field the query filters on, and how many days back does it go. That one check is what most often kills or reshapes the hunt, and everything else can be refined as the hunt runs.
  • Why check field-level coverage rather than just that the log source is present?
    A source can be present and still not answer the question. Execution telemetry with only a process name cannot answer what a scheduled job ran, because the payload is in the arguments. Normalisation can also drop or rename a field between the host and the search index, so the field should be checked in the records you will actually query, not in the vendor's schema.

Before reviewing the tape you check that the camera was recording, was pointed at the door, and that the disk still holds last month. An empty screen from a camera that was never switched on is not proof that nobody walked in.

saying these in an interview costs you the question

  • Treats an empty result as proof the behaviour did not happen
  • Assumes every host in the fleet runs the same logging configuration
  • Confirms "we have Linux logs" without checking record types
  • Starts the hunt and discovers the collection gap four days in
  • Believes a DNS query log shows what the process received back

context