skip to content

Threat Hunting

Proactive search for an adversary no alert has fired on: forming a hypothesis, learning what the estate normally does, reading the rare tail, and banking what the hunt produced.

on this pageshow

explore

questions

page 1 of 2

A threat hunt across your ESXi hosts returns no hits — what does that negative result establish?

level: juniorimportance: must knowfreq 56%

answer

  1. a statement about the search, not the estate
  2. three ways a zero happens
  3. telemetry, logic, scope
  4. ESXi hosts carry no endpoint agent
  5. no evidence found, in what was searched

basics

~20 s

A negative hunt establishes only that the logic you ran, over the telemetry you held, for the hosts that were reporting, in the window retained, matched nothing. It is a statement about the search, not proof the estate is clean.

solid answer

~50 s

It establishes that the logic I ran, against the telemetry I had, for the hosts that were forwarding, over the window actually retained, matched nothing. That is a bounded statement about my search, not a statement about the environment. A zero has three possible parents: the behaviour never happened; it happened in a form my logic did not match; or it happened somewhere my logic could not see. On a virtualisation cluster the third is the live risk, because ESXi hosts sit outside the endpoint-agent estate and only appear in the search if each was configured to forward `auth.log` and `shell.log` off the host. So I bank the negative with its scope attached — technique identifiers, the host list, the sources, the date range, and the hosts that were silent. A negative with a scope is coverage; a negative without one is just a sentence.

go deeper

for a junior

Be ready to say, in one sentence, that a hunt with no hits establishes only that your queries matched nothing in the records you actually searched. Naming the three limits — telemetry, logic, scope — is what a screener is listening for.

for a middle

An interviewer expects you to explain how a zero is manufactured by the pipeline rather than the adversary: sources that were never forwarded, retention shorter than the time picker, and logic that encodes one variant of a technique.

for a senior

Show that you would not write the negative up until you had tested the zero itself, and that you treat the discovery of eleven blind hosts as the hunt's real output rather than an aside.

for a principal

Own the framing risk: a bounded negative is read as an assurance by everyone above you unless the write-up prevents it. Decide what your organisation is allowed to conclude from an empty hunt, and say so in the standard.

## The claim a hunt can and cannot make Threat hunting is a deliberate search for adversary behaviour that no alert fired on. Most hunts end the same way: nothing. That outcome is normal and it is still worth something — but only if you are precise about what it means, because "we hunted for it and found nothing" is the single easiest sentence in security operations to over-read. A hunt result is the output of a pipeline with three independent stages, and every one of them must have worked before the zero says anything about an adversary at all: 1. **The telemetry existed.** The records describing the behaviour were generated, shipped, parsed and retained for the period you claim to have covered. 2. **The logic matched.** Your query expressed the behaviour in the form the adversary would actually have used, not just the one form the write-up illustrated. 3. **The scope covered the ground.** The hosts, accounts and time range you searched include the place and moment the behaviour would have occurred. If any stage failed, the zero is about your pipeline, not about the estate. That is why a bare "no findings" is close to worthless: it does not say which of the three it is asserting. ## Why a virtualisation cluster makes this concrete Suppose a public write-up describes an actor pivoting from stolen administrator credentials into vCenter, enabling SSH or the ESXi Shell on the hosts, and running commands directly on the hypervisors. You hunt it. Every one of the three stages is fragile here: - ESXi is not a general-purpose operating system running your endpoint agent, so the richest telemetry you have everywhere else simply does not exist on these machines. What you have is host syslog — authentication events, shell command records, the host agent and vCenter agent logs — and only if each host was pointed at a collector. - Host-local ESXi logs live on a small scratch area and can roll quickly or not survive a reboot, so a host that was not forwarding has no recoverable history at all. - Retention for that syslog stream may be far shorter than the retention your search interface implies for other sources. A ninety-day time picker over a fourteen-day stream returns zero rows for the seventy-six days that were never there, and returns them without an error. So a truthful negative from that hunt might be: *no records matching three named techniques were found in forwarded ESXi authentication and shell logs from 37 of 48 hosts, over the 14 days retained*. Every clause in that sentence is load-bearing, and dropping any of them turns a defensible statement into an indefensible one. ## Getting the direction of the claim right The reliable formulation is that the hunt found no evidence of the behaviour in the records searched. It did not establish the behaviour's absence, and it certainly did not establish the absence of an adversary — an intruder who never touched a hypervisor is entirely compatible with your result, and so is one who did it on a host that was never forwarding. Nor does a zero say anything about your detection rules. A hunt query is an ad-hoc search a human wrote and ran once; a deployed rule is a different artefact with a different failure mode. Conflating the two is a common wrong answer. ## Why the negative is still worth banking Three reasons a properly recorded empty hunt earns its keep: - **It is coverage evidence.** Six months later, someone will ask whether this technique was ever looked for in the cluster. A scoped record answers that; a memory does not. - **It exposes the gap.** The most useful output of that hunt is often not the security result at all — it is the discovery that eleven hosts are invisible. That gap is actionable in a way the zero is not. - **It is a baseline.** When the same hunt runs next quarter with more hosts forwarding and longer retention, you can say the estate's answerability improved, which is a real claim about the defence. ## The failure mode on the reading side A hunter writes a careful, bounded negative and an executive, an auditor or a project sponsor reads it as "we checked, we're fine". That collapse happens by default unless the write-up prevents it in plain language, near the top, in the same sentence as the result. Put the gaps where the conclusion is, not in an appendix nobody opens. ## The one thing a zero does prove It proves something about you rather than the adversary: that a named hypothesis was tested, to a stated depth, over stated ground, on a stated date, by a named person. That is a modest claim, and it is the only one the evidence supports.

  • Someone says the hunt proved the cluster was not compromised by that technique. How do you correct them without making the hunt sound worthless?
    I would say the hunt tested a specific hypothesis over specific ground and found nothing there, which is genuinely useful and worth recording. Then I would name the ground: fourteen days of forwarded logs from thirty-seven of forty-eight hosts. The correction is not that the hunt failed, it is that its conclusion has edges, and the eleven unsearched hosts are the next piece of work rather than a footnote.
  • Does an empty hunt tell you anything about whether your detection rules for that technique work?
    No. A hunt query is an ad-hoc search a human ran once; a deployed rule is a separate artefact running continuously with its own inputs and its own ways of silently breaking. A zero from the hunt says nothing about whether the rule would fire, and a rule that has never fired says nothing about whether the technique occurred. They are separate questions and get answered separately.
  • The hunt was based on one public write-up of the technique. What does that limit?
    It limits the logic stage. The write-up shows one observed implementation, and my query almost certainly encodes that implementation's specifics — a particular command form, a particular sequence. An actor who reached the same objective through the vCenter API rather than a host shell would produce no match. I record which variants the logic covered so the negative is not read as covering the technique in general.

Searching three rooms of a twelve-room house, with the lights off in nine of them and the last hour of the CCTV overwritten, and reporting "nobody in the house". The search was real; the conclusion is bigger than the search.

saying these in an interview costs you the question

  • Treats no hits as proof the environment is clean
  • Reports 'no findings' with no host list or date range
  • Assumes the search window equals the data's retention
  • Thinks an empty hunt validates the detection rules
  • Drops silent hosts from the count instead of recording them

context

open as a page

Your hunt finds suspicious activity that no detection rule ever alerted on. Does that make it less likely to be malicious?

level: juniorimportance: must knowfreq 55%

basics

~10 s

No. A rule set only covers behaviour someone wrote a rule for, over sources someone connected. Silence measures your detection coverage, not the activity's intent. Judge the behaviour on its own evidence.

open as a page

What must a hunt query gain before it can run unattended as a detection rule?

level: juniorimportance: must knowfreq 62%

basics

~20 s

It has to stand without its author: logic narrowed from browse-everything to one defensible claim, an explicit threshold and evaluation window, a named owner who answers when it misfires, and triage notes saying what benign matches look like and what the analyst does next.

open as a page

What makes a threat-hunting hypothesis falsifiable, and why does 'are we breached?' fail?

level: juniorimportance: must knowfreq 72%

basics

~20 s

A falsifiable hunting hypothesis names one adversary behaviour, the telemetry where that behaviour would leave a record, and the result that would kill it. 'Are we breached?' names no behaviour and no observation, so no query can end it.

open as a page

When does a worry about intruder-installed remote-access tools become a hunt, a rule, or a ticket?

level: juniorimportance: must knowfreq 62%

basics

~20 s

Route by cost. A hunt is a one-off, time-boxed search that answers the question once. A standing rule creates triage work on every match forever, so it needs recurrence, precision and an owner. A ticket is for when only remediation is left.

open as a page

What telemetry preflight checks come before the first threat-hunt query?

level: juniorimportance: must knowfreq 60%

basics

~20 s

Check which sources actually record the behaviour you are hunting, on which hosts and build images, which fields those records carry, and how far back they reach. A query over a source that never collected the record returns zero rows and proves nothing.

open as a page

What does a C2 beacon look like in TLS-encrypted flow and proxy metadata?

level: juniorimportance: must knowfreq 72%

basics

~20 s

Repeated connections from one host to one destination at a near-constant interval, each session short and roughly the same size, often with more bytes out than in. Metadata carries the rhythm and the volume, never the content.

open as a page

What does a parent-child process baseline give a threat hunter who has no alert in hand?

level: juniorimportance: must knowfreq 68%

basics

~20 s

It records which processes ordinarily spawn which on this estate, so a hunter judges a pair rather than a process. A command shell is unremarkable everywhere; a command shell spawned by a document reader is not.

open as a page

In threat hunting, what is stack counting, and what does a rare value actually prove?

level: juniorimportance: must knowfreq 62%

basics

~20 s

Stack counting groups records by a chosen field, counts every value, and reads the smallest counts first. A rare value proves only that it is unusual in this population over this window: a prioritisation signal, never a verdict.

open as a page

Why does hunting for LSASS credential dumping by tool name miss most of the technique?

level: juniorimportance: must knowfreq 62%

basics

~20 s

One technique has many procedures. LSASS memory can be read by comsvcs.dll's MiniDump export, a renamed dumping utility, a vulnerable signed driver, direct syscalls or a process clone. A tool-name hunt catches only whoever did not rename it.

open as a page

A vendor report names OAuth consent persistence — how do you turn it into a testable hunt hypothesis?

level: middleimportance: must knowfreq 58%

basics

~20 s

Restate the report's technique as a claim about your own tenant, name the record that would carry it — the identity provider's consent and delegated permission-grant entries — and fix the result that would kill the claim before querying anything.

open as a page

Two hundred agentless printers beacon to one CDN-hosted name every 60 seconds — how do you reach a verdict?

level: seniorimportance: must knowfreq 50%

basics

~20 s

Prevalence decides it: behaviour shared by a whole device population points at shipped firmware, not per-host compromise. Attribute the destination and the change that started it, then close as a benign true positive with the discriminator recorded.

open as a page

A service account's baseline shows Sunday-night logons only, and last night it authenticated to eighty hosts at 03:00 — how do you reach a verdict?

level: seniorimportance: must knowfreq 58%

basics

~20 s

Decompose the deviation — hour, host count, logon type, source address — then hunt the disconfirming record: change tickets, the patch job's own logs, the source hosts. The owner's explanation is a hypothesis to test, not a close.

open as a page

Your LSASS handle-access hunt returns 30,000 benign rows — why is that fine in a hunt but not a rule?

level: seniorimportance: must knowfreq 50%

basics

~20 s

A hunt's output is reviewed once, in bulk, by an analyst already holding the context, so noise can be aggregated and filtered interactively. A standing rule's noise recurs forever, lands on whoever is on shift, and costs a fresh triage every firing.

open as a page

What must an empty threat hunt's write-up record for the negative to count as coverage later?

level: middleimportance: should knowfreq 44%

basics

~20 s

Record the hypothesis and technique identifiers, the exact logic, the sources it ran against, the date range bounded by real retention, the hosts in scope versus searched versus silent, and what a hit would have looked like.

open as a page

A hunt hit lands in a self-hosted CI runner's job logs. What do you preserve before you start pivoting?

level: middleimportance: should knowfreq 48%

basics

~20 s

Export the matched records with their source, event time and collection time, and capture the query provenance: exact text, source searched, time range and timezone. Build-platform run history expires and the next job wipes the workspace.

open as a page

Sysmon Event IDs 19, 20 and 21 all record WMI activity — which one shows persistence is armed?

level: middleimportance: should knowfreq 48%

basics

~20 s

Event ID 21, the WmiEventConsumerToFilter binding. A filter (19) is a trigger with nothing attached and a consumer (20) is an action nothing calls; only the binding wires them together. Even then, 21 proves the subscription was registered, not that it has ever run.

open as a page

Why can a signal be too noisy for a standing detection rule and still be worth hunting?

level: middleimportance: should knowfreq 48%

basics

~20 s

A rule is judged one firing at a time and pays triage cost on each, so a common behaviour buries the queue. A hunt reviews the whole population at once, stacking by rarity and joining to a source of truth.

open as a page

In a hunt over auditd logs, how do you confirm execve records exist on a Linux fleet?

level: middleimportance: should knowfreq 46%

basics

~20 s

Sample a representative host per build image and confirm an execve-matching audit rule is loaded and producing SYSCALL/EXECVE record pairs in the search index. Without a matching rule the kernel writes nothing, so the source is silent rather than empty.

open as a page

How do you measure a beacon's check-in interval when the implant jitters its sleep?

level: middleimportance: should knowfreq 55%

basics

~10 s

Work on the distribution of inter-arrival gaps per source-destination pair, not the mean. Jitter randomises around a base sleep, so gaps still cluster in a band; measure that band's tightness relative to its centre.

open as a page

Why does a threat-hunting baseline expire, and what should you record alongside one?

level: middleimportance: should knowfreq 52%

basics

~20 s

A baseline is a snapshot of an estate that is under continuous authorised change, so it decays as the estate moves. Record the learning window, the population covered, the assumptions it rests on, and the events that would invalidate it.

open as a page

Your stack of Kubernetes pod-create records by container image is all count 1 - what went wrong?

level: middleimportance: should knowfreq 48%

basics

~20 s

The counted value carries per-build entropy: image tags and digests change on every build, so each deploy becomes a distinct value. Normalise before counting - strip tag and digest and group on registry host plus repository path.

open as a page

Which LSASS dumping variants does a sweep of Sysmon Event ID 10 process access miss?

level: middleimportance: should knowfreq 42%

basics

~20 s

Sysmon Event ID 10 records one process opening a handle to another, with the access mask granted. It catches user-mode dumpers, including direct-syscall ones, but not a signed driver reading memory from kernel mode, nor a read aimed at a forked clone.

open as a page

Your ESXi hunt query returned zero rows — how do you establish the zero is real before writing it up?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Prove the data was there before concluding the behaviour was not. Check which hosts delivered events and when, check real retention, check the filtered fields are still populated, and run a broader search that must return known-benign activity.

open as a page

The release manager says your build-runner finding was an engineer debugging a failing job. How do you settle it?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Settle it on evidence, not assertion: reconcile the commands against the repository's pipeline definition, find the authenticated identity and how it reached the runner, look for a change record, and check whether anything left the host. Unreconciled means escalate.

open as a page

A hunt enumerated existing WMI subscriptions estate-wide; the promoted rule watches creation events. What coverage was lost?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Everything already in place. The hunt asked what exists right now; the rule only sees what is created from the moment it goes live, so the entire existing backlog — including anything planted before telemetry reached that host — is invisible to it forever. Promotion has to ship a recurring state sweep alongside the rule.

open as a page

Your OAuth consent hunt found only sanctioned grants — what does that disprove, and what does it not?

level: seniorimportance: should knowfreq 44%

basics

~20 s

It disproves exactly the claim you wrote: over the grants you queried, in the window you queried, none fell outside the sanctioned set. It says nothing about other persistence routes, other windows, or whether the tenant is clean.

open as a page

You have twelve hunt ideas and one week a month to hunt — how do you rank them and close each out?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Gate on telemetry: an idea whose data cannot cover the window becomes a collection request, not a hunt. Then rank by relevance, by sceptically assumed coverage, and by whether a hit is routable. Time-box each one and write it up.

open as a page

Your preflight shows half the Linux fleet has no execve auditing — what do you deliver?

level: seniorimportance: should knowfreq 41%

basics

~20 s

Run the hunt over the hosts that do report, publish the result with its denominator attached, and file a named collection requirement for the dark half. The hypothesis is deferred, not answered, and the write-up must say so.

open as a page

Stacking pod images leaves a 4,000-row tail of singletons - do you work it or abandon the hunt?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Test whether rarity carries information before spending the hours: measure what share of records sits in the tail. If self-service registries make every image rare by construction, regroup on a field with convention or abandon and write it up.

open as a page

showing 1–30 of 39