What makes a threat-hunting hypothesis falsifiable, and why does 'are we breached?' fail?
answer
- a hunt has to be able to end
- name the behaviour, not the fear
- where would the record appear
- state the killing result in advance
- 'are we breached' is unanswerable by query
basics
~20 sA falsifiable hunting hypothesis names one adversary behaviour, the telemetry where that behaviour would leave a record, and the result that would kill it. 'Are we breached?' names no behaviour and no observation, so no query can end it.
solid answer
~50 sA hunting hypothesis is a claim about your own estate that a finite piece of telemetry can disprove. Three parts make it falsifiable: a specific adversary behaviour ("a third-party application holds a delegated mail-read scope in our tenant that nobody sanctioned"), the records that would carry it (the identity provider's application consent and permission-grant audit entries), and a stated kill condition — the result that ends the hunt as a no ("every delegated grant in the population resolves to an application on the sanctioned integration list"). 'Are we breached?' fails all three: it names no behaviour, so no query follows from it; it names no observation surface; and no result can ever settle it, because a clean query speaks only about what it queried. Writing the kill condition down before you query is what stops a hunter widening the search until something finally looks suspicious.
go deeper
Be ready to state the three parts out loud: the behaviour, the record that would carry it, and the result that would kill the claim. Expect to be handed a vague worry and asked to sharpen it into one sentence.
Explain why a clean result speaks only about the population and window you queried, and why fixing the kill condition in advance is what lets the hunt end at all.
Show that you hold the line when a hunt returns nothing, reporting the disproof and its exact scope rather than widening the query until something looks odd enough to escalate.
Own the framing standard for the team: every hunt in the backlog carries a written claim and a written kill condition, so a quarter of disproofs reads as measured coverage rather than as nothing found.
## Why a hunt needs a hypothesis at all Threat hunting is looking for an adversary in your own environment when nothing has alerted. Because nothing has alerted, nothing bounds the search: the estate is large, the telemetry is larger, and "go look for evil" can absorb unlimited time and end whenever the hunter gets tired or finds something odd enough to write up. The hypothesis exists to bound the work. It is a claim about **your** environment, written before the first query runs, that a finite piece of telemetry can **disprove**. The word that matters is *disprove*. A hunt that can only succeed — that can only end by finding something — has no stopping rule, and a hunt with no stopping rule is not a piece of work, it is a mood. ## The three parts **1. A behaviour, not a fear.** The hypothesis has to name something an adversary would *do*, stated concretely enough that you could recognise it in a record. "An unsanctioned third-party application holds a delegated mail-read scope in our tenant, granted through user consent" is a behaviour. "Someone is stealing our email" is a fear. The behavioural framing also survives the adversary swapping tooling, which a hypothesis pinned to one domain or one application identifier does not. **2. The record that would carry it.** Name the observation surface before you start: which log, from which system, over which population. For the consent example that is the identity provider's consent and delegated-permission-grant audit records — no endpoint telemetry is involved at all, which is worth saying out loud because hunters reach for process telemetry by reflex. If no source you hold would carry the behaviour, you do not have a hunt yet; you have a collection requirement, and confirming the source is present and covers the period is a separate preflight step before the hunt is scheduled. **3. The result that would kill it.** State, in advance, the outcome you would accept as a *no*: "every delegated grant in the population resolves to an application we own or contract for, with scopes matching its documented purpose." This sentence is the one worth arguing over with whoever brought you the worry, because it is the sentence that decides when the hunt is finished. Agreeing it afterwards is not agreeing it: a hunter who finds nothing and has not pre-registered a kill condition will keep widening the query, and eventually everything looks a little strange. ## Why 'are we breached?' fails It fails on all three counts, but the deep failure is the third. There is no query result that answers it. If the query returns nothing, the honest reading is "nothing matched this query over this population in this window" — which is a statement about the query, not about the estate. Turning that into "we are not breached" is the single most common wrong move in this domain, and it is the reason the claim has to be narrow enough that a clean result actually means something. ## Direction of the claim Get the direction right in both directions. A record that a consent was granted proves that an authorization was recorded — not that the application ever used it, and not that a credential was stolen. The absence of a matching record proves that nothing matching your query was recorded — not that the behaviour did not occur, and certainly not that the estate is clean. Every clean hunt result is bounded by the population, the window and the source that produced it, and the write-up carries those bounds or it will be quoted without them. ## The two ways framing goes wrong *Unfalsifiable*: "no evidence of malicious application consent exists in the tenant." The word doing the work — *malicious* — is the very thing under test, so no query settles it, and the hunt runs forever. *Indicator lookup*: "no application with client identifier `8e2f...` appears in our tenant." This one is decidable, but it is not a hunt; it is a lookup of one artefact, and it is dead the moment the adversary registers a different application. A hypothesis should sit at the level of the behaviour, so that changing infrastructure does not invalidate it. ## Where a hypothesis comes from Three honest sources, all converging on the same skeleton. A **report** naming a technique someone else observed; a **technique identifier** from a public catalogue, used as a label for a behaviour rather than as content; or a **crown-jewel asset** — start at the thing worth stealing and ask what someone who wanted it would do that you would see. In every case you still have to write the claim about your own estate and fix the kill condition yourself; none of the three sources hands you one. ## In the interview Expect to be handed a vague worry and asked to sharpen it in one sentence, then asked what result would make you stop. A candidate who answers the second question with "when I've checked everything" has not understood the exercise. A candidate who says "the hunt ends as a no when every grant in the window resolves to a sanctioned application, and I'd agree that with the person who raised it before I query" has.
- A colleague proposes 'an attacker is in our SaaS tenant' as this quarter's hunt. How do you sharpen it?Ask two things: which behaviour, and which record would carry it. Push it down to something like "a third-party application holds a delegated mail-read scope granted by consent to an application outside our sanctioned list". Then agree what result would kill it before anyone queries. If neither of you can name a record that would carry the behaviour, it is not a hunt yet, it is a worry.
- If the hunt disproves the hypothesis, was the week wasted?No. A disproof is a result: it narrows where the adversary can be, and it leaves behind a reusable query, an agreed population and a documented scope. The wasted hunt is the one with no kill condition, which either runs forever or ends when the hunter loses interest. Report the disproof in its own right, stating exactly which claim it disposed of.
- How specific should a hypothesis be before specificity starts hurting?Specific enough to name an observable, general enough to survive the adversary changing tooling. "The application with this client identifier consented in our tenant" is an artefact lookup that dies the moment the identifier changes. "Any application outside the sanctioned list holding a mail-read scope" names the behaviour and stays true across tooling, while still being decidable by one query.
It is a lab experiment rather than a search party. You write down in advance the result that sends you home, otherwise you keep running the experiment until the data eventually agrees with you.
saying these in an interview costs you the question
- Says the hypothesis is proved when the hunt finds nothing
- States the goal as 'look for anything unusual'
- Names a threat-actor group instead of an observable behaviour
- Makes a single domain or client identifier the hypothesis
- Decides the kill condition only after seeing results