A prevalence lookup shows 412 of 6,000 hosts ran the alerted certutil download command in 30 days: what can you conclude?
answer
- compared with which population?
- common is not the same as authorised
- 412 of which 6,000 hosts
- count the argument shape, not the image
- safe to lower suspicion, never to raise a verdict
basics
~20 sOnly that the behaviour is common in that population. Prevalence measures a base rate, not authorisation. It can lower your suspicion cheaply, but only if those 412 hosts are comparable to the alerting one, and it never proves anything is malicious.
solid answer
~50 sPrevalence answers "how ordinary is this here", which is a statement about a population, not about this event. A count of 412 in 6,000 says the behaviour is not rare, so the prior that any single instance is an intrusion drops — but the denominator has to be the right one. If those 412 are packaging and imaging workstations and my alert is on a finance laptop, the estate-wide figure is the wrong comparison; within that peer group the prevalence is near zero and the same command line escalates. Two more constraints: high prevalence can itself be the finding, because a tool deployed by an adversary across the estate is also common, and low prevalence never proves malice — new tooling, a migration or a red-team exercise are all rare and benign. And count the process *plus its argument pattern*: `certutil.exe` runs everywhere; `certutil -urlcache -split -f http://...` is a much sharper cut.
go deeper
Know what a prevalence lookup is: a count of how many hosts ran a given process and argument over a window. Be ready to say it measures how common something is, not whether it is allowed.
Explain why the denominator decides the answer — which population you counted over, and how a peer group such as packaging workstations versus finance laptops makes the same number mean two different things.
Show the two-minute judgement: which direction prevalence can safely push a verdict, why it never produces a malicious conclusion on its own, why the argument pattern is the unit to count, and what you record so a later hunt can reopen what you closed.
Own the risk of letting a base rate close alerts at scale. An adversary who adopts the estate's own tooling is exactly the one prevalence hides, so argue for compensating hunts, queryable closure reasons and periodic review of what those reasons have been used to dismiss.
## What prevalence is, and what it is a statement about A prevalence or frequency lookup counts how many distinct hosts (or accounts, or both) produced a given observation over a window: *how many machines ran this process with this argument pattern in the last 30 days, and how many ran it exactly once*. It is the estate's own base rate for a behaviour. On a two-minute triage budget with sixty unworked items behind you, it is one of the highest-value lookups available, because it is a single query and it resolves the largest category of alerts — behaviour that is simply how this company works. The crucial direction check: prevalence is evidence about a **population**, not about **this event**. "412 of 6,000 hosts do this" is compatible with the 413th being an adversary. What the count changes is your prior, not your verdict. ## The denominator is the whole game Estate-wide is almost never the population you want. Suppose the 412 hosts break down like this: ``` packaging / imaging workstations 398 of 430 hosts (92%) build agents 11 of 240 hosts (5%) finance laptops 1 of 2,000 hosts (0.05%) ``` Read estate-wide, the behaviour is "common" at 6.9%. Read against the alerting host's peer group, the same number supports two opposite and equally correct verdicts. On a packaging workstation, whose entire job is fetching and repackaging vendor installers, this is what the machine is for — close it as normal for this estate, with the reason recorded. On a finance laptop it is a near-unique event in its peer group and the prevalence lookup has just made the alert *more* interesting, not less. This is the core skill the question tests: choosing a comparison group — role, business unit, image, OS build — before reading the count. An analyst who quotes an estate-wide percentage without saying which population it came from has not actually used the evidence. ## What a high count does not mean **It does not mean authorised.** Common and permitted are different properties. A remote-access utility installed by an adversary across four hundred machines is highly prevalent and entirely unauthorised. Widespread is a reason to look at the *shape* of the spread: did it appear on all four hundred hosts within the same two hours, or has it been steady for a year? A behaviour with no history that is suddenly everywhere is a finding in its own right, and the prevalence lookup is what surfaces it. **It does not mean benign for this instance.** An adversary who wants to survive triage will deliberately choose the tooling your estate already runs. Blending into the base rate is a design goal of the tradecraft, which means the alerts prevalence closes are exactly the alerts a careful intruder wants closed. That is not an argument against using the cut — with sixty items in the queue you have no alternative — but it is an argument for two habits: record *what* you closed on, as a structured reason rather than free text, so a later hunt can reopen every alert closed for that reason; and keep something else watching the same behaviour that does not depend on rarity. ## What a low count does not mean Symmetrically, rare is not malicious. First-of-its-kind events are produced constantly by benign change: a new admin tool, a one-off data migration, a contractor's unfamiliar workflow, an authorised red-team exercise nobody told the SOC about. Prevalence is safe to use in the down-weighting direction and unsafe in the up-weighting one. A rare event earns *more lookup*, not a verdict. ## Count the argument, not just the binary `certutil.exe` executes on thousands of hosts for certificate work that has nothing to do with downloading. Counting the image name alone will tell you the behaviour is ubiquitous and will be true and useless. The unit that matters is the process plus the argument pattern that made the rule fire — here the URL-cache download form (ATT&CK `T1105`, ingress tool transfer). Normalise before you count: strip the variable path and URL, keep the flag shape. Two analysts counting different units will reach different verdicts on the same alert and both believe they checked prevalence. ## Putting it together on a two-minute budget 1. Normalise the observation to process plus argument shape. 2. Choose the peer group — the host's role, not the whole estate. 3. Read both numbers: how many hosts in that group, and how long it has been happening. 4. Down-weight on high, steady, in-group prevalence; escalate on out-of-group rarity or a sudden estate-wide appearance. 5. Record the closure reason in a form you can query later. That is a defensible verdict in two minutes, and it survives the question a lead will ask afterwards: *what would have had to be different for you to have escalated this?*
- The same 412 hosts all first ran that command within the same two-hour window last Tuesday. Does that change your reading?Completely. Prevalence with no history is not a base rate, it is a mass event. A behaviour that appears on four hundred machines in two hours is either a deployment somebody made or an adversary pushing tooling at scale, and both need identifying. Always read the count alongside its time distribution: steady over months supports normal-for-this-estate, a sudden simultaneous appearance is itself the finding.
- You closed the alert on prevalence and it later turns out to have been the intruder. How do you limit that damage in advance?Record the closure as a structured reason — this rule, this argument pattern, this peer group — rather than free text, so that when the reason is invalidated you can reopen every alert closed under it in one query. Pair that with a hunt that does not depend on rarity, because an adversary deliberately adopting common tooling is precisely the case prevalence hides.
- Would a per-account prevalence tell you anything the per-host count does not?Often more. Host prevalence answers whether the machine population does this; account prevalence answers whether this identity does. A packaging workstation running the command is normal, but the same command under a finance user's account on that workstation is a different observation. Where an identity can move between hosts, the account count is usually the sharper of the two.
A word's frequency tells you whether it is unusual in a language. It never tells you whether this particular sentence is a lie.
saying these in an interview costs you the question
- Treats high prevalence as proof the activity is authorised
- Quotes an estate-wide count for a host of a different role
- Calls rare activity malicious because it is rare
- Counts the binary name and ignores the arguments
- Never considers that widespread can mean widely compromised
- Closes on prevalence without recording a queryable reason