skip to content

The platform owner refuses your 12-month LSASS sweep on compute cost — how do you still run the hunt?

level: principalimportance: nice to knowfreq 30%

answer

  1. sample first, then extrapolate cost
  2. cheapest filter before the expensive one
  3. one variant per query
  4. buy the window where yield is high
  5. the claim shrinks with the window

basics

~20 s

Scope before arguing: sample a slice to estimate hit density and cost, run one procedure variant at a time with the cheapest filter first, project only the fields you need, and spend the full window only on the low-volume variants.

solid answer

~50 s

I would stop asking for one enormous query and turn the sweep into a series of cheap ones. Sample a few days to measure hit density and scan cost, then extrapolate — now the conversation has numbers in it. Run variant by variant, cheapest discriminator first: filter on the access-mask read bit and exclude the four known-good source images before touching anything expensive like call-stack matching, and project only the fields I need. Chunk the window and run it off-peak. Then decide where the full window is worth buying: low-base-rate variants get twelve months, the broad noisy one gets thirty days. What I will not trade is the honesty of the claim — a thirty-day sweep supports "not seen in thirty days", not a clean bill for the year, and whether the rest stays unexamined is the risk owner's call, not the platform owner's.

go deeper

for a junior

Know that searching a long history costs real money and time on a shared platform, and that narrowing the fields and the filter before running matters.

for a middle

Be able to describe concrete cost reductions: sampling to estimate, applying the cheap indexed filter before an expensive one, projecting a few fields, splitting the window into off-peak chunks.

for a senior

Show the trade being made deliberately — full window for low-base-rate variants, a short window for the broad one — and state precisely what claim each choice supports.

for a principal

Own the boundary between scoping and accepting risk: you engineer the sweep down and price the options, but a risk owner, not the platform owner, decides whether an unexamined eleven months is acceptable, and you make sure they see that choice.

## Why the full window is worth wanting A hunt is retrospective by nature. Its value is a statement about a *period*: this behaviour did or did not appear across the estate over some span of time. Intrusions are routinely discovered long after initial access, and the procedure you are sweeping for may have run months before anyone wrote a rule for it. A seven-day sweep answers a question nobody needed asked; the whole reason to reach for the full searchable window is to make a claim that covers the time an intruder could plausibly have been present. That is also exactly why it is expensive. Sweeping every process-access record on a large fleet for a year is one of the heaviest things a hunter can ask a search platform to do, and the platform owner refusing is a legitimate act of stewardship, not obstruction. ## Engineer the sweep down before you negotiate Most of the cost is usually self-inflicted, so exhaust that first. - **Sample and extrapolate.** Run the sweep over a few days on a slice of hosts. That gives you the hit density, the scan volume and the wall time, and turns "I need twelve months" into "twelve months costs roughly this much and, at the observed hit rate, will surface roughly this many rows for review". Nobody can negotiate against a request with no number in it. - **One variant at a time.** The technique expands into several procedures; each has its own query and its own cost. Running them as separate jobs spreads load, produces partial results early, and lets you stop when one of them finds something. - **Cheapest predicate first.** Filter on the narrow, indexed thing — the target image and the memory-read bit in the access mask — and exclude the small set of known-good source images *inside* the query, before applying anything expensive such as call-stack pattern matching. The ordering can change the cost by an order of magnitude. - **Project, do not export.** Ask for the five fields the analysis needs, not whole records. - **Chunk and schedule.** Split the window into slices run off-peak. The same total work spread across nights is a very different proposition for a shared cluster than one query at 10:00 on a weekday. - **Move the data once.** If the platform genuinely cannot support the scan, a single narrow extraction into somewhere cheap, analysed offline, is sometimes the pragmatic answer — with the obvious care about where a copy of that data now lives. ## Then trade window against variant What survives is a real trade, and the right shape is usually asymmetric. Variants with a low base rate — an unusual source image, an unbacked call stack, a driver load on a workstation fleet — are cheap per unit of window, so buy them the full twelve months. The broad variant that matches every management agent in the estate is expensive and low-yield per row, so buy it thirty days and say so. ## The decision that is not yours The scoping is your job. Accepting the residual is not. If the affordable sweep covers thirty days, then the organisation has an unexamined eleven months for that technique, and someone with a risk mandate — not the person who owns the search cluster — decides whether that is acceptable. Your obligation is to put the choice in front of them in the terms they can act on: what the sweep would cost, what claim each budget buys, and what remains unknown at each price point. This is also where precision of language matters most. "We hunted for credential dumping and found nothing" is heard as a clean bill of health for the enterprise. "Across thirty days of process-access records, on the ninety-four percent of hosts reporting, the five procedures on our list produced no unexplained hits; the preceding eleven months were not searched" is a sentence a decision can be built on. Never let the first version be the one that reaches an executive summary. ## Make the next negotiation shorter Instrument the hunt itself. Record scanned volume, wall time, rows returned, rows reviewed and the verdicts they produced. Over a handful of hunts that gives you a cost-per-hunt figure with outcomes attached — including outcomes that are not intrusions, such as a discovered telemetry gap, an unmanaged host or a misconfigured agent, which are real findings that the platform owner's own team usually cares about. A standing, predictable compute allocation for hunting is far easier to win with that history than by arriving with a large one-off request each quarter. And when the platform owner sees findings that improve their estate coming out of the spend, the negotiation stops being adversarial.

  • Why sweep the full retention window at all rather than the last week?
    Because a hunt's product is a claim about a period, and intrusions are commonly discovered long after initial access. The procedure may have run before any rule for it existed. A week-long sweep answers a question the existing detections already cover; the value of the retrospective sweep is precisely the time it reaches back over.
  • How do you make the next compute negotiation easier?
    Instrument the hunts: scanned volume, wall time, rows reviewed, and the verdicts produced — including non-intrusion findings such as hosts that stopped reporting or agents misconfigured. A cost-per-hunt figure with outcomes attached buys a standing allocation, which is far cheaper to administer than a large one-off request each quarter.
  • What do you tell leadership when only 30 days were affordable?
    State the claim exactly as narrow as it is — this behaviour did not appear in thirty days of records on reporting hosts — and name the residual: eleven months unexamined for this technique. Then let the risk owner decide whether to fund the rest. The failure here is the summary sentence that quietly drops the scope.
  • The platform owner offers a fixed nightly compute budget instead. Is that better than one large grant?
    Usually yes. A predictable recurring allocation lets you chunk long windows across nights, re-run sweeps as new variants are published, and plan hunts rather than beg for them. It also converts an adversarial one-off negotiation into a stable arrangement both sides can size their capacity against.

saying these in an interview costs you the question

  • Runs the whole window unfiltered and blames the platform
  • Reports a 30-day sweep as a clean bill for the year
  • Cannot estimate the sweep's cost before requesting it
  • Treats compute cost as somebody else's problem entirely
  • Lets the platform owner decide what residual risk is acceptable

context