skip to content

An intel report lands with 40 indicators and a 12-month sweep bills per terabyte scanned. How do you scope it?

level: seniorimportance: nice to knowfreq 38%

answer

  1. the bill is bytes scanned
  2. fidelity before volume
  3. anchor the window to the campaign
  4. one pass, forty predicates
  5. a hit re-prices the expensive pass

basics

~20 s

Cut on three axes before spending: indicator fidelity, a time window anchored to the campaign, and the cheapest source that can carry the artefact. Then test all forty in one pass, because the bill is bytes scanned, not searches run.

solid answer

~50 s

I would not run forty indicators against every source over twelve months. First cut by **fidelity**: sweep the high-context artefacts — the trojanised installer's file hash, full domains, distinctive URL paths — before bare addresses on shared hosting, which cost most and match least usefully. Second, cut by **window**: the report says when the poisoned update shipped, so sweep that window plus a margin and widen only on a hit. Third, cut by **source**: the build system's job, dependency and registry logs are tiny and directly on point for a poisoned build tool, so they go before a year of proxy data and long before rehydrating a cold archive. Then make the query cheap — the bill is bytes scanned, so one pass carrying all forty predicates beats forty passes. I agree the byte budget and an off-peak slot up front, and a hit on a cheap pass is what buys the expensive one.

go deeper

for a junior

Know that a wide historic search costs real money and platform capacity, and that the scope gets agreed with whoever owns the platform before it runs.

for a middle

Explain what actually drives the bill — bytes read, not searches issued — and how restricting fields, sources and time window reduces it, while parallelism does not.

for a senior

Show a defensible priority order across indicator fidelity, campaign window and source cost, batching predicates into one pass, and escalating the moment a cheap pass produces a hit.

for a principal

Own the standing envelope: what an intelligence report is worth in query spend before anyone asks permission, who arbitrates when the sweep competes with production, and when a sweep is allowed to stop.

## The bill is bytes scanned On any usage-priced or capacity-constrained search backend, the cost driver is **how much data the engine had to read**, not how many searches you ran. That single fact reorganises the whole exercise: - Forty searches, one per indicator, over the same year of proxy logs, read that year forty times. - One search carrying forty predicates reads it once. - Running the forty in parallel improves wall-clock time and makes the peak load worse; it does not reduce a single byte. - A full-text regex over every field of every source is one pass, but the widest possible one — restricting to the fields that can actually carry the artefact (the query name, the URL, the hash column) is often an order of magnitude cheaper than matching anywhere. A related trick: **materialise once, sweep repeatedly**. If a campaign will produce more indicators over the coming weeks, one expensive pass that extracts a narrow projection — timestamps, client, queried name, URL, hash — into a small dataset lets every subsequent indicator be swept against something cheap. ## Three cuts, in order **Fidelity.** Not all forty artefacts are worth the same. A file hash of the trojanised installer is unambiguous and lives in small indexes. A full domain is specific. A URL with a distinctive path is specific. A bare address on shared or cloud hosting is the worst of both worlds: it sits in the highest-volume source you own, and a match may mean nothing more than that somebody used the same CDN. Sweep down that gradient, and be willing to declare the low-fidelity tail not worth scanning at all. **Window.** A supply-chain report gives you dates: when the poisoned version shipped, when it was pulled, when the second stage was first observed. Sweeping twelve months when the campaign occupies four is paying triple for the same answer. Anchor to the campaign window plus a margin either side, and widen only if something lands. The instinct to "search everything, just in case" is what turns a defensible request into one the platform owner refuses. **Source.** Rank by cost per useful answer, not by comprehensiveness. In an engineering estate hit through a build-tool update, the build system's own records — job logs, dependency resolution, artefact provenance, and the package registry's record of which version each pipeline pulled — are small, cheap, precisely targeted, and often settle the exposure question outright. Resolver and proxy logs come next. A cold archive that must be rehydrated before it can be searched at all comes last, and usually only after something cheaper has already justified it. ## The two people you have to bring with you This is where the question stops being technical. Somebody pays the query bill, and somebody owns the platform and will not let a wide scan run in business hours next to production workloads. Both are answered the same way: **go in with a scoped plan rather than an open-ended request.** Say which indicators, over which window, against which sources, at what estimated cost, in what order, with a stop condition. Take the off-peak slot for the expensive passes and run the cheap ones immediately. Agree in advance what happens if it hits, so nobody is negotiating priority at 02:00. ## When a hit changes the economics A hit on a cheap pass converts the exercise from a speculative check into incident scoping, and speculative checks and incidents are funded differently. That is the moment to go back and ask for the expensive pass — the cold-archive rehydration, the wider window — and it is normally granted immediately, because you are no longer asking whether something happened. So the sequencing is not merely thrifty; it is how you earn the budget for the expensive half. Front-loading the cheap, high-fidelity, narrow-window passes maximises the chance of holding a hit in your hand when you make the larger request. ## What good looks like under pressure A weak answer runs everything against everything and hopes. A better answer prioritises. The strongest answer also names what it chose **not** to sweep and why, before anybody asks — the bare shared-hosting addresses, the months outside the campaign window, the source that would need rehydrating — and states the condition that would reverse each of those decisions. That is what a scoped, defensible sweep looks like, and it is why the platform owner says yes the second time.

  • The platform owner refuses any wide scan during business hours. How do you work with that?
    Split the sweep. The narrow, high-fidelity passes over small sources are cheap enough to run immediately and do not threaten anything; the wide passes get an agreed off-peak slot with a byte budget and a stop condition. I would hand over the scope in writing before it runs. And I would agree the exception up front: if a cheap pass hits, this stops being a check and becomes an incident, and the wide pass runs when it needs to.
  • Why not just run all forty indicators in parallel to finish sooner?
    Parallelism reduces wall-clock time, not bytes read. Forty concurrent searches still scan the same data forty times, so on a usage-priced backend the bill is identical to running them sequentially, while the peak load on the platform is far worse — which is exactly what the platform owner is protecting against. Batching the predicates into a single pass is the change that actually reduces cost.

saying these in an interview costs you the question

  • Runs every indicator against every source with no ordering
  • Runs one search per indicator over the same data
  • Believes running searches in parallel reduces the bytes scanned
  • Sweeps bare shared-hosting addresses first because they look alarming
  • Requests a full-year rehydration before any cheap pass has run

context