Why does a 40,000-entry indicator list in a SIEM search cost more to run than one behavioural clause?
answer
- not the comparison, the rows touched
- selectivity is what buys cheap searches
- any host could be in the list
- cost recurs on every scheduled run
- match at ingest, not in the search
basics
~20 sThe list match has no selective predicate: every record must be tested against it on every scheduled run, so cost tracks total traffic volume and grows with the list. A behavioural clause narrows the records first.
solid answer
~50 sThe expensive part is rarely the per-row comparison, because most SIEM backends hash a lookup table and test a row in constant time. The expensive part is how many rows you touch. A list match on egress proxy records has nothing to narrow on: the destination host could be any of forty thousand values, so the search reads every proxy record in the window, on every schedule, forever. A behavioural clause is normally selective — restrict to non-browser process images, or to records with no preceding name resolution — and throws away the overwhelming majority of rows before the costly work starts. Two costs ride along: the list must be distributed and held in memory on every search node, and a bare match carries no context, so each hit reaches the queue as an alert containing nothing but a string. Applying the list at ingest as an enrichment field is usually the cheaper shape.
go deeper
Know that a search over every egress record is expensive because of how many records it reads, not because of how many entries it compares against. Be able to say what a selective filter is.
Explain the three cost lines separately — authoring, running, triaging — and name at least one way to keep a large list without paying the scheduled-scan cost, such as matching it at ingest.
Be ready to reason about window width, refresh behaviour of a large lookup, and the queue cost of context-free alerts, and to propose the ingest-enrichment shape rather than refusing the list outright.
Frame the cost in terms your team feels: recurring analyst hours and search capacity, and set the standard for which lists are allowed to alert at all.
## Where the cost actually lives Engineers guess wrong about this. The instinct is that comparing a value against forty thousand entries must be forty thousand times slower than comparing it against one. In practice, a SIEM's lookup or join against a static set is hashed, so the per-record test is near constant time. The cost that matters is different, and it is structural. **Rows touched.** A search is only as cheap as the number of records it has to read. Selectivity is what buys that. A behavioural clause such as "the initiating process image is not one of the browsers or the patch agent" or "the destination was reached with no prior DNS resolution from this host" throws away nearly every record early — often before decompression, if the field is indexed. An indicator list has no equivalent narrowing step: any host value in the source might be in the list, so the search must examine *every* egress record in the window. The cost is therefore proportional to your total egress volume, not to how interesting the traffic is, and it recurs on every scheduled execution. **Distribution and memory.** A large lookup has to reach every search node and be held there. Refreshes are not free, and a list that changes constantly means the searches running during a refresh see different content — a subtle source of results that do not reproduce. **Window arithmetic.** Run a five-minute search over a five-minute window and the cost is one pass. Widen the window to catch delayed events, or run the same list retrospectively across a longer period, and the same non-selective scan multiplies. ## The cost that is not machine time The run cost people notice is the search. The cost that actually hurts a small team is the queue. A behavioural rule that fires arrives with structure: a process, a parent, a user, a destination class, and a reason the clause exists. A bare list match arrives with a string. The analyst then has to reconstruct why the entry was ever in the list before they can say anything about the host — and on a small team that reconstruction is a large fraction of the alert's total cost. Multiply by however many entries in a very large list happen to overlap with shared hosting, content delivery ranges, or domains that changed hands, and the run cost you should be quoting to the person asking for the list is analyst-hours per week, not CPU. ## Where the large list stays cheap The answer is not "never use a big list". It is to move it out of the scheduled-search path. - **Match at ingest, store as a field.** Enrich each record at write time with a boolean or a tag, and searches read a pre-computed, selective field instead of joining. The comparison happens once per record ever, rather than once per record per search execution. - **Match on a low-volume source.** The same list against DNS query logs or authentication logs is a very different scan from the same list against every full egress record. - **Use it as context rather than as a trigger.** The list adds a badge to alerts that fired for another reason. That costs nothing when nothing happens and is genuinely useful when something does. - **Keep the alerting subset small.** A few hundred entries that someone will re-confirm, with a date on them, behaves nothing like forty thousand permanent ones. ## What to say when asked A good answer separates three cost lines and does not conflate them: **authoring** (the list wins, decisively), **running** (the list scans everything, the behavioural clause narrows first), and **triage** (the list arrives without context, the behavioural rule arrives with a story). Then it names the shape that keeps the first advantage without paying the other two: match the big list at ingest as enrichment, and reserve scheduled alerting for a small list with an owner and an expiry date.
- If the lookup is hashed and constant time per record, why does the list size matter at all?It matters for distribution and memory rather than comparison time: the table has to reach and sit on every search node, and refreshes create windows where concurrent searches see different content. The dominant cost is still the non-selective scan, which is set by traffic volume, but list size is what makes the table itself an operational object.
- What changes if you match the list at ingest instead of at search time?The comparison happens once per record for the life of that record, and searches then filter on a pre-computed field, which is selective and cheap. You trade a permanent, repeating scan cost for a one-off write cost, at the price of only tagging records against the list as it stood when they arrived.
saying these in an interview costs you the question
- Assumes cost scales with entries compared per row
- Ignores that the search runs on every schedule
- Counts only CPU and never analyst time
- Thinks a big list is fine because storage is cheap
- Cannot name a selective predicate for a behavioural clause