skip to content

As SIEM platform owner, how do you answer a detection engineer who wants a correlation window widened to six hours?

level: principalimportance: nice to knowfreq 28%

answer

  1. answer with a price, not a verdict
  2. state = first-stage rate x window x size
  3. rarer anchor before longer window
  4. streaming memory versus scheduled re-query
  5. eviction is an invisible miss: meter it

basics

~20 s

Answer with a price and alternatives, not yes or no. Held state scales with the first stage's event rate times the window, so ten minutes to six hours is roughly thirty-six times more. Seek the same coverage cheaper first.

solid answer

~50 s

A streaming sequence rule must remember every unmatched first stage until its window expires, so its state is roughly first-stage rate times window times record size, across however many rules do this. Going from ten minutes to six hours is a thirty-six-fold multiplier on that rule's share of memory, and if the anchor is every build minting a deploy credential, that is a large number. So I ask three things. Can the anchor be rarer, matching only privileged deploy roles? Can the join key be narrower? Does it have to be streaming at all, or can a scheduled join re-run over a six-hour lookback, moving the cost from resident memory to repeated query cost, at the price of latency and needing deduplication? If it must be wide and real-time, then something else on the platform gets displaced, and I want that trade named, funded and signed rather than quietly absorbed.

go deeper

for a junior

Know that a sequence rule must remember every unmatched first event until its window ends, so a longer window costs memory as well as precision.

for a middle

Be able to estimate the cost as event rate times window times retained size, and explain why making the first stage rarer helps far more than trimming fields.

for a senior

Walk the alternatives ladder — rarer anchor, narrower key, scheduled re-query, tiered rules — and insist on metering evictions so a cap never silently kills coverage.

for a principal

Own the trade explicitly: whose budget funds a wide window, which single sequence earns it, what the platform gives up, and where the accepted miss is written and signed.

## Why this is a budget conversation, not a configuration change A streaming correlation engine holds **in-flight window state**: for every first-stage event that has not yet found its partner, a record is retained until the window expires. Multiply that across rules and you get the platform's real memory profile — often millions of partially matched sequences. The rough model is: > state per rule ~= (rate of unmatched first-stage events) x (window length) x (bytes retained per pending match) Every term is negotiable, and window length is the one the requester reached for first because it is the only one visible in the rule file. A move from ten minutes to six hours is a factor of thirty-six on the second term. If the anchor stage is common — a build pipeline minting a deploy credential fires on every merge — the absolute number is what matters, not the ratio, and it can be the largest single line item on the engine. ## Answer with a ladder, not a verdict A flat refusal loses the detection; a flat yes silently taxes every other rule. Work down the terms: 1. **Make the anchor rarer.** Restrict the first stage to the variant that actually matters — credentials minted for privileged deploy roles rather than every role. This usually beats every other lever, because it attacks the rate term directly and often loses no coverage of the behaviour you care about. 2. **Trim what is retained.** Keep the join key, the timestamp and the few fields the second stage needs, not the whole record. A modest win, but free. 3. **Narrow the join key.** A high-cardinality key spread across many pending values is fine; a coarse key that fans out into a huge per-key match space is not. Cap per-key state and know what the cap does. 4. **Change the execution model.** A scheduled join that re-runs every fifteen minutes over a six-hour lookback holds no long-lived memory at all; it pays in repeated query cost against the log store, in detection latency, and in needing deduplication because the same pair will be found by several overlapping runs. For a six-hour behaviour this is very often the right answer: nobody is responding in seconds to something whose stages are hours apart. 5. **Tier the rule.** Keep a narrow, cheap, real-time version for the fast case and run the wide version on the batch path. Two rules, two latencies, one behaviour covered. Only when all of that is exhausted is the question actually "do we buy more memory", and by then both sides know what they are buying. ## The trap: silent eviction Most engines respond to a state cap by evicting. An evicted pending match is a **false negative that produces no record of itself** — the rule looks healthy, the dashboard is green, and coverage has quietly gone. As the platform owner this is your instrumentation to own: emit an eviction counter per rule, alert when it is non-zero, and treat sustained eviction as a broken detection rather than a capacity metric. A cap without a meter is worse than a smaller window honestly chosen, because at least the smaller window is written down. ## The argument that a wider window is not safety The engineer's strongest card is a real case where the stages were two hours apart. That is a good argument for funding *that* sequence generously; it is not an argument for widening everything. And it has a limit worth saying out loud: **any finite window can be waited out.** Six hours is a better bet than ten minutes, not a guarantee, so a behaviour that matters should also have a detection that does not depend on the two stages being close — a property of the second stage alone, or a hunt over stored data. Widening is a purchase with diminishing returns, and pretending otherwise is how a platform ends up spending its whole budget on one rule's tail. ## What gets written down Whatever you settle on, three things belong on the record: the window each rule got and the evidence behind it; the state budget that decision protects and what was displaced to fund any exception; and the behaviour the estate has knowingly accepted it will not catch, with the detection owner's name against it. That last one is the point of the whole exercise. A window is an accepted miss, and an accepted miss that nobody signed is indistinguishable, six months later, from coverage.

  • The engineer says the case they are chasing had two hours between the stages. Does that settle it?
    It justifies a wide window for that sequence, not for the estate. Fund it with the rarest anchor you can construct and keep other rules at their measured lengths. Also say the uncomfortable part: any finite window can be waited out, so that behaviour needs a second detection that does not depend on the two stages being close together.
  • How do you stop a state cap from silently destroying detections?
    Meter it. Every eviction of window state is a false negative that leaves no trace in the rule's output, so emit a per-rule eviction counter, alert when it is non-zero, and treat sustained eviction as a broken detection rather than a capacity statistic. An unmetered cap makes the platform look healthy while coverage disappears.
  • When is moving the rule to a scheduled batch join clearly the right call?
    When the behaviour's own tempo is hours, so no one was going to respond in seconds anyway, and when the log store can answer a six-hour lookback cheaply enough on a schedule. You trade resident memory for repeated query cost and added latency, and you must deduplicate because overlapping runs will find the same pair more than once.

saying these in an interview costs you the question

  • Treats window length as free because storage is cheap
  • Approves the change without pricing the extra retained state
  • Refuses outright instead of offering a cheaper equivalent
  • Adds a hard state cap with no eviction metric
  • Argues that a wide enough window eventually makes the rule safe

context