In a SIEM correlation rule requiring event A then event B, what does the time window control, and what does widening it cost?
answer
- the rule's assumption about tempo
- two failure directions, not one
- short window: patient sequence walks past
- long window: unrelated pairs get joined
- measure real gaps, document the miss
basics
~20 sThe window is the largest event-time gap the rule will join two stages across; outside it the pair is never matched. Too short misses a patient sequence; too long joins unrelated events, raising false alarms and the state held.
solid answer
~50 sA sequence rule says A, then B, by the same entity, within N minutes, and N is the rule's whole assumption about how fast the behaviour happens. Take a rule where a build job mints a short-lived cloud deploy credential and that credential's first API call then comes from an address outside our egress ranges. If the observed gap in real cases is a few minutes, ten minutes catches it. Cut N to two minutes and an adversary who simply waits is never matched, and the rule's silence looks exactly like a clean estate. Stretch N to six hours and the nightly blue/green deploy, which mints the same credential and later egresses from a newly provisioned ASN, lands inside the same pattern most days, so precision collapses and pending state is held far longer. Pick N from measured gaps, then write down what it deliberately misses.
go deeper
Be ready to state in one sentence what the window does: it caps the gap between the two stages, and outside it the rule never joins them. Know both failure directions, not just noise.
Explain how you would pick the number from observed gaps rather than a default, and why the join key and the ordering matter as much as the length does.
Show that you would defend the number with measured data, and that a noisy sequence gets fixed by tightening the join before anyone touches the window.
Own the fact that the chosen window is an accepted miss. Decide which sequences are worth a long one, and make that limitation written and owned rather than implicit in a rule file.
## What a sequence rule actually states A correlation rule (a *sequence* rule, in most engines) states something of the form: **event A, then event B, involving the same entity, within N units of time.** It has three separable parts, and candidates who blur them get the rest wrong: - **The stage predicates** — what counts as A and what counts as B. Each is usually a filter over one log source. - **The join key** — the field that makes the two records *the same story*: a session name, a user principal, a host, a job id. Without it you are joining any A to any B. - **The window (N)** — the maximum gap the rule will accept between the two stages, measured on the events' own timestamps. The window is not a formality inherited from a vendor sample. It is a **claim about the adversary's tempo**: you are asserting that when this behaviour is real, the two stages happen within N of each other. ## The worked case A continuous-integration job mints a short-lived cloud deploy credential; the build system records that it happened, and the cloud control plane records that credential's subsequent API calls, including the source address. Stage A is the mint; stage B is the first use of that credential's session from an address outside the estate's known egress ranges. The join key is the session name, which carries the build job id. The window encodes the belief that whoever steals a token out of a build pipeline uses it promptly, because it expires. ## The two failure directions **Too short.** The stages of the real behaviour fall further apart than N — the intruder pauses, or the credential is stolen and used at the start of the next shift — and the pair is never joined. Nothing is emitted. This is the direction people underrate, because its symptom is *silence*, and a rule's silence is indistinguishable from an estate where the behaviour never happened. You cannot compute a false-negative rate from a detection's own output; the only way to know a window is long enough is to execute the behaviour and see whether the rule matched. **Too long.** The window stops carrying information. Two events that are six hours apart are barely evidence of each other, and the benign population moves in. In the case above, the routine blue/green deploy mints exactly the same kind of credential and its traffic egresses from a newly provisioned address range; at ten minutes those two facts never coincide, at six hours they coincide most days, and the rule becomes a daily benign true positive — the events are real, the sequence genuinely occurred, and it is authorised. Length also costs the platform: a streaming engine must remember every unmatched first stage until its window expires, so held state scales roughly with the first stage's event rate times N. ## Choosing N Measure rather than guess. Gather the gaps you have evidence for — purple-team executions of the behaviour, confirmed past cases — and the gaps in the benign population that shares the same stages. Choose N above the malicious gaps you intend to catch and, where possible, below the typical benign gap. If the two distributions overlap heavily, that is the finding: **the window alone cannot separate them**, and you must tighten the join instead — exclude the deploy service principal, require the second stage from outside a known egress list, require the credential to carry a privileged role. That is also the answer to "the rule is noisy, shall we shorten the window?" Shortening trades away real detections to fix a precision problem the window did not create. Shorten only when the measured gaps say N was always more generous than the behaviour needs. ## Things the window is not - It is **not the search range**. A scheduled search may scan an hour of data while the rule still refuses any pair more than ten minutes apart. - It is **not a grace period**. A grace period delays *when* the rule reads the data so that a slow feed has arrived; it does not loosen what counts as a match. - It **does not degrade gracefully**. A pair one second past N is not a weaker alert; it is no alert. - It does not enforce ordering on its own. "A then B" requires the engine to compare the events' timestamps; a co-occurrence join that ignores order will happily match the reverse sequence, which is often an entirely different, benign thing. ## Write down the miss Once N is chosen, the rule's documentation should say what evidence set it, which behaviour it knowingly will not catch (the same sequence executed slowly), and who accepted that. Otherwise the limitation is invisible, and a future reader treats six quiet months as coverage.
- If the false alarms all come from the nightly deploy, why not just shorten the window?Because the deploy sits inside the window legitimately, so shortening buys precision by throwing away real detections. Fix the join instead: exclude the deploy service principal, require a privileged role on the minted credential, or require the second stage from outside known egress ranges. Shorten the window only if measured gaps show it was always more generous than the behaviour needs.
- Does the window measure the gap between the events or between their arrival at the SIEM?Between the events, using their own timestamps; arrival time only decides when the rule can see them. A rule that measures the gap on arrival will happily join two records an hour apart because a batch shipped them together, and will refuse a genuine pair that arrived on different feeds.
- What do you record once you have settled on ten minutes?The evidence — the distribution of gaps you measured — the behaviour the window deliberately does not catch, and who accepted that miss. That turns the limit into an owned, reviewable decision. Without it, the rule's silence reads as an all-clear to the next person, when it may only mean the adversary was slower than your number.
A window is a stakeout with a fixed shift. Go home after ten minutes and the second visitor who arrives at minute eleven is never connected to the first; stay for a week and everyone who walks past twice starts to look like a pair.
saying these in an interview costs you the question
- Treats the window as an arbitrary default copied from a vendor sample
- Says a longer window is always the safer choice
- Assumes no alerts means the sequence did not happen
- Measures the gap on arrival time rather than event timestamps
- Fixes a false alarm by shrinking the window instead of tightening the join