skip to content

A two-stage correlation rule never matches because one source ingests four minutes late. What do you change?

level: middleimportance: should knowfreq 46%

answer

  1. in time, but read too early
  2. two clocks: event gap versus arrival delay
  3. delay evaluation, do not widen the match
  4. size it from the lag tail
  5. paid for in latency and state

basics

~20 s

Add a grace period: hold each window open, or delay evaluating it, until the slow feed has landed, and match on event timestamps. That buys correctness with detection latency and state held longer. Do not widen the correlation window instead.

solid answer

~50 s

The pair was in time; the rule read the data too early. A build system's job log lands within seconds, while the cloud control-plane feed arrives in roughly four-minute batches, so a scheduled search over the last ten minutes of event time runs before the second stage is visible — and if the buckets do not overlap, that record's event time falls in a range already evaluated and is never revisited. The fix is a grace period: evaluate a window only once `now` is past its end plus the expected lag, or keep pending partial matches alive past window end so a late second stage can still complete them. Size it from the measured arrival-lag distribution at a high percentile, not the mean, and say what happens to anything later. You pay in alert latency and in state held for window plus grace. The sequence then fires late, but in time.

code

text · 11 lines
text
rule:     ci_credential_mint THEN cloud_api_call_from_unknown_egress WITHIN 10m
schedule: every 10m, scanning the previous 10m of EVENT time, no overlap, no delay

event_time   ingest_time   source        record
10:07:03     10:07:06      ci-runner     job 8841: deploy role assumed, session "job-8841"
10:09:41     10:13:35      cloud-audit   session "job-8841" first API call, src 203.0.113.44
              ...           (cloud-audit is delivered in ~4 minute batches)

run at 10:10  scans event_time 10:00-10:10  -> stage 2 not ingested yet: no match
run at 10:20  scans event_time 10:10-10:20  -> stage 2 present, but its event_time is
                                               outside the scanned range: never revisited

go deeper

for a junior

Know that a rule can only match records that have arrived. If one feed lands minutes after the fact, a perfectly correct rule can still produce nothing at all.

for a middle

Explain the grace period as delaying evaluation rather than loosening the match, and name what it costs: later alerts and partial matches held longer.

for a senior

Show that you would measure the lag distribution, pick a tail percentile, and say out loud what happens to anything that arrives past it.

for a principal

Decide which detections deserve real-time latency at all, and route the chronically slow feeds onto a batch path instead of making every alert on the platform wait for the worst source.

## Two different clocks, two different knobs Every correlation rule deals with two independent quantities: - **The gap between the events**, measured on their own timestamps. The correlation *window* constrains this. - **The delay before a record is visible to the rule.** The *grace period* (sometimes a watermark, or an evaluation delay) accommodates this. Confusing them is the most common mistake on this topic. Widening the window to compensate for a slow feed changes what the rule considers *related* — it starts joining events genuinely twenty minutes apart — when the actual problem is that the events were three minutes apart and the reader was early. ## How the miss happens Take a sequence over two very different feeds: a build system's job log, shipped within seconds, and a cloud control-plane audit feed delivered in batches every few minutes. Stage A is a job minting a short-lived deploy credential; stage B is that credential's first API call from an address outside the estate's egress ranges; the join key is the session name carrying the job id. Suppose the rule is scheduled every ten minutes and scans the previous ten minutes of **event** time, with no overlap and no delay. A run covering 10:00-10:10 fires at 10:10 and cannot see a stage-B record whose event time is 10:09 but which will not be ingested until 10:13. The next run covers 10:10-10:20; by then the record exists, but its event time is outside the scanned range. The gap between the two stages was under three minutes, comfortably inside a ten-minute window, and the rule still produced nothing — permanently. ## The two mechanisms that fix it **Delay evaluation.** Do not evaluate a time bucket until `bucket_end + grace` has passed in wall-clock terms. The 10:00-10:10 bucket is evaluated at 10:15 instead of 10:10, by which time the slow feed's records for that period have landed. This is what stream processors call a watermark, and it is the cleanest version because each event-time bucket is still evaluated exactly once. **Hold the partial match longer.** In a stateful streaming engine, keep an unmatched stage A alive past the end of its window by the grace period, so a stage B that arrives late but whose event time falls inside the window can still complete it. The match condition is unchanged; only the retention is extended. A third pattern, overlapping lookbacks — re-running every five minutes over the last thirty — also works, but it re-evaluates the same period repeatedly, so it will emit the same alert several times unless you deduplicate on the matched tuple. ## Sizing it Measure the arrival lag per source: for each record, the difference between when it happened and when it became queryable. Use a high percentile of that distribution — p99, not the mean — because the mean hides exactly the tail you are trying to survive. Then answer the question that everyone forgets: **what happens to records that arrive later than the grace period?** They are silently dropped from matching. That is an accepted miss and it belongs in the rule's documentation, or it belongs on a slower second-pass rule that re-runs over a longer lookback and accepts producing a next-hour finding rather than a page. ## What it costs - **Latency.** Every alert from the rule is now at least the grace period later than it could have been. On a sequence whose response is to cut off a credential, minutes matter, and that is a real trade rather than a free correctness fix. - **State.** Pending partial matches are held for window plus grace instead of window, so a four-minute grace on a ten-minute window is a forty per cent increase in retained state for that rule. - **Review.** The grace period is pinned to a pipeline property that changes. When a feed's delivery mode changes, a grace period sized for the old behaviour quietly starts dropping matches again. ## The claim you can and cannot make While this rule was misfiring, its output was empty, and empty output is not evidence that the behaviour did not occur. The honest statement after the fix is that the rule now matches the sequence with roughly four minutes of added latency, that the period before the fix has no coverage from this rule, and that anything arriving beyond the grace percentile is still missed. If the estate needs assurance for that earlier period, it comes from re-running the logic retrospectively over stored data, not from the fact that nobody was paged.

  • Why not simply widen the correlation window from ten minutes to twenty instead?
    Because that changes what counts as related. A twenty-minute window joins events genuinely twenty minutes apart and brings in benign pairs, while the real problem was that the events were three minutes apart and the rule read the data before one of them arrived. A grace period delays reading; it does not loosen matching.
  • How do you choose the grace period's length?
    From the measured arrival-lag distribution of the slowest source in the rule, at a high percentile rather than the mean, since the mean hides the tail you are trying to survive. Then state explicitly what happens to arrivals beyond it: silently dropped, or picked up by a slower second-pass rule. Re-measure when the pipeline changes.
  • What if the slow feed's lag is hours rather than minutes?
    Do not buy that with grace on a real-time rule, because every alert would then be hours late. Split the logic: a fast rule over the feeds that arrive promptly, and a scheduled join with a long lookback for the slow feed, whose output is accepted as a next-hour finding rather than a page. Deduplicate where the two overlap.

Two photographs were taken three minutes apart, but one of them came back from the lab four minutes later. Waiting for the lab is the fix; claiming the photographs were taken further apart is not.

saying these in an interview costs you the question

  • Widens the correlation window to compensate for ingest lag
  • Blames the detection logic when the data had not arrived yet
  • Sets the grace period from the average lag rather than a tail percentile
  • Assumes overlapping lookbacks need no deduplication
  • Never says what happens to arrivals later than the grace period

context