skip to content

SIEM & Detection Engineering

How an adversary behaviour becomes a rule that survives real data: picking the observable, writing correlation and rarity logic over it, and keeping it firing after it ships.

on this pageshow

explore

questions

page 1 of 2

In a SIEM correlation rule requiring event A then event B, what does the time window control, and what does widening it cost?

level: juniorimportance: must knowfreq 68%

answer

  1. the rule's assumption about tempo
  2. two failure directions, not one
  3. short window: patient sequence walks past
  4. long window: unrelated pairs get joined
  5. measure real gaps, document the miss

basics

~20 s

The window is the largest event-time gap the rule will join two stages across; outside it the pair is never matched. Too short misses a patient sequence; too long joins unrelated events, raising false alarms and the state held.

solid answer

~50 s

A sequence rule says A, then B, by the same entity, within N minutes, and N is the rule's whole assumption about how fast the behaviour happens. Take a rule where a build job mints a short-lived cloud deploy credential and that credential's first API call then comes from an address outside our egress ranges. If the observed gap in real cases is a few minutes, ten minutes catches it. Cut N to two minutes and an adversary who simply waits is never matched, and the rule's silence looks exactly like a clean estate. Stretch N to six hours and the nightly blue/green deploy, which mints the same credential and later egresses from a newly provisioned ASN, lands inside the same pattern most days, so precision collapses and pending state is held far longer. Pick N from measured gaps, then write down what it deliberately misses.

go deeper

for a junior

Be ready to state in one sentence what the window does: it caps the gap between the two stages, and outside it the rule never joins them. Know both failure directions, not just noise.

for a middle

Explain how you would pick the number from observed gaps rather than a default, and why the join key and the ordering matter as much as the length does.

for a senior

Show that you would defend the number with measured data, and that a noisy sequence gets fixed by tightening the join before anyone touches the window.

for a principal

Own the fact that the chosen window is an accepted miss. Decide which sequences are worth a long one, and make that limitation written and owned rather than implicit in a rule file.

## What a sequence rule actually states A correlation rule (a *sequence* rule, in most engines) states something of the form: **event A, then event B, involving the same entity, within N units of time.** It has three separable parts, and candidates who blur them get the rest wrong: - **The stage predicates** — what counts as A and what counts as B. Each is usually a filter over one log source. - **The join key** — the field that makes the two records *the same story*: a session name, a user principal, a host, a job id. Without it you are joining any A to any B. - **The window (N)** — the maximum gap the rule will accept between the two stages, measured on the events' own timestamps. The window is not a formality inherited from a vendor sample. It is a **claim about the adversary's tempo**: you are asserting that when this behaviour is real, the two stages happen within N of each other. ## The worked case A continuous-integration job mints a short-lived cloud deploy credential; the build system records that it happened, and the cloud control plane records that credential's subsequent API calls, including the source address. Stage A is the mint; stage B is the first use of that credential's session from an address outside the estate's known egress ranges. The join key is the session name, which carries the build job id. The window encodes the belief that whoever steals a token out of a build pipeline uses it promptly, because it expires. ## The two failure directions **Too short.** The stages of the real behaviour fall further apart than N — the intruder pauses, or the credential is stolen and used at the start of the next shift — and the pair is never joined. Nothing is emitted. This is the direction people underrate, because its symptom is *silence*, and a rule's silence is indistinguishable from an estate where the behaviour never happened. You cannot compute a false-negative rate from a detection's own output; the only way to know a window is long enough is to execute the behaviour and see whether the rule matched. **Too long.** The window stops carrying information. Two events that are six hours apart are barely evidence of each other, and the benign population moves in. In the case above, the routine blue/green deploy mints exactly the same kind of credential and its traffic egresses from a newly provisioned address range; at ten minutes those two facts never coincide, at six hours they coincide most days, and the rule becomes a daily benign true positive — the events are real, the sequence genuinely occurred, and it is authorised. Length also costs the platform: a streaming engine must remember every unmatched first stage until its window expires, so held state scales roughly with the first stage's event rate times N. ## Choosing N Measure rather than guess. Gather the gaps you have evidence for — purple-team executions of the behaviour, confirmed past cases — and the gaps in the benign population that shares the same stages. Choose N above the malicious gaps you intend to catch and, where possible, below the typical benign gap. If the two distributions overlap heavily, that is the finding: **the window alone cannot separate them**, and you must tighten the join instead — exclude the deploy service principal, require the second stage from outside a known egress list, require the credential to carry a privileged role. That is also the answer to "the rule is noisy, shall we shorten the window?" Shortening trades away real detections to fix a precision problem the window did not create. Shorten only when the measured gaps say N was always more generous than the behaviour needs. ## Things the window is not - It is **not the search range**. A scheduled search may scan an hour of data while the rule still refuses any pair more than ten minutes apart. - It is **not a grace period**. A grace period delays *when* the rule reads the data so that a slow feed has arrived; it does not loosen what counts as a match. - It **does not degrade gracefully**. A pair one second past N is not a weaker alert; it is no alert. - It does not enforce ordering on its own. "A then B" requires the engine to compare the events' timestamps; a co-occurrence join that ignores order will happily match the reverse sequence, which is often an entirely different, benign thing. ## Write down the miss Once N is chosen, the rule's documentation should say what evidence set it, which behaviour it knowingly will not catch (the same sequence executed slowly), and who accepted that. Otherwise the limitation is invisible, and a future reader treats six quiet months as coverage.

  • If the false alarms all come from the nightly deploy, why not just shorten the window?
    Because the deploy sits inside the window legitimately, so shortening buys precision by throwing away real detections. Fix the join instead: exclude the deploy service principal, require a privileged role on the minted credential, or require the second stage from outside known egress ranges. Shorten the window only if measured gaps show it was always more generous than the behaviour needs.
  • Does the window measure the gap between the events or between their arrival at the SIEM?
    Between the events, using their own timestamps; arrival time only decides when the rule can see them. A rule that measures the gap on arrival will happily join two records an hour apart because a batch shipped them together, and will refuse a genuine pair that arrived on different feeds.
  • What do you record once you have settled on ten minutes?
    The evidence — the distribution of gaps you measured — the behaviour the window deliberately does not catch, and who accepted that miss. That turns the limit into an owned, reviewable decision. Without it, the rule's silence reads as an all-clear to the next person, when it may only mean the adversary was slower than your number.

A window is a stakeout with a fixed shift. Go home after ten minutes and the second visitor who arrives at minute eleven is never connected to the first; stay for a week and everyone who walks past twice starts to look like a pair.

saying these in an interview costs you the question

  • Treats the window as an arbitrary default copied from a vendor sample
  • Says a longer window is always the safer choice
  • Assumes no alerts means the sequence did not happen
  • Measures the gap on arrival time rather than event timestamps
  • Fixes a false alarm by shrinking the window instead of tightening the join

context

open as a page

In Splunk SPL, why can a hunt search with no index or time bounds report a false 'nothing found'?

level: juniorimportance: must knowfreq 68%

basics

~20 s

An unbounded search can be finalized or truncated before it reads everything, so an empty result proves only that the search stopped early, not that nothing happened. Pin the index and the time range in the base search.

open as a page

In a first-seen detection, what does 'first seen' actually assert about the value it fired on?

level: juniorimportance: must knowfreq 62%

basics

~10 s

Only that this value never appeared in the rule's lookback window for that entity. It is a claim about your recorded history, not evidence that the value is malicious or even genuinely new.

open as a page

Before writing a detection rule, what must be true of your telemetry for an adversary behaviour to be detectable?

level: juniorimportance: must knowfreq 72%

basics

~20 s

A record the estate actually writes must carry at least one field whose value differs between the behaviour and the same action done legitimately. With no separating field, no rule can match it, however the rule is written.

open as a page

Why does a detection rule declare the log source and fields it needs, not just the search?

level: juniorimportance: must knowfreq 58%

basics

~20 s

A rule only works if the records it queries exist and carry those fields. Declaring the source and field names makes that assumption checkable before deployment; a rule over missing data returns nothing and looks exactly like a quiet estate.

open as a page

In a SIEM, how does a rule matching a known-bad domain differ from one matching a behaviour?

level: juniorimportance: must knowfreq 72%

basics

~20 s

An indicator rule matches a fixed artefact such as a domain or hash, so it is cheap to write and dies when that artefact changes. A behavioural rule matches how the activity looks: costlier to tune, but it survives.

open as a page

What does an ATT&CK technique ID attached to a detection rule actually claim?

level: juniorimportance: must knowfreq 70%

basics

~20 s

It claims what adversary behaviour the rule was written to catch. The label sits on the rule, not on any alert: a firing is not proof of malice, and the tag does not mean every variant of that behaviour is caught.

open as a page

What is a Sigma rule, and why can't you run one directly in your SIEM?

level: juniorimportance: must knowfreq 72%

basics

~20 s

Sigma is a YAML format that describes a detection - a log source, field matches and a condition - independently of any product. It runs only after a converter and a field-mapping pipeline compile it into SPL, KQL or another backend query.

open as a page

Your DNS-tunnelling detection has raised no alerts in five months - what does that silence prove?

level: juniorimportance: must knowfreq 60%

basics

~20 s

Nothing on its own. Zero alerts has at least three causes: there was no such activity, the rule is broken or muted by a pipeline change, or the data never reached the rule. Silence is a question, not an assurance.

open as a page

Why does allowlisting the host that triggered an EDR alert cost more visibility than narrowing the rule?

level: juniorimportance: must knowfreq 64%

basics

~20 s

Allowlisting a host silences every future firing of that rule there, including a real intrusion. Narrowing the logic removes only the benign pattern, and it removes it everywhere, so the rest of the estate keeps its coverage.

open as a page

Why does a fixed 500-downloads-a-day rule miss a user whose SaaS downloads tripled?

level: middleimportance: must knowfreq 55%

basics

~20 s

A single global threshold has to sit above the loudest legitimate account, so it sits far above a typical one. A user going from six downloads a day to forty has tripled and is still nowhere near the number. Comparing each account to its own history is what makes the change visible.

open as a page

A converted Sigma rule runs in your SIEM without error and matches nothing. Why?

level: middleimportance: must knowfreq 55%

basics

~20 s

Usually the mapping, not the logic: the query was pointed at the wrong table, or a Sigma field was mapped onto a column that exists but carries something else - a bare file name where the rule expects a full path.

open as a page

A two-stage correlation rule never matches because one source ingests four minutes late. What do you change?

level: middleimportance: should knowfreq 46%

basics

~20 s

Add a grace period: hold each window open, or delay evaluating it, until the slow feed has landed, and match on event timestamps. That buys correctness with detection latency and state held longer. Do not widen the correlation window instead.

open as a page

In Splunk SPL, why do detection engineers rewrite a join between two indexes as stats?

level: middleimportance: should knowfreq 58%

basics

~20 s

Because join runs a subsearch that is capped and truncated without an error, so the correlation can silently omit the account you are hunting. Reading both sources in one search and aggregating with stats by the shared key is complete.

open as a page

Which record do you write a Kerberos ticket-harvesting rule over — the domain controller's Windows Security 4769, endpoint process creation, or network flow?

level: middleimportance: should knowfreq 52%

basics

~20 s

The domain controller's Kerberos ticket-issuance records: the only surface every requester in the forest must pass through. The cost is a thin vocabulary — no process, binary or parent at alert time — and an enormous benign base rate.

open as a page

What must be true on a Linux host for an auditd watch on authorized_keys to produce records?

level: middleimportance: should knowfreq 41%

basics

~20 s

The watch must be loaded into the running kernel ruleset, not just present as a file; the audit daemon running and collected; the rule unshadowed by earlier rules; and the fields the detection reads arriving with those names.

open as a page

When should a detection rule carry an ATT&CK sub-technique ID rather than the parent?

level: middleimportance: should knowfreq 66%

basics

~20 s

Label at the granularity the logic can actually separate. A sub-technique ID claims the rule distinguishes that procedure from its siblings. If the rule fires identically on all of them, the parent is the honest label; if it only ever catches one, the parent hides that.

open as a page

How can a parser field rename silently stop a SIEM rule from matching, with no error?

level: middleimportance: should knowfreq 50%

basics

~20 s

Detection backends treat an unknown field as absent, so a predicate on it is simply false, not invalid. The search still runs, still reports success, and returns an empty result set - a legal answer, not a fault.

open as a page

You ship an exclusion to a live detection rule — how do you find out later what it hid?

level: middleimportance: should knowfreq 47%

basics

~20 s

An exclusion suppresses matches silently. Enforce it in the rule's logic rather than at the sensor, so the records still reach the log store, and record its scope and start date so you can search that gap later.

open as a page

First-seen SaaS rules fired 900 alerts overnight after IT migrated every user to a new tenant. What now?

level: seniorimportance: should knowfreq 47%

basics

~20 s

Confirm the flood is the migration by checking that the novel values are all the same expected new tenant on the change window, then suppress narrowly on that value with an expiry, keep every other detection live, and fix the cause with a cold-start learning period before any first-seen rule may fire for an entity.

open as a page

Your purple team Kerberoasted the forest and nothing fired — which observable can a detection actually build on?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Only one thing is forced: a domain controller must issue a service ticket for the targeted service account, recorded as a 4769. The encryption type, the tool, the source host and the request rate are all the adversary's choice, so detections built on them are optional to evade.

open as a page

How do you prove which hosts a new auditd-based rule actually covers across a drifting fleet?

level: seniorimportance: should knowfreq 47%

basics

~20 s

From evidence in received telemetry, not from desired state. Query recent records carrying the watch's tag per host, join that against the asset inventory for the uncovered remainder, and scope the rule's claimed coverage to the proven hosts.

open as a page

Your CDN-fronted C2 rule has been rewritten three times this month — what must its replacement prove?

level: seniorimportance: should knowfreq 55%

basics

~20 s

That its clause still fires when the fixed artefacts are deliberately changed. Re-run the behaviour with a fresh hostname, certificate and TLS stack; replaying the stored record proves nothing, because it carries the artefacts you already know.

open as a page

A Kubernetes rule firing on any pods/exec audit record is tagged T1609, T1611 and T1078 — which labels survive?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Only the container administration one, T1609. The API audit record proves a command channel was opened into a container. It shows no break-out to the node, so T1611 is a hypothesis, and every exec in the estate uses an accepted credential, so T1078 is true of everything and therefore says nothing.

open as a page

A Sigma rule's count aggregation won't convert and an intrusion is live. What do you ship?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Ship the selection without the threshold, scoped narrowly and routed to an investigation queue rather than the pager, and record what was dropped at deploy time. You have traded precision for coverage deliberately, which is defensible; doing it silently is not.

open as a page

An incident lead asks the exact date your DNS rule stopped matching - how do you establish it?

level: seniorimportance: should knowfreq 36%

basics

~20 s

Date the data, not the story: plot the daily count of events populating the field the rule filters on, find the step change, and correlate it with the parser deploy record. Bound the claim to what retention supports.

open as a page

An intruder renames their tool to your excluded profiler's filename and path — what did your exclusion get wrong?

level: seniorimportance: should knowfreq 51%

basics

~20 s

The exclusion keyed on attacker-controllable attributes — any file can take that name and sit in that directory. Scope exceptions on code signature and parent chain, and alert whenever the excluded name appears without them.

open as a page

Your intel analyst wants a whole subscribed indicator feed alerting in the SIEM and you have one analyst — what do you agree to?

level: principalimportance: should knowfreq 44%

basics

~20 s

Take the feed as enrichment on records and alerts by default, and let only a small, high-confidence subset alert, with an expiry date and a named owner who re-confirms it. The constraint is your analyst's week.

open as a page

Your per-user download threshold came from four weeks of data and now fires every quarter-end. Why?

level: middleimportance: nice to knowfreq 32%

basics

~10 s

The baseline window never contained a quarter-end, so the rule encoded four ordinary weeks as the definition of normal. A threshold derived from a period is only valid for periods that look like it.

open as a page

showing 1–30 of 39