skip to content

A Sigma rule's count aggregation won't convert and an intrusion is live. What do you ship?

level: seniorimportance: should knowfreq 40%

answer

  1. search converts, grouping often does not
  2. clean conversion is not complete conversion
  3. replace the threshold with scope
  4. route away from the pager
  5. write the loss down at deploy time

basics

~20 s

Ship the selection without the threshold, scoped narrowly and routed to an investigation queue rather than the pager, and record what was dropped at deploy time. You have traded precision for coverage deliberately, which is defensible; doing it silently is not.

solid answer

~50 s

First be precise about what was refused. Everything before the pipe - the field matches - converts almost everywhere; the aggregation after it often does not, because many backends implement search only and have no grouping construct wired up. That leaves two options. Reimplement the grouping natively, `stats count by ... | where count > 5` in SPL or a `summarize` in KQL, which is right when someone fluent is on the bridge and you have the minutes. Or degrade on purpose: ship the selection alone, narrow it by scope rather than by threshold - one business unit, a named host set, a bounded window - route hits to a triage queue with an owner watching instead of to the pager, and write down at deploy time exactly which clause went, what volume you expect and when it gets revisited. The count clause was the noise control, so its loss lands on an analyst queue that is already working an intrusion. Scoping is what keeps that survivable.

code

yaml · 15 lines
yaml
title: Mshta Launching A Remote Scriptlet
logsource:
  category: process_creation
  product: windows
detection:
  selection:
    Image|endswith: '\mshta.exe'
    CommandLine|contains:
      - 'http://'
      - 'https://'
  condition: selection | count() by User > 5
timeframe: 10m
level: high
tags:
  - attack.t1218.005

go deeper

for a junior

Know that a Sigma rule can contain logic the target platform cannot express, and that the converter's success message does not mean the whole rule shipped.

for a middle

Explain which constructs travel - field matches and booleans - and which do not: aggregations, correlations, temporal proximity and regex dialects.

for a senior

Show the judgment: degrade deliberately with scope and routing as the replacement noise control, and record the loss at deploy time rather than after the incident.

for a principal

Be ready to defend a policy on degraded detections - who may ship one during an incident, what must be recorded, and how they get revisited before they become invisible permanent furniture.

### What survives conversion and what does not The useful mental split is *search* versus *everything else*. Field matches, string modifiers, boolean combinations and wildcards convert to essentially every backend. What travels badly: - **Aggregations and thresholds.** The legacy `condition: selection | count() by User > 5` form, and the newer correlation rule types that count events, count distinct values or require temporal proximity, depend on backend support that is frequently absent. The converter either refuses outright or you get the search half only. - **Temporal proximity.** Anything meaning "these two things within N minutes" needs a correlation construct in the target, and a plain search backend has none. - **Regular expressions.** The `re` modifier assumes a dialect. Character classes, anchors and lazy quantifiers do not behave identically across engines, and some backends decline regex entirely. - **CIDR and case semantics.** Sigma defines case-insensitive matching; whether the operator the backend picks is case-sensitive varies, as does native CIDR support. - **Field-less keyword searches.** A `keywords` block needs full-text search over the raw event, which some platforms do not offer, or offer at a cost you will notice. - **The metadata.** `level`, `tags` and `falsepositives` are not part of the query at all. If your deployment tooling does not carry them, they are lost regardless of backend. An interviewer is checking whether you know that a *clean* conversion and a *complete* conversion are different claims. ### The decision under time pressure An incident lead with a live intruder wants a published rule running estate-wide in twenty minutes. The engineer knows the count clause will not survive. Reimplement or degrade. **Reimplement** when the aggregation is the detection - when the raw selection is so common that ungrouped it is meaningless - and someone can write and sanity-check the native grouping inside the window you have. You get fidelity, at the cost of an artefact your team now maintains by hand. **Degrade** otherwise, and degrade *deliberately*. Three controls make it defensible: 1. **Scope instead of threshold.** You have lost the noise control the author built in, so replace it with a different one. Restrict to the affected business unit, the host set in scope, or a time window - narrower blast radius, same behaviour covered. 2. **Route away from the pager.** Hits go to a triage queue with a named owner watching it, not to whoever is on call. During an incident the queue you would flood is the same queue working the intrusion. 3. **Write the loss down at deploy time, not later.** Which clause was dropped, why, the expected volume, the queue, the owner, and a date to revisit. This note is what stops a degraded rule becoming permanent furniture that everyone assumes still has a threshold on it. The outcome is not hypothetical: a degraded rule is still a rule. Shipped ungrouped and scoped, it can fire on a second host and that firing is what turns a suspicion into a confirmed second compromised system - which is worth far more than a perfectly faithful rule shipped after the intruder finished. ### The cost you take on when you reimplement A native grouping written by hand is no longer generated from the portable source. The next version the author publishes will not carry your grouping, and re-running the conversion will not reproduce it. That is often the right trade, but it has to be visible at the moment you make it rather than discovered a year later by someone wondering why the deployed query and the rule they are reading do not agree. ### The habit to demonstrate Read the converter's output, not just its exit status. Know before you deploy whether the query you are about to ship is the whole rule or the part of it that happened to be expressible. Everything else in this answer follows from that one habit.

  • Why not just alert on every match and let the analysts sort it out?
    That is what you are doing - but knowingly, with the blast radius controlled, is the difference. An unscoped selection on a signed script host runs across every workstation in the estate, and the queue it floods is the same one working the intrusion. Scope it, route it off the pager, and put a name against it.
  • The count clause was the only thing keeping the rule quiet. What exactly do you record?
    At deploy time: which clause was dropped and why, the volume you expect, the queue it routes to, the owner watching it, and a date to revisit. That note is the difference between a deliberate degradation and a rule everyone later assumes still carries a threshold it lost months ago.
  • You reimplement the aggregation natively instead. What did you just lose?
    The link back to the portable source. The native query becomes an artefact your team maintains by hand; the next published version of the rule will not carry your grouping, and re-converting will not reproduce it. Acceptable, but it has to be visible at the moment you decide rather than discovered later.
  • Which parts of a rule are most likely to convert cleanly but behave differently?
    Regular expressions, case sensitivity and wildcard escaping. Sigma defines matching as case-insensitive, but the operator a backend selects may not be, and regex dialects differ on anchors and quantifiers. Those produce a valid query with subtly different semantics, which is harder to notice than an outright refusal.

saying these in an interview costs you the question

  • Assumes a clean conversion means the whole rule converted
  • Ships the unscoped selection straight to the on-call pager
  • Blocks on perfect fidelity while the intruder is still active
  • Thinks Sigma aggregations run on every backend
  • Leaves the dropped clause undocumented after the incident

context