skip to content

What is a Sigma rule, and why can't you run one directly in your SIEM?

level: juniorimportance: must knowfreq 72%

answer

  1. a portable format, not an engine
  2. logsource is abstract, not an index
  3. compiled by a backend plus a pipeline
  4. metadata is not part of the search

basics

~20 s

Sigma is a YAML format that describes a detection - a log source, field matches and a condition - independently of any product. It runs only after a converter and a field-mapping pipeline compile it into SPL, KQL or another backend query.

solid answer

~50 s

A Sigma rule is a YAML document with a `logsource` block (an abstract pointer such as `product: windows`, `category: process_creation`), a `detection` block holding one or more named selections plus a `condition`, and metadata such as `level`, `tags` and `falsepositives`. It is a source format, not an executable query - there is no Sigma engine anywhere. A converter, today `sigma-cli` driving a pySigma backend, compiles it into SPL, KQL or Lucene, and a *processing pipeline* supplies the environment-specific half: which index or table that logsource actually lives in, and which of your columns each Sigma field name maps to. That split is the point - the author publishes behaviour, you supply the schema - but it also means a rule that is correct as published can be wrong in your tenant, and the metadata never rides along into the alert unless your pipeline puts it there.

go deeper

for a junior

Be ready to name the parts of a Sigma rule - logsource, detection, condition, metadata - and to say plainly that it is compiled into a platform query rather than executed as one.

for a middle

Explain the two halves of conversion: the backend produces the query dialect, the processing pipeline supplies the index and field mapping. Interviewers probe whether you know which half owns what.

for a senior

Expect to discuss what a published rule guarantees once it lands in your tenant, and how you verify that rather than trusting a clean conversion exit code.

for a principal

Own the call on whether the estate authors detections portably at all, and be able to price the abstraction tax against the value of importing other people's rules.

### What the format is Sigma exists to solve one problem: a detection idea is portable, but every query language that could express it is not. One team writes `mshta.exe` with an `http` URL on its command line as SPL, another as KQL, a third as a Lucene query, and none of them can share the result. Sigma is a YAML description of the detection itself, from which each of those queries can be generated. The document has a small, fixed shape. `title` and an optional `id` name it. `status` says whether it is experimental or stable. `logsource` declares the kind of data it needs - a `category` (for example `process_creation`, `network_connection`, `dns_query`), a `product` (`windows`, `linux`, `aws`), and sometimes a `service` (`security`, `sysmon`). `detection` contains one or more named search identifiers - each a map of field names to values, with modifiers such as `contains`, `startswith`, `endswith` or `re` - and a `condition` that combines them with `and`, `or`, `not` and `all of`. Around that sit `falsepositives` (free-text notes on what benign work will trip it), `level` (the author's severity), and `tags` (commonly ATT&CK technique identifiers such as `attack.t1218.005`). ### Why it is not a query language A query language has an engine that executes it. Sigma has none. It is compiled, and the compilation has two halves that people routinely conflate. The **backend** produces the dialect: it knows how to write a wildcard, a string comparison, a list membership and a boolean combination in SPL or KQL. The **processing pipeline** produces the environment binding: it resolves `product: windows` + `category: process_creation` to the actual table or index where you keep that data, and it renames Sigma's taxonomy fields (`Image`, `CommandLine`, `ParentImage`, `User`) to whatever your platform calls them. Run a rule through a backend with no suitable pipeline and you get, at best, a refusal - and at worst a syntactically valid query naming columns your schema does not have. Three consequences follow directly, and they are what interviewers are actually testing. First, **the same rule is two different queries on two platforms, and they are not guaranteed to be semantically identical**. Sigma's matching is defined as case-insensitive; backends differ on whether their operators are. Wildcard and escaping rules differ. A rule that behaves one way in one tenant can behave slightly differently in another without anybody editing it. Second, **what the backend cannot express is refused or dropped**. Search logic converts almost everywhere; anything that groups, counts or correlates across events does not, because many backends implement search only. Third, **the metadata is not part of the search**. `level`, `tags` and `falsepositives` are properties of the rule document, not of the compiled query. Unless your deployment tooling copies them onto the alert definition, the analyst who receives the hit at 03:00 gets a bare search result with no severity, no technique context and none of the author's notes about what benign activity looks like. That is a real triage cost, paid on every firing. ### What it is not Sigma is not YARA. YARA matches content - strings and byte patterns inside a file or a memory image - and is evaluated by a scanner against that content. Sigma matches fields in log records and is compiled into a search that a log platform runs. They answer different questions: *is this artefact the malware* versus *did this behaviour appear in my telemetry*. ### What portability actually buys It buys consumption. When a vendor or a community project publishes a rule for a signed Microsoft script host being used to fetch a remote scriptlet, you can take that rule the day it appears instead of reimplementing it from a blog post. It buys an exit if you ever change SIEM. What it does not buy is correctness in your data: a published Sigma rule is a claim about adversary behaviour, not a claim about your columns, your agents or your coverage. Turning the first into the second is the work.

  • What does the logsource block actually name?
    A category, a product and sometimes a service - for example `product: windows`, `category: process_creation`. Those are taxonomy labels, not index names. The processing pipeline decides that this means one tenant's `DeviceProcessEvents` table and another's Sysmon index. Without that mapping the converter has nothing to point the generated query at.
  • How is Sigma different from YARA?
    YARA matches content - strings and byte patterns in a file or memory image - and a scanner evaluates it against that content. Sigma matches fields in log records and is compiled into a search a log platform runs. One asks whether an artefact is the malware; the other asks whether a behaviour appeared in your telemetry.
  • The rule says level high and carries ATT&CK tags, but the alert shows neither. What went wrong?
    Nothing in the search. That metadata is not part of the compiled query - it lives in the rule document. Your pipeline or deployment tooling has to copy it onto the alert definition. Skip that and analysts get a bare search hit with no severity and no technique context, which you pay for on every single firing.

Sigma is source code and the backend is the compiler; the processing pipeline is the linker that binds abstract field names to the symbols your SIEM actually exports.

saying these in an interview costs you the question

  • Calls Sigma a query language with its own engine
  • Thinks conversion is a plain syntax translation
  • Assumes logsource names an index in your SIEM
  • Expects level and tags to reach the alert automatically
  • Confuses Sigma with YARA content matching

context