skip to content

Why does a detection rule declare the log source and fields it needs, not just the search?

level: juniorimportance: must knowfreq 58%

answer

  1. a query plus a claim about the world
  2. zero hits has more than one cause
  3. silence is what success looks like too
  4. the assumption should be falsifiable
  5. source, prerequisite, field names, scope

basics

~20 s

A rule only works if the records it queries exist and carry those fields. Declaring the source and field names makes that assumption checkable before deployment; a rule over missing data returns nothing and looks exactly like a quiet estate.

solid answer

~50 s

A rule is a query plus a claim about the world: that some source is shipping records of a particular kind, and that those records carry the fields the logic reads. If either half is false the rule returns zero rows forever, and zero rows is indistinguishable from an estate where the behaviour never happened. Declaring the inputs turns that silent assumption into something a human or a script can verify at deploy time. Concretely, a rule looking for an SSH public key appended to a service account's `authorized_keys` on Linux servers should say which source it expects (audit records from a file watch on the `.ssh` paths), which fields it reads (`file.path`, `process.executable`, `user.effective.id` in Elastic Common Schema terms), and what must already be configured on the host for those records to exist at all. Without that, you have deployed a claim you cannot check.

go deeper

for a junior

Be ready to say what a rule assumes beyond its logic: a source that is shipping, fields that are populated, and something enabled on the host. Know that zero hits does not mean nothing happened.

for a middle

An interviewer expects you to explain how a field can be defined by a schema and still be empty for a given source, and to name the deploy-time checks that separate a broken rule from a quiet estate.

for a senior

Show that you write the input contract into the rule so it can be checked mechanically across an estate, and that you refuse to claim coverage for hosts where the prerequisite was never proven.

for a principal

Own the standard: every rule in the library declares its source, prerequisite and fields, and coverage is reported as a host set. Be able to argue why that discipline is worth the authoring friction it adds.

## The rule is not just the logic Most people first meet a detection rule as its search: a condition over some records. That is only half of it. The other half is a set of assumptions about the world, and those assumptions are what actually decide whether the rule ever produces an alert: 1. **A source exists and is shipping.** Some agent, forwarder or platform is emitting records of the right kind, from the hosts you care about, into the store the rule runs over. 2. **The records carry the fields.** The logic reads named fields. Those fields must be present and populated for *this* source, not merely defined in the schema. 3. **A host-side prerequisite is satisfied.** Many sources emit nothing about a behaviour unless something was deliberately enabled: an audit rule loaded into the kernel, a configuration section that turns on a particular event class, an audit policy subcategory switched on. The telemetry does not exist by default. A rule that names all three is a *contract*. A rule that names none of them is a guess that looks like a control. ## Why the failure mode is so dangerous Detection has an asymmetry that most engineering does not. When a web service is broken it returns errors; when a detection is broken it returns **silence**, and silence is exactly what a healthy detection produces most of the time. There is no error to page on. A rule pointed at a field nobody populates, or at hosts where the required audit configuration was never applied, will sit in the console at zero hits looking precisely like a rule doing its job in a clean estate. So the direction of inference matters: **a rule producing no alerts tells you nothing about the estate.** It could mean the behaviour did not occur, or that the records never arrived, or that they arrived with different field names, or that the logic never matched. Declaring the inputs is what lets you rule out the boring explanations before you start believing the interesting one. ## What a declaration looks like in practice Take persistence by writing an attacker-controlled SSH public key into a service account's `authorized_keys` on a Linux server fleet. There is no EDR on the older half of that fleet, so the observation surface is the Linux audit subsystem. A rule for this behaviour should state: - **Source:** Linux audit records from hosts in the server fleet, shipped by the log agent. - **Prerequisite:** a file watch covering the `.ssh` directories, loaded into the running kernel audit ruleset and tagged with a known key, plus the audit daemon running and its output being collected. - **Fields read:** the file path written to, the executing process image, and the acting user identity - in ECS terms `file.path`, `process.executable` and `user.effective.id`. - **Scope:** which hosts this is asserted to be true for. Each line is falsifiable. You can query the last few days of received records for that watch key per host and see which hosts have ever produced one. You can check whether the field the logic reads is populated on real records rather than merely legal in the schema. Neither check requires an adversary to cooperate. ## Schema presence is not population A common junior error is to treat a shared schema - Elastic Common Schema, OCSF - as a guarantee. It is not. A schema defines what a field is *called* and what it *means*; it says nothing about whether any particular source fills it in. Cloud audit records, endpoint process records and file-integrity records all live in the same schema and populate wildly different subsets of it. Writing `process.executable` into your logic is a bet that the source you picked actually carries the executing process, and plenty of file-oriented sources do not. ## What this buys you Declaring inputs is what makes three later things possible: deploying a rule only where its prerequisite is proven, stating detection coverage as a host set rather than a hope, and reviewing the rule when the platform changes underneath it. It is also the cheapest honesty available in detection engineering - it costs a few lines of metadata and it removes the single most common way a detection is wrong, which is that it was never able to fire in the first place.

  • A rule references the ECS field file.path and the field exists in your schema. What have you actually verified?
    Only that the field is legal. A schema defines a field's name and meaning; it does not say which sources populate it. The check that matters is looking at real received records from the intended source and confirming the field is present and non-empty on them. Endpoint, cloud and file-integrity sources share a schema while filling in very different subsets of it.
  • Your rule was tested successfully in a lab host and merged. Why is that not evidence it works in production?
    The lab host had the prerequisite enabled because you enabled it by hand. Production hosts get their configuration from config management, drift from it, and in an older estate may never have received it. The lab proves the logic matches the behaviour; it proves nothing about whether the records the logic needs are being produced anywhere else.
  • Where should the declared inputs live so they are useful?
    With the rule itself, in a form something can read: an explicit source and required-field list in the rule's metadata rather than a sentence in a wiki or a comment in a ticket. If it sits next to the logic it gets reviewed with the logic, and a script can walk every rule and ask which prerequisites the estate currently satisfies.

It is like a smoke alarm wired to a circuit nobody checked. It never sounds, which is exactly what you would expect in a house that never catches fire.

saying these in an interview costs you the question

  • Says no alerts in a month means the estate is clean
  • Assumes a field in the schema is populated by every source
  • Treats the search logic as the whole rule
  • Believes a rule deployed in the SIEM is deployed on the hosts
  • Verifies a rule only in a hand-configured lab host

context