skip to content

How do you prove which hosts a new auditd-based rule actually covers across a drifting fleet?

level: seniorimportance: should knowfreq 47%

answer

  1. desired state is a self-report
  2. look for the tag in received records
  3. benign writes are your proof of life
  4. join from the inventory, not the telemetry
  5. coverage is a host set that decays

basics

~20 s

From evidence in received telemetry, not from desired state. Query recent records carrying the watch's tag per host, join that against the asset inventory for the uncovered remainder, and scope the rule's claimed coverage to the proven hosts.

solid answer

~50 s

Config management tells you what it intended to do, and on a drifting estate that is not what the kernel loaded. The check that settles it is evidence in the data: for each Linux host in the inventory, has any record carrying the watch's tag arrived in the last few days? Benign writes to those paths - config management rotating a deploy key, an engineer adding their own - are frequent enough to serve as proof of life, and their absence over a long enough window says the watch is not loaded. Join from the asset inventory, not from the hosts already sending logs, or you only ever see what you have. Then deploy with coverage stated as that host set, re-run the check on a schedule, and treat the remainder as a named gap rather than rounding it up to the fleet.

go deeper

for a junior

Understand that a rule can be live in the search platform while doing nothing on most hosts, and that the way to find out is to look for records those hosts should already be producing.

for a middle

Be able to describe the evidence query: which hosts have produced records carrying the watch's tag recently, and why the asset inventory rather than the telemetry has to define the denominator.

for a senior

Demonstrate judgment about the evidence window against the base rate of benign writes, about grouping silent hosts to tell a rollout gap from a measurement artefact, and about re-checking as the estate drifts.

for a principal

Own what coverage means as a reported figure. Decide whether the organisation reports detections deployed or hosts demonstrably covered, and be ready to defend the smaller, truer number.

## The question is not "is the rule deployed" but "where can it fire" A detection deployed to the search platform is deployed once. The prerequisite it depends on is deployed hundreds of times, by a different team, through config management, onto hosts that drift, that were built from older images, that were excluded from a rollout for a reason nobody remembers. On an estate with no endpoint sensor on the older servers, that audit configuration is the only path to the behaviour, so the rule's real coverage is exactly the set of hosts where the watch is live. ## Desired state is not evidence The tempting answer is to ask config management. It will happily report the audit rule as applied on 900 hosts. That is a statement about a file it wrote, made by a system that measures its own success. It does not know whether the ruleset was rebuilt, whether an earlier suppression or immutability directive prevented the load, whether the audit daemon is running, or whether the forwarder is shipping. Every one of those failures leaves the file exactly where config management put it. The same trap applies to the asset inventory alone, which tells you a host exists, and to "the host is sending logs", which tells you *a* source is alive on it - very often a different one from the one your rule needs. ## Evidence in the received data The check that means something is: **which hosts have produced any record carrying this watch's tag recently?** A file watch tagged with a key stamps that tag onto every matching record, including entirely legitimate ones. Writes to SSH key files are not rare in a managed fleet - key rotation, onboarding, a deployment tool refreshing an automation account's access - and those benign writes are the proof you need. You are not looking for the adversary; you are looking for the watch to have spoken at all. Two details make this honest rather than comforting: - **Join against the inventory, not against the telemetry.** If you list hosts *in the data* you can only ever enumerate hosts you already receive from. The uncovered set is exactly the hosts that appear in the asset inventory and are absent from the evidence, so the inventory has to be the left-hand side of the join. - **Choose the window against the real base rate.** If legitimate writes to those paths happen weekly, a seven-day window will misclassify covered hosts as uncovered. Where the base rate is too low or too uneven, the alternative is to have config management or an inventory agent report the *loaded* ruleset as a periodic fact and consume that as a second, weaker line of evidence - weaker because it is a self-report from the host rather than a record that travelled the same path your alerts will travel. That last point is the one people miss. The value of tag-carrying records as proof is that they exercise the entire pipeline the detection depends on: kernel, daemon, forwarder, parser, index. An inventory fact proves only the first link. ## Then scope, and say so Once you have the covered set, two things follow. First, the rule ships scoped to it, or at minimum its coverage metadata names it - a rule silently spanning the fleet while only functioning on part of it is the failure this whole exercise exists to prevent. Second, the uncovered complement becomes a named, owned gap with a route to closing it, not a rounding error. It is entirely legitimate for a detection to be inert on part of an estate **by design and on the record**; it is not legitimate for that to be a surprise. Re-run the check on a schedule, because coverage is not a property you establish once. Hosts get rebuilt from old images, an audit configuration gets reverted during an unrelated hardening change, a new region comes online without the rollout. Coverage decays quietly and in the same direction every time. ## Where this stops This is a deployment-time and periodic-coverage question: establishing where a rule can fire and stating it. Working out why a rule that *was* firing has stopped is a different investigation with a different starting point, and it belongs to the people who own rule health monitoring. What you owe here is that the answer to "where does this rule work?" is a host set derived from evidence, and that nobody downstream is told the fleet is covered when 40 percent of it never loaded the watch.

  • Why is querying received records better evidence than a host reporting its loaded audit ruleset?
    Because a received record has travelled the whole path the alert will travel: kernel, daemon, forwarder, parser, index. A host self-reporting its ruleset proves only the first link, and every link after it can be broken independently. The self-report is useful as a second, weaker signal, especially where legitimate writes are too rare to serve as proof of life.
  • You find 40 percent of the fleet has produced no record with the watch's tag in 30 days. What is the first thing you check?
    Whether those hosts have a plausible reason to be silent before assuming they are broken. Group them by image, build date, region and config-management group: a clean split along one of those lines points at a rollout that never reached them, while a scatter across all groups points at your evidence window being too short for the base rate of legitimate writes.
  • How should the rule express the fact that it only covers part of the fleet?
    As explicit metadata on the rule: the host set or the selection criterion it is valid for, and the prerequisite that defines it. Scope the query to that set where the platform allows it, so the rule is inert elsewhere by design. Then the uncovered remainder can be tracked as a gap with an owner instead of being silently absorbed into a coverage claim.

saying these in an interview costs you the question

  • Treats config-management desired state as coverage
  • Counts hosts that send any logs as covered
  • Enumerates hosts from telemetry, so misses the silent ones
  • Establishes coverage once and never re-checks it
  • Reports fleet-wide coverage for a partially loaded prerequisite

context