skip to content

Your emulation miss traces to absent auditd rules on one host image — who owns the finding, and how do you make that stick?

level: principalimportance: nice to knowfreq 28%

answer

  1. the label picks the owner
  2. a rule over absent data cannot fire
  3. never fired looks like a quiet estate
  4. scope it to the build class
  5. acceptance is a re-execution, not a config

basics

~20 s

The image and platform owners, not the detection team. Make it stick by scoping the finding to the host class with control-host evidence, naming an owner who can change the image, and setting the acceptance test as re-executing the technique.

solid answer

~50 s

Classify it as a **collection gap on a named host class**, because the classification decides the owner and the fix. A detection gap goes to the detection engineers and closes cheaply; a collection gap goes to whoever builds the image, is slow, and carries an audit-volume cost — so it is the label people quietly want to change. Resist that: a rule over records that are never produced cannot fire, and a rule that has never fired is indistinguishable from one watching a quiet estate, so the reclassification manufactures the appearance of coverage. Make it stick with the control-host artefact, a scope statement counting how much of the fleet shares that build, and an acceptance test that re-executes the technique on a rebuilt host rather than inspecting configuration. If the platform team declines on cost, write it as a scoped accepted blind spot, not a closed finding.

go deeper

for a junior

Know that a detection rule cannot fire on records that are never produced, so a missing log source is not something the detection team can fix for you.

for a middle

Be able to explain why an unfireable rule is worse than an open gap: a rule that never fires looks the same as a rule watching a quiet estate.

for a senior

Demonstrate the evidence package — control host, scoped host class, named owner — and an acceptance test that re-executes the technique rather than inspecting configuration.

for a principal

Own the classification discipline across the programme, including the decision to convert a refused finding into a scoped accepted risk instead of escalating until you win.

## Why the label decides everything By the time you have walked the chain back, you hold a fact: on hosts built from a particular image, the behaviour you executed produced no audit record at all, while a control host of a different build produced it immediately. You now have to write that down, and the two available labels are not equivalent. - **Detection gap** — the estate saw the behaviour and had no rule for it. Owner: detection engineering. Cost: a rule, days, no infrastructure change. Closes fast and looks like progress on the quarterly slide. - **Collection gap** — the estate did not see the behaviour. Owner: whoever owns the host image and its audit configuration. Cost: an image change, a fleet rebuild cycle, a conversation about audit volume and its performance and storage bill. Closes slowly and looks like an unresolved risk for a quarter. The organisational gravity is entirely towards the first label, and it will be applied by well-meaning people who are trying to be helpful. A senior purple-team lead's real job here is to make the correct label survive contact with that gravity. ## The argument, stated so it cannot be waved away A rule over telemetry that is never produced **cannot fire**. And here is the property that makes it dangerous rather than merely useless: a detection produces no output when it is working perfectly in a quiet estate, and no output when it is dead. Those two states are indistinguishable from the outside, so an unfireable rule is not a null; it is a **positive false belief in coverage**, sitting on the dashboard as a closed finding, and it will be counted as coverage in the next assurance conversation. That is strictly worse than the open gap, because an open gap is at least visible. This is also why the ordinary way of validating rules does not save you. Replaying a captured input through a rule tests the rule's logic. It cannot tell you whether the estate would have produced that input in the first place — only executing the behaviour on a real host of that class can, and that is exactly what the exercise did. ## Making it stick **Evidence.** The finding carries the control-host record and the target's silence over the same window, so the negative is defensible rather than an assertion. Without that, the finding is one person's claim and will be renegotiated. **Scope.** You tested one or two hosts; the finding is about a class. Say how many machines share that build lineage and how you counted, and say plainly what you have not tested. An unscoped finding is either dismissed as an anomaly or inflated into a fleet-wide crisis, and both outcomes lose you credibility. **A real owner.** The finding must name a team with the authority to change the image, not a distribution list. If no such owner exists, that absence is itself the finding to escalate. **An acceptance test that is an execution.** "The audit rule is present in the build definition" is not acceptance; configuration present in a repository is not the same as records arriving from a running host. Acceptance is: rebuild a host of that class, execute the technique again, observe the record, then observe the rule, then observe the alert. Retest the whole chain, because fixing the first link tells you nothing about the three that were never exercised. ## When the answer is no Sometimes the platform team is right to refuse. Broad syscall auditing on a large fleet has a genuine volume and performance cost, and the honest response is not to escalate until you win. It is to convert the finding into an explicitly accepted risk: this technique, on this host class, is currently unobservable; the class contains this many machines running these workloads; here is a narrower rule that would cover the specific behaviour at a fraction of the volume, or a compensating observation on a different surface — the journal record of the unit starting, a file-integrity watch on the unit directory — that gives partial coverage. Then it goes on the risk register with a named accepting owner, and it comes back to the next exercise as a known blind spot rather than an open ticket. The general principle to state in an interview: **an emulation programme's output is not rules, it is correctly attributed and correctly owned findings.** Its credibility depends on every finding routing to the link that actually broke, and on never letting a cheap fix at one link be used to close a problem at another. The moment collection gaps start being closed with detection content, the programme's numbers become fiction, and everyone downstream — including the people who decide whether the SOC gets the audit budget — is reading a report that overstates what the estate can see.

  • What is the acceptance test for a finding attributed to absent telemetry on a host image?
    Rebuild a host of that class, re-execute the technique, and walk the whole chain again: record present, rule matches, alert queued. Configuration appearing in a build repository is not acceptance, because a setting committed is not a record arriving. And retesting only the first link leaves the three that were never exercised still unproven.
  • The platform team refuses on audit-volume grounds. What do you write?
    An accepted risk, not a closed finding: this technique on this host class is currently unobservable, the class is this many hosts, and here is the narrower audit rule or the compensating surface — the journal record of the unit starting, a watch on the unit directory — that would give partial coverage at lower cost. Name the accepting owner and carry it into the next exercise as a known blind spot.
  • Why is a rule written over uncollected data worse than leaving the gap open?
    Because an open gap is visible and an unfireable rule is not. A detection that never fires looks identical whether the estate is quiet or the rule is dead, so it will be counted as coverage on every assurance review while covering nothing. You have converted a known risk into a hidden one and taken credit for it.
  • How do you scope a finding from two tested hosts to a host class without over-claiming?
    State the property you are generalising over — the build lineage or image version whose audit configuration you inspected — count how many hosts share it, and say explicitly which classes you did not test. If you cannot enumerate the class, the finding is about the hosts you tested plus an open question, and saying so is what keeps the rest of the report credible.

saying these in an interview costs you the question

  • Accepts a new rule as the fix for uncollected telemetry
  • Reports a collection gap with no control-host evidence
  • Generalises from one host to the fleet without scoping
  • Treats a build-repo config change as the acceptance test
  • Escalates instead of writing a scoped accepted risk

context