skip to content

A fleet owner refuses your auditd watch on 40% of hosts. What do you ship, and what coverage do you claim?

level: principalimportance: nice to knowfreq 29%

answer

  1. attack the volume objection first
  2. narrow by path, never by writer
  3. a snapshot diff loses who and when
  4. inert by design and on the record
  5. the refuser accepts the residual risk

basics

~20 s

Ship the rule scoped to hosts where the prerequisite is proven, state coverage as that host set, and make the remainder a named gap with an owner. Try to buy coverage back by narrowing the watch.

solid answer

~50 s

Do not ship a rule that silently pretends to cover the fleet, and do not hold it hostage until the argument is won - partial detection you can describe beats none. First attack the objection on its merits: a broad syscall rule is expensive, a watch on a few specific key paths is not, and re-scoping often makes the volume complaint go away. Where it truly cannot ship, offer a weaker compensating signal - a periodic inventory of those files, diffed - and be explicit that it gives you changed state without who changed it or when. Then make the residual risk owned: the person who declined the control accepts the gap, not the security team. Record the covered host set as the rule's coverage, put a review trigger on the exception, and let no dashboard roll it up as fleet-wide.

go deeper

for a junior

Know that a detection can be deliberately deployed to only part of an estate, and that saying which part is as important as the rule itself.

for a middle

Be ready to explain what a periodic file snapshot gives you compared with an event-level audit record, and why one is not a drop-in replacement for the other.

for a senior

Show you can take a volume objection seriously, re-scope a watch to answer it with evidence, and refuse the concession that would create a bypass shaped like the attack.

for a principal

Own the coverage claim and the exception. Decide what gets reported upward, put the residual risk with the owner who declined the control, and attach a review trigger so the exception cannot quietly become permanent.

## The situation You have a detection for SSH key persistence on a Linux server fleet - an adversary appending a public key to a service account's `authorized_keys` to hold access. The behaviour is only visible through host audit records, and the older half of the estate has no endpoint sensor at all. The team that owns the fleet declines to load the audit watch on 40 percent of the hosts, citing syscall volume and the operational risk of touching the audit configuration on their busiest machines. They own that configuration and they can say no. This is not a technical problem with a technical answer. It is a decision about what you deploy, what you claim, and who carries what is left. ## First, test whether the objection is about your rule or about audit rules in general Volume objections are frequently correct about a different rule than the one you are proposing. A broad syscall-level rule really can be costly on a busy host. A watch on a small set of specific paths is orders of magnitude cheaper, because it is evaluated only when those objects are touched. Bring numbers from the hosts where it already runs: records per host per day, the actual added load. Offer to narrow the watch further - the service accounts that matter rather than every home directory. Very often the negotiation ends here, and re-scoping recovers most of the 40 percent. Be careful about one tempting concession: excluding the noisy legitimate writer by process, so that config-management writes are never recorded. That looks like a clean volume win and it is a hole shaped exactly like the thing you are detecting, since an adversary who can write the file can usually write it in a way that matches the exclusion. Narrow by path, not by whoever is writing. ## Second, price the alternatives honestly If the watch genuinely cannot ship, there are weaker options, and their weakness has to be stated rather than glossed: - **A periodic file inventory.** Config management or an inventory agent reads the key files on a schedule and reports their contents or a digest; you diff successive snapshots. This catches a key that is still present at snapshot time. It loses the acting process, the acting user and the moment of change, it cannot see a key added and removed between snapshots, and its detection latency is the snapshot interval. - **Detecting the consequence instead of the act.** Authentication records showing an interactive login as a service account that should never log in interactively catch some of the same intrusions, later and by a different route. That is a different rule with its own inputs, not a substitute for this one. Each of these is a real control; none of them is the control you wanted, and the coverage record must not treat them as equivalent. ## Third, ship what is real The rule goes out scoped to the hosts where the prerequisite has been proven from evidence, with the covered set - or the criterion that defines it - written into the rule's metadata. On the remainder it is inert **by design and on the record**. Two things follow from that phrasing. The rule is not a lie, because it never claimed those hosts. And the gap is not invisible, because "inert here, for this reason, since this date" is a sentence somebody can act on. ## Fourth, put the residual risk where it belongs This is the part that is genuinely a leadership call rather than an engineering one. A control declined by the team that would have to operate it is a risk accepted by that team, recorded like any other exception: what is not detected, on which hosts, what the compensating signal is, who accepted it, and what triggers a review - a security event on those hosts, a change to the fleet's baseline image, the next audit cycle. The failure mode to avoid is the security organisation quietly carrying a gap it did not choose, and reporting a coverage figure that includes hosts it knows are blind. ## Why the claim matters as much as the rule Coverage numbers are consumed by people who cannot check them: a risk committee, a customer questionnaire, a regulator's evidence request, the next incident's timeline. "Detection deployed for SSH key persistence" and "detection covering 540 of 900 Linux hosts, with the remainder under a named exception" are the same rule and completely different statements. The second is smaller, harder to say and the only one you can defend when the intrusion lands on one of the 360 - and being able to point at a dated, owned exception is a far better position than explaining why the dashboard said green.

  • Why not exclude the noisy config-management process from the watch to cut volume?
    Because the exclusion is a hole shaped like the attack. An adversary who can write the key file can usually arrange for the write to come from a path or identity that matches the exclusion, and you have removed exactly the events you would need. Narrow by the paths you care about instead, which cuts volume without creating a bypass.
  • The fleet owner asks why a periodic snapshot diff is not good enough. What do you say?
    That it detects a different thing. A snapshot shows the file changed between two points and gives you neither the process nor the identity nor the time of the change, and it misses a key added and removed inside one interval. It is a reasonable compensating control for hosts that cannot take the watch, and it should be recorded as weaker rather than as equivalent.
  • Who signs off on the gap, and what does that sign-off actually have to contain?
    The owner who declined the control, not the security team. It should state the behaviour not detected, the host set, the compensating signal and its limits, the date, and a review trigger such as a change to the fleet baseline or a confirmed incident on those hosts. Without a trigger it becomes permanent by default, which is how most exceptions die of old age.

saying these in an interview costs you the question

  • Ships the rule fleet-wide and lets the metadata imply full coverage
  • Refuses to deploy anything until the whole fleet complies
  • Excludes the legitimate writer by process to cut volume
  • Presents a periodic snapshot diff as equivalent coverage
  • Leaves the security team silently owning a gap it did not choose

context