The Linux platform owner refuses your auditd collection requirement on cost — now what?
answer
- the refusal is a budget decision, not obstruction
- shrink scope before escalating
- measure the overhead, do not assert it
- one hunt is a weak business case
- accepted gap, named owner, review date
basics
~20 sShrink the ask to targeted rules on a defined host set, settle the overhead argument with a measured pilot rather than assertion, and offer cheaper alternatives. If it is still refused, record an accepted gap with a named owner and expiry, and reflect it in hunt reporting.
solid answer
~60 sThe refusal is usually rational: continuous syscall auditing costs CPU on every host and adds ingest volume the platform team pays for, and "a hunter wanted it" is a weak reason. So I renegotiate the shape rather than repeat the ask. Narrow it — execution records plus two directory watches, on the crown-jewel host set, not the whole fleet. Replace the argument about overhead with a measurement: a time-boxed pilot on a canary group with the owner's own performance data as the acceptance criterion. Offer cheaper substitutes where they answer a weaker version of the question, such as a periodic inventory of unit and cron files through the existing configuration-management agent. Broaden the justification beyond my hunt to the detections that are also blind. And if it is still no, that is a legitimate outcome — I write it up as an accepted collection gap with the accepting owner named and a review date, and I make sure our reporting says that portion of the fleet is unhunted for this technique rather than clean.
go deeper
Know that telemetry has an owner and a cost, and that a collection request has to be specific enough for that owner to price rather than a general plea for more logging.
Be able to explain what to trade away first — fleet-wide scope, breadth of rules, retention — and why a measured pilot beats an argument about overhead.
Show that you can build the case beyond your own hunt to blocked detections and post-incident investigability, and that you offer honest cheaper substitutes with their limits stated.
Own the outcome when the answer stays no: an accepted gap with a named accepting owner and an expiry, plus reporting that stops crediting uninstrumented hosts as clean.
## Why the refusal deserves respect A collection requirement is a bill sent to somebody else. Continuous syscall auditing on a Linux fleet has a real cost: measurable CPU on every host, disk on hosts that buffer locally, and ingest and storage volume that the platform or observability budget carries. The owner refusing is not obstructing security; they are protecting a service budget against a request whose benefit was stated as "a hunt would have been possible". If the ask cannot be justified in their terms, it deserves to lose. That framing is what changes the play. The question is not how to escalate, it is how to make the smallest ask that unblocks the most. ## Shrink the ask Most refusals are refusals of the *fleet-wide, everything-on* version. Almost every hunt needs less: - **narrower selection**: execution records, plus write coverage on the unit and cron directories — not a broad hardening ruleset imported wholesale; - **narrower host set**: the hosts that matter, defined by what runs on them rather than by convenience — a crown-jewel subset of a large fleet is often a small fraction of the cost; - **narrower duration**: a defined retention for those records, agreed rather than assumed. A request that has visibly been cut down reads differently from one that has not, and it gives the owner something to say yes to. ## Settle the overhead with a measurement The overhead argument is usually conducted between two people who have not measured it. End that. Propose a time-boxed pilot on a canary group of the owner's choosing, with the acceptance criterion written in advance and taken from **their** performance data, not yours. Two outcomes, both useful: the overhead is acceptable and the objection dissolves, or it is not and you have learned something true about that workload instead of arguing. Offering the owner control of the canary set and the criterion is also what converts the conversation from a demand into a shared experiment. ## Offer the cheaper substitute honestly Sometimes a weaker source answers most of the question at a fraction of the cost. A periodic inventory of unit files and cron entries collected through a configuration-management agent that is already deployed shows the *installation* of scheduled persistence, though not what it executed or when. Saying so — including saying what it does not cover — buys credibility for the times you insist the stronger source is genuinely required. Do not present a substitute as equivalent when it is not. ## Broaden the justification beyond your hunt One hunter's blocked hypothesis is a weak business case. The same missing records usually block a set of standing detections and any future investigation on those hosts, which is a different sentence: this portion of the fleet cannot be investigated after an incident, not merely hunted. If an audit or contractual obligation touches those systems, that belongs in the ask too. Bring the count of hosts and what runs on them, so the decision is about business exposure rather than about a security team's preference. ## If the answer is still no A refusal after a good-faith negotiation is a legitimate organisational outcome, and the security failure mode is to absorb it quietly. Instead: - **write it down as an accepted gap**, in whatever risk register the organisation actually uses, with the record type, the host set, the hunts and detections it blocks, and the consequence — that this segment cannot be searched for this technique class; - **name the accepting owner**, a person with the authority to accept it, not the security team; - **set a review date**, so the acceptance expires rather than becoming permanent by silence; - **change your own reporting**, so that hunts and detection-coverage summaries mark those hosts unhunted or uncovered rather than counting them as clean. This is the part with lasting effect: it removes the false assurance the organisation was getting for free. Then let the next event do the arguing. When an incident touches an uninstrumented host and the investigation stalls, the accepted gap you wrote down converts the conversation from blame into a decision that was made, by a named person, with a date. That is usually when the requirement gets funded — and the difference between a security team that looks prescient and one that looks obstructive is entirely whether the gap was recorded beforehand.
- Why not escalate to leadership and have the ruleset mandated?Because a mandate you win over an unwilling owner tends to be implemented minimally, drift back, and cost you the relationship you need for the next twenty asks. Escalation is worth spending when the exposure is genuinely severe and the negotiation has honestly failed — and it lands far better when you can show the shrunken ask, the pilot measurement, and the refusal on record.
- What makes an accepted collection gap useful rather than paperwork?A named accepting owner with the authority to accept it, a specific statement of what cannot be answered, and an expiry date so silence does not renew it. Then the SOC's own reporting has to honour it: those hosts are marked unhunted for that technique, which removes assurance nobody had actually earned.
- How do you avoid making this a standing fight on every fleet?Move the requirement into the build image and the platform baseline, so coverage is inherited by new hosts rather than negotiated per fleet. Agreeing a minimum telemetry standard once, with the platform owners who will implement it, converts an endless series of individual asks into a property new systems arrive with.
saying these in an interview costs you the question
- Escalates immediately instead of shrinking the request
- Argues overhead by assertion rather than measuring it
- Justifies the ask on one hunter's blocked hypothesis
- Lets the security team accept the risk on the owner's behalf
- Keeps reporting the uninstrumented hosts as clean