skip to content

Your Linux fleet owner refuses a fleet-wide auditd execve rule on CPU and disk grounds - what now?

level: principalimportance: nice to knowfreq 30%

answer

  1. separate the CPU claim from the disk claim
  2. execve is rare next to file I/O
  3. your own automation makes the volume
  4. auid survives sudo
  5. make the loss explicit and owned

basics

~20 s

Split the objection. Execve-only auditing is far cheaper than the broad syscall rule sets people picture, and most of the volume is your own configuration management. Negotiate a measured rule, then get any accepted gap named and owned.

solid answer

~50 s

Take the two claims apart. **CPU**: auditd's cost scales with how often the audited syscall fires, and `execve` is rare next to file opens and connects, so an execve-only rule is a fraction of the broad rule packs people picture. **Disk**: usually true, and usually because of you — a configuration-management run executes dozens of processes per host, each producing a `SYSCALL` plus `EXECVE` plus `CWD` and `PATH` group. That is the fleet's benign true positive of mass remote execution. So bring a pilot ring with measured bytes, keep the kernel rule narrow rather than filtering in the kernel, and reduce volume downstream — while refusing to drop the automation account, because an intruder using that path vanishes with it. Then write down which questions the fleet cannot answer without the rule, and have the fleet owner accept that named risk.

code

text · 8 lines
text
type=SYSCALL msg=audit(1772630472.331:4412): arch=c000003e syscall=59 success=yes exit=0
  ppid=2841 pid=2903 auid=1002 uid=0 euid=0 ses=37 comm="python3.9"
  exe="/usr/bin/python3.9" key="exec"
type=EXECVE msg=audit(1772630472.331:4412): argc=3 a0="/usr/bin/python3.9"
  a1="/home/deploy/.ansible/tmp/ansible-tmp-1772630472.2-9931/AnsiballZ_command.py"
  a2="_ansible_check_mode"
type=CWD msg=audit(1772630472.331:4412): cwd="/home/deploy"
type=PATH msg=audit(1772630472.331:4412): item=0 name="/usr/bin/python3.9" ...

go deeper

for a junior

Know that Linux records nothing about process execution until an audit rule for the execve syscall exists, and that such a rule produces several joined records per execution rather than a single command line.

for a middle

Explain the record group — SYSCALL, EXECVE with arguments split across a0, a1, a2, plus CWD and PATH — and which fields identify the actor, particularly auid surviving su and sudo.

for a senior

Show you can tell your own configuration-management burst from mass remote execution, and that you check the audit backlog and lost counter before treating quiet hosts as clean.

for a principal

Own the negotiation and its aftermath: separate the CPU claim from the volume claim, bring measured pilot numbers, refuse the exclusions that create blind spots, and get any accepted gap named, owned and given a review date.

## Why this is a negotiation and not a configuration task On Linux there is no execution audit trail until someone writes a rule. The kernel audit subsystem records what its rules tell it to, so a rule like an `always,exit` filter on the `execve` syscall for the 64-bit architecture, tagged with a key, is the difference between a fleet you can investigate and one you cannot. That rule lives in a configuration the platform team owns, on hosts whose performance they are accountable for. Winning the argument on merit is the job. ## Taking the cost claim apart **The CPU claim.** Audit rules are evaluated in the syscall path, so cost scales with the *frequency of the audited syscall*. The reputation auditd has for being expensive comes from broad rule packs that audit file opens, writes, network calls and permission changes on busy servers. Process execution is orders of magnitude rarer than file I/O on a typical server. An execve-only rule is therefore not the thing being objected to — but it is what the fleet owner heard, because previous requests asked for everything. **The disk and wire claim.** This one usually survives scrutiny, and the reason is instructive: **your own automation dominates the volume.** A configuration-management run — Ansible, Salt, Puppet, a deployment pipeline — executes many short-lived processes per host per run, and each execution produces a record group: the `SYSCALL` record with `pid`, `ppid`, `uid`, `auid` and the result, an `EXECVE` record carrying the arguments split as `a0`, `a1`, `a2` (hex-encoded when they contain awkward characters), plus `CWD` and `PATH` records, all joined by one event serial number. Multiply by the fleet and the run cadence and the number is genuinely large. That same fact is the leaf's other lesson: **a configuration-management push is indistinguishable in shape from mass remote execution across the fleet.** It is this category's benign true positive, and if you do not know your own automation's identity and pattern, you will either chase it or, worse, learn to ignore the shape entirely. ## What to negotiate for - **Scope the kernel rule tightly, not the wrong way.** Ask for execve on the architectures actually in use, keyed for retrieval, and do not ask for broad file or network syscall auditing in the same breath. Conceding what you do not need is what makes the remaining ask credible. - **Reduce volume downstream, not at the kernel.** Once records exist you have choices about what to forward and how long to keep it. Filtering at the kernel means the evidence never existed. - **Do not simply drop the automation account.** It is the obvious saving and it is a hole you must name aloud: an intruder who runs through the same orchestration path, or with the same service identity, disappears with the filter. If you must reduce it, reduce by *volume sampling with full retention of anything unusual*, and say what you traded. - **Use `auid`.** The login uid is recorded per process and survives `su` and `sudo`, so a root-owned process still carries the identity of the human session behind it; daemon-spawned processes carry an unset value. It is the single most useful field for separating your automation from a person, and it is also what makes selective handling defensible rather than blind. - **Watch for silent loss.** The audit backlog has a limit; when it is exceeded, records are dropped and the `lost` counter rises. A gap you know about is manageable; a gap you do not is a false sense of coverage. Ask for the backlog setting and monitoring of that counter as part of the deal. - **Bring a measured pilot.** One ring of hosts, real bytes per host per day, real CPU deltas under the fleet's actual workload. Arguments lose to numbers, in both directions — and if the numbers say the objection was right, you have learned something worth knowing. ## The part that makes this a leadership question You may lose. The professional outcome then is not to shrug: it is to make the loss **explicit and owned**. Write down the questions the fleet cannot answer without execution records — what ran, under whose login, with what arguments, in what order — state that an intrusion on those hosts will be unreconstructable, and have the fleet owner accept that named risk with a review date. Attach it to the pilot measurements so the decision can be revisited on evidence rather than re-argued from scratch. The reason to insist on that ceremony is the failure mode. A missing execution rule does not fail loudly on the day it is refused. It fails months later, quietly, when there is an intrusion and nothing to reconstruct it from — and by then nobody remembers whose call it was.

  • Which auditd field tells you the human behind a process running as root?
    `auid`, the login uid. It is set when a session is established and is inherited across `su` and `sudo`, so a root-owned process still carries the account that logged in. Processes spawned by daemons carry an unset value, which is itself informative — a root shell with no login uid behind it is a very different thing from one with a name attached.
  • Why not simply exclude the configuration-management account's executions to cut the volume?
    Because that is the path an intruder would most like to use. Orchestration accounts run as root across the whole fleet, and a blanket exclusion makes anything running through them invisible. If volume forces a reduction, sample rather than exclude, keep anything that deviates from the automation's normal shape, and record the trade explicitly.
  • The fleet owner still refuses after your pilot. What do you actually do?
    Write the risk down in the terms the business understands: on these hosts, an intrusion cannot be reconstructed — no record of what ran, under which login, with which arguments. Get the fleet owner to accept that named risk with a review date, attach the pilot numbers, and revisit on evidence. Losing the argument is acceptable; losing it silently is not.

saying these in an interview costs you the question

  • Accepts the refusal and collects no execution telemetry at all
  • Claims execve auditing costs the same as full syscall auditing
  • Excludes the automation account wholesale to cut volume
  • Filters at the kernel rule rather than downstream
  • Assumes no records means no executions, ignoring backlog loss
  • Treats a configuration-management burst as mass remote execution

context