On a busy Kubernetes cluster, how do you balance audit coverage against API server latency and audit volume, and how do you choose between the log and webhook audit backends and their modes?
answer
- policy is the volume budget
- system chatter dominates, not humans
- measure with the audit metrics
- blocking favours completeness
- full buffer means silent drops
basics
~20 sBudget volume through the policy: drop or thin high-rate system traffic, omit RequestReceived, keep bodies only for rare writes. Then choose backend modes by what must win when the sink fails: blocking protects completeness, batch protects latency and may drop events.
solid answer
~40 sAudit cost is paid inside kube-apiserver on every request that matches a non-`None` rule, so the policy is the volume budget. I measure first with `apiserver_audit_event_total` and `apiserver_audit_level_total`, then thin what dominates: kubelet lease and status writes, controller reads, health probes - always narrowed by identity and resource together, never a whole group on its own. I omit `RequestReceived`, set `omitManagedFields`, and keep `RequestResponse` for rare writes. For the sink, the log backend writes to control-plane disk and defaults to `blocking`; the webhook backend defaults to `batch` with a 10,000-event buffer, which drops events when full. `blocking-strict` fails requests whose `RequestReceived` event cannot be recorded. I pick per requirement: an auditor who needs completeness gets blocking plus shipping; latency-sensitive clusters get batch plus alerting on `apiserver_audit_error_total`.
code
yaml · 24 linesapiVersion: audit.k8s.io/v1
kind: Policy
omitStages:
- RequestReceived
omitManagedFields: true
rules:
- level: Metadata
resources:
- group: ""
resources: ["secrets", "pods/exec"]
- level: None
userGroups: ["system:nodes"]
verbs: ["update", "patch"]
namespaces: ["kube-node-lease"]
resources:
- group: "coordination.k8s.io"
resources: ["leases"]
- level: Request
userGroups: ["system:nodes"]
verbs: ["update", "patch"]
resources:
- group: ""
resources: ["nodes/status", "pods/status"]
- level: Metadatago deeper
Recall that audit logging has a cost on the API server and that there are two backends: a local file and a remote webhook.
Explain the three modes, batch, blocking and blocking-strict, and what each does when the backend is slow or failing.
Show how to measure audit volume with the audit metrics, where system traffic dominates, and how to thin it by exact identity without losing high-signal rules.
Make the completeness-versus-availability call explicitly per cluster class, document it, and tie backend mode, buffer sizing and alerting to what the organisation must be able to prove.
## Where the cost lands Kubernetes audit logging runs **inside kube-apiserver's request path**. Every request that a policy rule assigns a level other than `None` is serialised once per non-omitted stage and handed to each configured backend. The audit filter wraps API Priority and Fairness, so flow control limits how many requests execute at once but does not reduce the audit work per request - and requests that flow control rejects with 429 are still audited. That is why audit-log volume is a capacity question for the API server, not just a storage bill. ## Where the volume comes from Take a 12-node GPU cluster serving models. Before any user does anything: - each kubelet renews its node Lease about every 10 seconds (a quarter of the default 40-second lease), so 12 nodes produce 72 Lease updates a minute, or 103,680 a day; - kubelets also write `nodes/status` and `pods/status`; - the scheduler and controllers list and watch Pods and Nodes constantly; - leader-election Leases for control-plane components renew continuously. Add a model-serving autoscaler resizing inference Deployments - each Pod requesting, say, a 0.35-core CPU request alongside its GPU - and writes multiply. At `RequestResponse`, each event also carries full objects, and list responses on large clusters can be megabytes. ## The policy as a volume budget 1. **Measure.** `apiserver_audit_event_total` and `apiserver_audit_level_total` show rate and level mix; the log backend's file growth shows bytes. 2. **Thin the known high-rate system traffic** by exact identity and resource - for example `level: None` for `leases` updates from `userGroups: ["system:nodes"]` in `kube-node-lease`, and `level: Request` rather than `RequestResponse` for node status writes. 3. **Omit `RequestReceived`** at policy level except on rules where in-flight visibility matters. 4. **Set `omitManagedFields: true`** so server-side-apply bookkeeping does not bloat bodies. 5. **Reserve `RequestResponse`** for rare, high-value writes such as RBAC changes. 6. **Never trade away the high-signal rules** - Secret reads, exec and RBAC stay above every exclusion. ## Backends and modes The two upstream backends can run together; events go to both. | Aspect | Log backend | Webhook backend | |---|---|---| | Enabled by | `--audit-log-path` | `--audit-webhook-config-file` (kubeconfig format) | | Default mode | `blocking` | `batch` | | Where data lives | Control-plane node disk, rotated by `--audit-log-maxsize` (default 100 MB) and `--audit-log-maxbackup` (default 100 files) | A remote receiver you operate | | Main risk | Disk fills or rotation discards history before it is shipped | Receiver slow or down; buffer overflows | | Oversize events | Written as-is | Optional truncation with `--audit-webhook-truncate-enabled` | The modes, set by `--audit-log-mode` and `--audit-webhook-mode`: - **`batch`** - events are queued (default buffer 10,000) and sent asynchronously; the webhook defaults send up to 400 events per batch, at least every 30 seconds, throttled to 10 batches per second with a burst of 15. When the buffer is full, events are **dropped** and counted in `apiserver_audit_error_total`. Latency is protected; completeness is not. - **`blocking`** - each request waits for the backend to process its events. A backend error is ignored and the request still succeeds, but a slow backend slows every audited request. - **`blocking-strict`** - like `blocking`, except that a failure to record the `RequestReceived` event fails the request with an internal error, counted in `apiserver_audit_requests_rejected_total`. Completeness wins over availability. ## Protecting the record itself - The audit log file lives on each control-plane node; with several API server replicas, each writes its own file, and a complete picture needs all of them. - Anyone with root on a control-plane node can edit the local file, which is another reason to get events off the node promptly. - The webhook receiver is a high-value target: it sees every audited request, so its endpoint needs TLS and its credentials, held in the kubeconfig-format config file, need the same care as the API server's own. ## Making the call There is no universally right answer; decide from what the organisation must prove and what it can tolerate: - **Regulated or forensic-critical clusters** - blocking modes, a sink sized for peak, alerting on errors, and a documented acceptance that an audit outage degrades the API. - **Latency-critical serving clusters** - `batch`, generous buffer, alerting on `apiserver_audit_error_total`, and a policy thin enough that the buffer never fills in normal operation. - **Managed control planes** - the provider usually owns the flags and the policy; the decision becomes whether its fixed policy answers your three investigation questions, and what you add around it. Whatever the choice, the log file on a control-plane node is not the archive: it must be shipped off the node, and that pipeline's durability belongs to the logging platform.
- Does API Priority and Fairness reduce Kubernetes audit-log volume when a client floods the API server?No. The audit filter sits outside flow control, so every request is audited whether flow control executes it or rejects it with 429. APF bounds concurrent execution, not the number of audit events. A flooding client therefore still generates events, and a policy rule at `None` or a thinner level for that identity is what reduces audit volume.
- How do you know a Kubernetes webhook audit backend in batch mode is losing events?Watch `apiserver_audit_error_total`, which counts events that failed to be audited, including those dropped when the batch buffer is full, and look for audit buffer errors in the kube-apiserver log. Alert on any sustained increase. Then either raise the buffer and throughput settings, fix the receiver, or thin the policy so normal load fits.
- Why run both the log and the webhook Kubernetes audit backends at once?They fail differently. The local file in blocking mode survives a receiver outage, while the batched webhook gets events off the control-plane node quickly without slowing requests. Together, a gap in one can be filled from the other, at the cost of doubled serialisation work inside the API server.
saying these in an interview costs you the question
- Human kubectl traffic is what drives most audit volume
- Batch mode never loses audit events
- Blocking mode fails the request whenever the backend errors
- API Priority and Fairness throttling also throttles audit logging
- Excluding the whole system:nodes group is a safe volume fix
- The audit log file on the control-plane node is the archive