skip to content

Why can a single Kubernetes API request produce several audit events, and why do audit Policies often omit the RequestReceived stage?

level: middleimportance: nice to knowfreq 25%

answer

  1. one request, several moments
  2. shared auditID joins them
  3. long-running calls add a middle stage
  4. final stage already has the outcome
  5. omitting trades in-flight visibility

basics

~20 s

kube-apiserver can emit one audit event per stage - RequestReceived, ResponseStarted for long-running calls, ResponseComplete, and Panic - all sharing one auditID. ResponseComplete already carries the outcome, so omitting RequestReceived roughly halves volume for ordinary requests.

solid answer

~40 s

An audit event records a request at a particular **stage**. `RequestReceived` fires as soon as the audit handler sees the request, before it is passed on to impersonation, flow control and authorization. `ResponseStarted` fires only for long-running requests such as a `watch` or `exec`, when the headers go out. `ResponseComplete` fires when the response body is finished, and `Panic` when handling crashed. Events for one request share an `auditID`. Most policies set `omitStages: [RequestReceived]` because `ResponseComplete` repeats everything and adds the response code, so the early event mostly doubles volume. The cost of omitting it: a request that never completes leaves no trace until it ends.

code

yaml · 20 lines
yaml
apiVersion: audit.k8s.io/v1
kind: Policy
rules:
  # Keep the early stage for interactive sessions
  - level: Metadata
    resources:
      - group: ""
        resources: ["pods/exec", "pods/attach", "pods/portforward"]
  # Bulk node status writes: final stage only
  - level: Request
    userGroups: ["system:nodes"]
    verbs: ["update", "patch"]
    resources:
      - group: ""
        resources: ["nodes/status", "pods/status"]
    omitStages:
      - RequestReceived
  - level: Metadata
    omitStages:
      - RequestReceived

go deeper

for a junior

Recall that one API request can show up as several audit records, one per stage, and that they share an auditID.

for a middle

Explain the four stages, which requests get ResponseStarted, and why ResponseComplete alone usually answers the question.

for a senior

Show that omitting RequestReceived hides long-running sessions until they end, and keep that stage for exec and RBAC rules while dropping it for bulk traffic.

for a principal

Weigh stage selection as part of the volume budget and of what the organisation must be able to reconstruct about requests still in flight during an outage.

## Stages in the request path The Kubernetes API server does not write one audit record per request. It writes one record per **stage** the request reaches, provided the policy level for that request is not `None` and the stage is not omitted. The stages are defined in `audit.k8s.io/v1`: | Stage | When it is emitted | What it lacks or adds | |---|---|---| | `RequestReceived` | As soon as the audit handler receives the request, before it goes further down the handler chain | No response status yet | | `ResponseStarted` | When response headers are sent but before the body - only for long-running requests | Status known, body still streaming | | `ResponseComplete` | When the response body is complete and no more bytes will be sent | Full outcome and latency | | `Panic` | When a panic occurred while handling the request | Marks a crashed request | All events for the same request carry the same `auditID`, so a consumer can join them. The `stage` field says which one a record is, and `stageTimestamp` says when it was written, alongside the fixed `requestReceivedTimestamp`. ## Where audit sits in the handler chain In kube-apiserver the audit filter runs right after authentication and **before** impersonation, API Priority and Fairness and authorization. Three consequences follow: - a `RequestReceived` event exists even for a request that authorization later denies; - a request that flow control rejects with HTTP 429 is still audited, and its `ResponseComplete` shows that code; - a request whose credentials fail authentication is audited through a separate failed-authentication handler, so rejected logins leave a trace too. ## Long-running requests For an ordinary `get` or `patch`, the gap between `RequestReceived` and `ResponseComplete` is milliseconds. For a long-running request it can be large: - a `watch` sends headers early (`ResponseStarted`) and only completes when the watch closes, which is often many minutes later; - `kubectl exec`, `attach` and `port-forward` complete only when the session ends. On a 12-node GPU cluster serving models, an operator who opens a shell into an inference Pod at 09:14 and leaves it open until 11:40 produces a `ResponseComplete` event stamped 11:40. If the policy omits `RequestReceived` and the investigator filters by the time of the incident, they may miss the session that was open during it. ## Why policies omit RequestReceived 1. **Volume.** For short requests `RequestReceived` duplicates who, what and when, and adds nothing `ResponseComplete` does not also carry. Omitting it roughly halves event count for those requests. 2. **Signal.** `ResponseComplete` includes the response code and the authorization annotations, which is what almost every question about a request needs. 3. **Defaults.** Policy-level `omitStages` is merged into every rule, so one line at the top covers the whole policy; a rule can add further stages of its own. What you give up: - **In-flight visibility.** A request that hangs, or a session still open, is invisible until it ends. - **Crash evidence.** If the API server process dies mid-request, no later stage is written for it. - **Strict mode meaning.** With the `blocking-strict` backend mode, a failure to record the `RequestReceived` event fails the request itself. If that stage is omitted, there is no event at that point for the strict check to act on. ## Reading a pair of events For a short `patch` on a Deployment with no stages omitted, the log holds two lines with the same `auditID`: - the first has `stage: RequestReceived`, the user, verb and `objectRef`, but no `responseStatus`; - the second has `stage: ResponseComplete`, the same identity fields, `responseStatus.code` (for example 200 or 403), the authorization annotations and, at `Request` level or above, the body. A consumer that counts lines as requests double-counts unless it filters on one stage. Picking `ResponseComplete` as the canonical record is the usual convention, with `RequestReceived` used only to spot requests that never finished. ## Choosing per rule - Keep `RequestReceived` for the few rules where an in-flight record matters, such as `pods/exec` or RBAC writes, by leaving it out of those rules' `omitStages` and not omitting it at policy level. - Omit it for bulk traffic such as node status updates and controller reads. - When querying, find long-running sessions by `requestReceivedTimestamp`, which every stage's event carries, rather than by when the event was written.

  • If a Kubernetes audit Policy sets omitStages at policy level and a rule sets a different stage, which stages does that rule omit?
    Both. When kube-apiserver loads the policy it merges the policy-level `omitStages` into each rule's own list as a union. A rule cannot re-enable a stage that the policy level omits, so a stage you want on some rules must be omitted per rule rather than globally.
  • Why does a Kubernetes watch request show a ResponseComplete audit event long after it started?
    A watch is a long-running request: the server sends headers immediately, which produces `ResponseStarted`, then streams changes until the watch is closed. `ResponseComplete` is written only when no more bytes will be sent. The event's `requestReceivedTimestamp` still records when the watch began, so query on that field when you need what was open at a given moment.

saying these in an interview costs you the question

  • Each audit rule that matches writes its own event
  • ResponseStarted is emitted for every request
  • RequestReceived is written only after authorization succeeds
  • Omitting RequestReceived loses the response code
  • Events for one request cannot be correlated