skip to content

At Metadata audit level, a Kubernetes entry shows a create on pods/exec into production — what can you establish?

level: middleimportance: should knowfreq 52%

answer

  1. levels: None, Metadata, Request, RequestResponse
  2. metadata includes the request URI
  3. exec arguments ride in the query string
  4. the session is a stream, not a body
  5. benign true positive, not false positive

basics

~20 s

That an authenticated identity was allowed to open an exec session into a named pod and container, when and from where, plus the command argv in the request URI. Nothing typed inside the session is audited.

solid answer

~50 s

Metadata level records everything about the request except bodies: the authenticated user and groups, verb `create`, the `objectRef` naming the namespace, pod and `exec` subresource, source IPs, user agent, timestamps and the response status. For exec specifically the `requestURI` query string still carries the requested command, container and `tty` flag, so you can usually see that an interactive shell was asked for. What you cannot get is the session itself — exec upgrades to a bidirectional stream, so keystrokes and output are outside the audit model entirely, and raising the policy to RequestResponse would not capture them. That leaves a verdict built on identity: who the account belongs to, whether a change ticket or deployment covers the window, and whether the operator confirms it. When they do, close it as a benign true positive — the rule was right and the activity was authorised.

code

json · 15 lines
json
{
  "kind": "Event",
  "apiVersion": "audit.k8s.io/v1",
  "level": "Metadata",
  "stage": "ResponseComplete",
  "verb": "create",
  "requestURI": "/api/v1/namespaces/payments/pods/api-7d9f-cq2mx/exec?container=api&command=%2Fbin%2Fbash&stdin=true&stdout=true&tty=true",
  "user": { "username": "[email protected]", "groups": ["platform-sre", "system:authenticated"] },
  "sourceIPs": ["10.14.3.22"],
  "userAgent": "kubectl/v1.31.2 (linux/amd64)",
  "objectRef": { "resource": "pods", "subresource": "exec", "namespace": "payments", "name": "api-7d9f-cq2mx" },
  "responseStatus": { "code": 101 },
  "requestReceivedTimestamp": "2026-05-14T02:41:07.113Z",
  "stageTimestamp": "2026-05-14T02:49:31.902Z"
}

go deeper

for a junior

Know the four audit levels by name and that Metadata means no request or response body. Be able to point at the user, verb, objectRef and source IP in an audit event and say what each one is.

for a middle

Explain why exec is different from an ordinary create: the arguments travel in the request URI, and the session itself is a stream that no audit level captures. Say what body capture costs when Secrets are involved.

for a senior

Show how you reach a verdict without session content — identity, change correlation, the operator's own account — and why the durable fix is a recording broker or runtime telemetry, not a higher audit level.

for a principal

Own the tradeoff across the estate: which namespaces justify body capture, who accepts the secret-exposure and volume cost, and whether production exec should require an approval path at all.

## Kubernetes audit levels in one paragraph The API server evaluates every request against an audit policy whose rules select a **level**: `None` (record nothing), `Metadata` (record the request's metadata but neither request nor response body), `Request` (metadata plus the request body), and `RequestResponse` (metadata plus both bodies). Events are also emitted at **stages** — `RequestReceived`, `ResponseStarted`, `ResponseComplete`, `Panic` — so one call can produce more than one record. Most production clusters run Metadata as the default with narrow exceptions, for two reasons: bodies multiply log volume, and bodies of `Secret` and `ConfigMap` writes would put the secret material itself into the audit log, turning the audit trail into a credential store. ## What the Metadata-level exec entry gives you For a `kubectl exec` the API server records a `create` on resource `pods` with subresource `exec`. From that single record you can state: - **Who**: `user.username` and `user.groups` as authenticated by the API server, plus `impersonatedUser` if `--as` was used. - **Where from**: `sourceIPs` and the self-reported `userAgent` (`kubectl/v1.31.2`, a CI client, a controller). - **What target**: `objectRef` naming the namespace, pod name and container. - **When**: `requestReceivedTimestamp` and `stageTimestamp`, which bracket how long the session ran. - **Whether it was allowed**: `responseStatus`, and for a successful exec the protocol upgrade rather than a 403. - **What was asked for**: this is the part people miss. Because `kubectl exec` passes its arguments as **query parameters on the request URI**, and `requestURI` is metadata, the initial argv survives at Metadata level — you can usually read `command=%2Fbin%2Fbash`, the container name and `tty=true`, which is enough to say an interactive shell was requested rather than a one-shot command. ## What no audit level will give you The exec session is not a request body. The call **upgrades the connection to a bidirectional stream** carrying stdin, stdout and stderr, and that stream is not part of the audited request. So: - **Raising the rule to `Request` adds nothing** for exec — there is no JSON body to capture. - **Raising it to `RequestResponse` adds nothing either**, while paying the full cost of body capture across everything else the rule matches. Anyone who answers "I would turn up the audit level to see the commands" has just accepted a large volume and secret-exposure cost for zero additional evidence on the exact behaviour they cared about. That is the trap in the question. ## Where the missing evidence actually lives To reconstruct what happened inside the session you need a different observation surface: - **In-container process telemetry** from a runtime sensor watching process creation inside the pod. - **Node-level process records**, if the node still exists and its telemetry was shipped. - **A session-recording access broker** in front of the cluster, which is the only control that captures keystrokes by design. And here the ephemerality bites: if the pod has since restarted or the node was scaled away, none of those exist retroactively. The audit entry may be the only account of the session that survives, which is why the policy decision — which namespaces get audited at what level, and whether exec is routed to a detection — has to be made before the incident. ## Reaching a verdict With the session content unavailable, the verdict is built on identity and corroboration: map the username to a person or a workload, check whether a change record, incident bridge or deployment covers the timestamp window, look at what else that identity did in the same window across the cluster and the identity provider, and ask the named operator. In the common case an on-call engineer opened a shell in a production pod during a live degradation, which looks exactly like an intruder and is not one. That is a **benign true positive**: the detection was correct about the behaviour, and the behaviour was authorised. Closing it as a false positive would be wrong, and tuning the rule away because of it would delete a genuinely valuable detection. The durable improvement is not a higher audit level — it is making exec into production rare and attributable: an approval path, short-lived access, and a recorded broker session so the next entry arrives with the transcript already attached.

  • Would raising the policy rule to RequestResponse have captured the shell session?
    No. Exec upgrades to a bidirectional stream carrying stdin and stdout, and that stream is never part of the audited request or response body. RequestResponse captures bodies for ordinary create and update calls — which is why it also drags Secret contents into the audit log — but for exec, attach and port-forward it adds nothing at all.
  • The engineer confirms it was her, during a live outage. Do you tune the detection out?
    No. This is a benign true positive, not a false positive: the rule correctly identified a shell into production. Tuning it away removes the only signal you have for the malicious version of the same action. Reduce the volume instead by suppressing execs that correlate with an open change or incident record, and keep the rule firing for everything else.
  • You want a transcript of the next such session. What do you change?
    Not the audit level — it cannot capture streams. Put a recording access broker in front of the cluster so interactive sessions are captured by design, or run a runtime sensor inside the pod that logs process creation within the container. Both are decisions to make before the incident, since neither can be applied retroactively to a pod that has already restarted.

saying these in an interview costs you the question

  • Claiming RequestResponse level would capture the shell keystrokes
  • Assuming Metadata level hides the requested command entirely
  • Closing an authorised production exec as a false positive
  • Proposing full body capture without mentioning Secret exposure
  • Planning to image the pod later, after it has already restarted

context