An emulation proved an exec into a privileged pod went unseen — do you detect it or prevent it?
answer
- ask whether the behaviour is ever legitimate
- the verb is legitimate, the pod shape is not
- the audit record ends where the shell begins
- exemptions are where detection belongs
- rollout time is an interim-rule window
basics
~20 sPrevent, primarily. A pod that runs privileged and mounts the node filesystem has almost no legitimate use, so admission control can make the technique impossible, while a detection only tells you it already happened. Detect the narrow exception paths the policy has to allow.
solid answer
~50 sThe deciding question is whether the behaviour has a legitimate use in this estate. Running a pod that is privileged and mounts the node root filesystem has essentially none, so the honest destination is a preventive change owned by the platform team: admission control refuses that pod shape, and the technique stops existing. A detection over the Kubernetes audit log would only tell you, minutes later, that someone already had the node. It is also thin evidence: the audit record proves an exec request was authorised and the connection upgraded, not what was typed into the shell afterwards. But `exec` itself is legitimate — engineers debug with it — so you cannot block the verb; you block the pod shape. Then route a second, smaller row to detect: exec into the pods the policy still exempts, such as node agents and storage drivers. Prevent the shape, detect the exception, and give each row a different owner.
code
json · 13 lines{
"kind": "Event",
"apiVersion": "audit.k8s.io/v1",
"stage": "ResponseComplete",
"verb": "create",
"user": { "username": "svc-deploy", "groups": ["system:authenticated"] },
"objectRef": { "resource": "pods", "subresource": "exec",
"namespace": "prod", "name": "debug-tools-x7q2" },
"requestURI": "/api/v1/namespaces/prod/pods/debug-tools-x7q2/exec?container=app&command=%2Fbin%2Fsh&stdin=true&tty=true",
"responseStatus": { "code": 101 },
"annotations": { "authorization.k8s.io/decision": "allow" },
"requestReceivedTimestamp": "..."
}go deeper
Be ready to say that a detection tells you the exec happened while an admission-time configuration change stops the privileged pod existing at all, and that only the second removes the technique.
Expect to run the legitimacy test out loud: exec is a normal engineering action, a privileged pod with a host mount is not, so you route the pod shape rather than the verb.
Demonstrate that you take both destinations for the right reasons — exemptions and rollout time — and that you arrive at the platform team with report-only data rather than a demand.
Own the cost argument: every gap routed to a rule instead of a configuration change adds permanent triage load, so the routing default across the programme is a strategic choice, not a per-finding one.
## The finding During an emulation, an operator created a pod that ran with `privileged: true` and a `hostPath` volume mounting the node root, then opened a shell inside it and read files off the host. Nothing alerted. The exercise produced one clean piece of evidence: the cluster audit log recorded the request. ## What the record proves, and what it does not A Kubernetes audit event for an exec carries the identity that made the request, the source address, the target namespace and pod, the `exec` subresource, the initial command from the request URI, the authorisation decision, and the response status. A successful shell shows a protocol upgrade rather than a normal 200. What it does **not** carry is the session. Everything typed after the shell opens travels inside the upgraded connection and never reaches the audit log. So a detection built here proves that a shell was opened in a pod, not that anything bad was done in it. That asymmetry matters for routing: the detection you could build is low-fidelity and would need a human to reconstruct intent afterwards, while the preventive control acts on the pod specification itself, which is fully described in the create request the API server already evaluates. ## The routing test Ask one question first: **does this behaviour have a legitimate use in this estate?** - If yes, prevention will break real work, and detection with good context is usually the right destination. - If no, prevention removes the technique outright and every future alert it would have produced. Here the answer splits. `exec` in general is legitimate — engineers debug live workloads with it, and blocking the verb would be an outage waiting to happen. But *a privileged pod with the node filesystem mounted* is not a shape any ordinary workload needs. So you do not route the verb, you route the shape: admission control rejects pod specifications that request privileged execution or a host mount outside an explicit exemption list. That is a platform-team change, not a security-team change, and it removes the technique rather than watching it. ## Why you usually take both Routing is not exclusive, and a good answer says so: 1. **Exemptions.** Real clusters must permit some privileged workloads — node agents, storage drivers, monitoring daemon sets. Those exempted pods are exactly where the same technique still works, so an exec into them is a small, high-signal detection worth owning. 2. **Rollout time.** If the policy takes a quarter to reach every namespace, you are exposed for a quarter. An interim detection covers the window and is retired when the policy lands. 3. **Attempt visibility.** A rejected pod creation is itself worth knowing about; a policy that silently denies teaches you nothing about who keeps trying. ## When you route to accept instead If the cluster is a legacy build where half the workloads currently violate the shape, prevention is a migration programme, not a change. The honest outcome may be: an interim detection now, a prevention row with a real date owned by the platform team, and — if that date cannot be committed — an explicitly recorded acceptance with a named person, rather than a prevention row that quietly never moves. ## The failure modes an interviewer listens for - **Routing everything to a rule** because the security team can land a rule without asking anyone. That converts a solvable configuration problem into a permanent triage obligation. - **Claiming the audit log shows what the attacker did.** It shows a request was authorised, not the contents of the session. - **Blocking `exec` outright** and discovering the on-call engineers now cannot debug production. - **Handing the platform team a policy with no blast-radius data.** Running the policy in a report-only mode first and bringing the list of pods that would have been rejected is what turns a demand into a plan. ## The shape of the answer Prevent the pod shape, detect the exempted residual, cover the rollout window with an interim rule, and give the prevention row to the team that owns the cluster with a date and a re-execution booked to prove it.
- What can you conclude from that exec audit record, and what can you not?You can conclude that a named identity requested a shell in a specific pod, that authorisation allowed it, and that the connection upgraded, so a session opened. You cannot conclude what was run inside it: keystrokes and output travel in the upgraded stream and are never written to the audit log. Anything about behaviour inside the container has to come from container or node telemetry instead.
- The platform team says the admission policy will break existing workloads. What do you bring back?Data, not insistence. Run the policy in a report-only mode across the cluster and return the list of pods that would have been rejected, by namespace and owner. That converts an argument about risk into a scoped migration with a known size, and it usually shrinks the objection to a handful of genuine node-level agents that become the exemption list your detection then covers.
- Would you ever route this gap to a detection only?Yes, when the preventive change cannot land soon: a legacy cluster where many workloads still use the shape, or a platform team with no capacity this quarter. Then the detection is explicitly interim, the prevention row stays open with an owner and a date, and if that date never gets committed the residual is recorded as an accepted gap rather than left implied.
saying these in an interview costs you the question
- Routes it to a detection because the security team can ship a rule alone
- Blocks the exec verb entirely and breaks legitimate debugging
- Claims the audit log shows the commands typed inside the shell
- Treats detect and prevent as mutually exclusive
- Hands the platform team a policy with no blast-radius data
- Forgets that exempted privileged pods keep the technique alive