skip to content

The Engine's Blast Radius

A mutator that matches every Pod can add a container to every workload, and the engine receives whatever it matches. Interviewers ask what your enforcement plane is worth to an attacker.

on this pageshow

questions

4

Why is a mutating admission rule that injects a sidecar into every Pod riskier than a validating rule?

level: juniorimportance: must knowfreq 60%

answer

  1. one says no, the other rewrites
  2. what runs is not what was submitted
  3. cluster-wide, no manifest change
  4. injection primitive if compromised

basics

~20 s

A validating rule can only accept or reject. A mutating rule rewrites the object that actually runs, so whoever controls it can add a container, volume or environment variable to every Pod it matches, cluster-wide.

solid answer

~50 s

A validating rule returns a verdict; a mutating rule returns a patch, and the patched object is what gets stored and run. Teams reach for mutation because it drives adoption: an egress guardrail that injects a proxy sidecar and the matching `NO_PROXY` environment into every Pod means no team has to edit a manifest. Inverted, that same rule is a cluster-wide arbitrary-workload-injection primitive. Whoever controls what the mutator returns chooses what image runs beside every matched Pod, what it mounts and what it can see. It is also silent: the injected container is absent from the developer's manifest and from Git, and the stored object carries no field naming the webhook that added it. The plane that would normally notice a rogue container is the plane that added it. Mutation only acts on requests, so existing Pods pick it up as they are recreated.

go deeper

for a junior

Be ready to state the difference in one line: a validating rule returns allow or deny, a mutating rule returns a patch and the patched object is what runs. Then say why that makes the mutator the more valuable target.

for a middle

Explain the mechanics: the patch is applied before the object is persisted, mutation happens before validation, and the stored object records no attribution for who patched it. Be able to say why the injection only reaches Pods as they are recreated.

for a senior

Show you would treat an operating mutator as a privileged workload: pin what it injects, narrow what it matches, and arrange for drift between live objects and source to be detected by something other than the engine itself.

for a principal

Own the tradeoff. Mutation buys estate-wide adoption without asking hundreds of teams to change anything, and pays for it with a cluster-wide injection primitive. Be able to say which guardrails are worth that and what compensating controls make it acceptable.

## Two kinds of rule, two very different kinds of power An admission rule sits between a request to the Kubernetes API server and the object being persisted. A **validating** rule can only return a verdict: allow, or deny with a message. A **mutating** rule returns a patch, and the API server applies it — the object that ends up in etcd, and therefore the workload that actually runs, is the patched one. Mutating admission runs first; validating admission then sees the already-patched object. ## Why teams choose mutation anyway Take a real guardrail: all outbound traffic from workloads must leave through an egress proxy. Written as a validating rule, it rejects any Pod without the proxy sidecar — correct, but it hands work to every team in the estate, and the rollout is a migration with hundreds of pull requests. Written as a mutating rule, the platform injects the proxy sidecar and the matching `NO_PROXY` environment into every Pod at admission. Nobody edits a manifest, adoption is instant, and compliance is a property of the platform rather than of each team's diligence. This is a genuinely good reason to mutate, and it is why mutating rules are common. ## The same rule, inverted Now read that capability as an attacker would. The rule is a mechanism that inserts an attacker-chosen container into every Pod that matches, across every namespace, with whatever image, mounts and environment the patch specifies. Compromise the engine — or the content of the rules it serves — and you do not need access to a single team's repository, a single ServiceAccount, or a single node. You get workload execution wherever the match reaches. Three properties make this worse than the equivalent compromise of a normal deployment: - **It is silent.** The injected container is not in the submitted manifest and not in the source repository. It *is* in the stored object, but the stored object has no field saying "a webhook added this". Someone comparing live state to source will see the difference; nobody comparing a Pod to its own spec will. - **The watcher is the thing that changed it.** Validating rules that might have flagged an unexpected container usually come from the same engine, and they run after the mutation. A plane cannot be trusted to police itself. - **It spreads on ordinary operations.** Admission acts on requests, so already-running Pods are untouched until they are recreated. A node drain, a rollout or a scale event is what carries the injection outward — which also means it can arrive hours or days after the compromise. ## Validating rules are not harmless A validating rule cannot change what runs, but it still receives the full body of every object its match rules select, and a rule that denies can deny broadly. So the distinction is specific: mutating rules add **write into every matched workload**; both kinds already carry **read of every matched object**. ## Shrinking the blast radius The questions to ask about any mutation you operate are: does this rule need to mutate at all, or would rejecting with a clear message do? Does it need an external engine, or can it be an expression the API server evaluates itself — a `ValidatingAdmissionPolicy` runs CEL in-process, cannot make network calls, and can only reject, never modify, so there is no remote endpoint to compromise. And how narrow is the match? The set of requests the API server sends to the engine is the set of objects the engine can read and, for a mutator, the set of workloads it can write into. Narrowing that match is least privilege for the engine, not throughput tuning. Finally, the mutation content itself — the injected image, its registry, its digest, the volumes it mounts — deserves the review you would give any privileged workload, because that is exactly what it is.

  • So a validating rule carries no risk at all?
    It cannot change the stored object, but it still receives the full body of every object its match selects, so it is a read path over everything it matches. And a rule that denies is a denial primitive: broaden its deny condition and nothing ships. What it cannot do is choose what code runs.
  • If the mutator is compromised, why wouldn't your validating rules catch the injected container?
    Validating admission runs after mutation, so it does see the injected container — but the validating rules almost always come from the same engine, and they were written to allow the sidecar the platform injects. The plane checking the object is the plane that changed it. Independent detection has to come from outside, by comparing live objects to their source.
  • What limits how far a mutator you have to keep can reach?
    Its match rules. The API server only sends the engine requests that match on resource, namespace and object selectors, so that match defines both which workloads can be injected into and which object bodies the engine ever sees. Treat narrowing it as least privilege for the engine, not as a performance setting.

A validating rule is a bouncer who can only turn people away. A mutating rule is a bouncer who can also slip something into every bag on the way in.

saying these in an interview costs you the question

  • Says mutating rules only add defaults, so they are low risk
  • Assumes the injected container shows up in the team's manifest
  • Treats validating and mutating rules as equally privileged
  • Expects the same engine's validating rules to catch its own mutation
  • Thinks a compromised mutator instantly changes already-running Pods

context

open as a page

In a Kubernetes AdmissionReview, what does the API server actually send an engine whose match includes Secrets?

level: middleimportance: should knowfreq 52%

basics

~20 s

The full object body — for a Secret, its complete data — plus the previous version on an update, the requesting user and groups, the operation and dryRun. It carries no other object and no cluster state.

open as a page

Your admission engine may be injecting an undeclared sidecar into every Pod: how do you confirm it and contain it?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Compare live objects with the manifests that produced them: the extra container is in the stored object but not the source. Then remove the mutating registration so the API server stops calling the engine, and roll the workloads.

open as a page

Your admission rule needs data from other cluster objects: should the engine get a standing cluster-wide read?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Only if the rule is worth the credential. An admission request carries no other object, so the engine must fetch state itself, and a standing cluster-wide read turns the enforcement plane into a cross-namespace reader and an escalation target.

open as a page