skip to content

Your Gatekeeper Assign is installed but the Pods come out unchanged — how do you debug it?

level: seniorimportance: should knowfreq 43%

answer

  1. no denial means no diagnostic
  2. check the Pod, not the Deployment
  3. core group is the empty string
  4. server dry-run shows the stored shape
  5. pathTests skip without saying so

basics

~20 s

A mutator that matches nothing fails silently, so work outward: confirm which object is actually admitted, then applyTo's group/version/kind, then the match block, then the location path and pathTests — and verify on a server-side dry-run of a real object.

solid answer

~50 s

There is no error to read, so the debugging is a checklist rather than a log hunt. First, check you are looking at the right object: a mutator on Pod does fire for Pods created by a Deployment, but the Deployment's own YAML never changes, so `kubectl get deploy -o yaml` will always look untouched. Second, `applyTo` — core kinds need `groups: [""]`, and `"core"` or `"v1"` there matches nothing without complaint. Third, `match` — namespace lists, label selectors, `scope`, and the namespaces the Gatekeeper install exempts by default. Fourth, the `location` path: a glob over a list that has no matching entries, or the wrong list entirely. Fifth, `pathTests`, which skip silently by design. Confirm each fix with `kubectl apply --dry-run=server -o yaml` and diff the result against the manifest — server dry-run runs admission, so mutations show up without anything being created.

go deeper

for a junior

Know that a mutator never reports failure, so an unchanged object is not evidence of anything by itself. Being able to name the three selectors — applyTo, match, location — is what is expected here.

for a middle

Be ready to give the checklist in order and to say what each step rules out, including why a mutator on Pod leaves the Deployment YAML looking untouched.

for a senior

Demonstrate falsifiable debugging: establish that the mutator is being consulted at all before touching the path, and check every hypothesis against a server-side dry-run rather than a redeploy.

for a principal

Own the observability gap. Decide what the platform runs so a broken mutator surfaces on its own — mutation logging, coverage measured cluster-wide — rather than being found by whoever eventually notices.

## Why this is hard Every other part of Gatekeeper talks back. A Constraint that denies returns a message; a Constraint in `dryrun` records violations in its own status; a broken ConstraintTemplate reports its compilation error. A mutator does none of this. Matching nothing and matching everything produce the same observable outcome when the field was already correct, and a rule with a typo in it is indistinguishable from a rule that had nothing to do. So you debug by narrowing, in the order that costs least. ## The checklist **1. Are you inspecting the object that was admitted?** The most frequent false alarm. You applied a mutator to `Pod`, deployed a `Deployment`, and then read the Deployment back. Deployments are admitted as Deployments; the Pod is created later by the ReplicaSet controller and *that* create goes through admission and gets mutated. Look at a Pod. Conversely, if you targeted the Deployment, your `location` has to walk through `spec.template.spec`, not start at `spec`. Also remember that Pods running before the mutator existed were admitted under the old rules. A rollout is what puts them through admission again. **2. `applyTo`** Three fields, all exact-match lists: `groups`, `versions`, `kinds`. Core objects such as Pod live in the empty group — `groups: [""]`. Writing `"core"` or `"v1"` there is the single most common typo, and it produces no validation error: the mutator is accepted, stored, and never selects anything. **3. `match`** `applyTo` says which kinds, `match` says which instances. Check `namespaces` and `excludedNamespaces`, `labelSelector` and `namespaceSelector`, `name`, and `scope`. Then check the exclusions you did not write: a Gatekeeper installation is normally configured to leave system namespaces such as `kube-system` and its own namespace alone, and a test workload parked in one of those will never be touched no matter how correct the rule is. **4. `location`** Walk the path by hand against a real admitted object. A glob over a list selects existing entries only; a path naming `containers` ignores `initContainers`; a mistyped key filter matches nothing. Missing intermediate map fields are created, so a path that is merely deep is fine — it is the list segments that fail quietly. **5. `pathTests`** A failing test skips the mutation with no record. `MustNotExist` in particular will make a rule look broken exactly when it is working: the field was already set, so nothing was written. **6. The plumbing** Is Gatekeeper's mutating webhook installed and healthy at all? Some installations disable mutation, and a mutator applied to a cluster whose mutating webhook is not running is inert but perfectly valid YAML. Check the controller pods and their logs. Gatekeeper can also be configured to log each mutation it applies and to annotate mutated objects with the mutators that fired — turn that on before you start guessing, because it converts the whole problem from inference back into reading. ## The verification loop The technique that ends most of these sessions is server-side dry-run. Submitting a real manifest with `--dry-run=server -o yaml` sends the request through the API server's full admission chain and returns the object as it would have been persisted — mutations applied — without storing anything. Diff that against what you submitted and you can see, per iteration, whether your last edit made the mutator fire. Client-side dry-run does not do this: it never reaches the API server, so no admission runs and no mutation appears. For coverage rather than a single object, pair the mutator with a Constraint that checks the same field and leave it in a non-blocking mode; its violation list tells you which workloads the mutator is not reaching, across the whole cluster, instead of one manifest at a time. ## What a strong answer sounds like Order and falsifiability. A weak candidate starts editing the location path because that is the interesting part. A strong one establishes first that the mutator is even being consulted — right object, right kind, right namespace, webhook alive — and only then argues about the path, checking each hypothesis against a dry-run rather than against a redeploy.

  • How do you see the mutated object without creating it?
    Submit it with `--dry-run=server -o yaml`. The request goes through the API server's admission chain, mutations included, and the object comes back as it would have been stored, with nothing persisted. Client-side dry-run never contacts the server, so it shows you your own manifest and tells you nothing.
  • You target Pod but everything is deployed as Deployments — does the mutator fire?
    Yes. The ReplicaSet controller creates the Pod, and that create is an admission request, so the Pod is mutated. What confuses people is that the Deployment object itself is untouched, so reading the Deployment back shows no change and the rule looks dead.
  • What in applyTo goes wrong most often?
    The group. Core objects like Pod use `groups: [""]` — the empty string — and writing `"core"` or `"v1"` there matches nothing. The mutator is still valid YAML and is accepted without warning, so the only symptom is silence.
  • How do you find every workload the mutator is failing to reach?
    One object at a time does not scale. Pair the mutator with a Constraint checking the same field, left in a non-blocking mode, and read its violation list — that gives you a cluster-wide inventory of what is still unrepaired instead of a per-manifest guess.

saying these in an interview costs you the question

  • Looks for a denial message from a mutator
  • Checks the Deployment instead of the created Pod
  • Writes core or v1 in applyTo groups
  • Assumes an unmatched location raises an error
  • Uses client-side dry-run to check a mutation

context