skip to content

Your admission engine may be injecting an undeclared sidecar into every Pod: how do you confirm it and contain it?

level: seniorimportance: should knowfreq 42%

answer

  1. crosses team boundaries, follows Pod churn
  2. live object versus declared source
  3. submitted object versus persisted object
  4. removing the registration reverts nothing
  5. rotate whatever the match selected

basics

~20 s

Compare live objects with the manifests that produced them: the extra container is in the stored object but not the source. Then remove the mutating registration so the API server stops calling the engine, and roll the workloads.

solid answer

~50 s

First establish the shape: the same container appearing in namespaces owned by unrelated teams, arriving when Pods were recreated rather than when anything was merged, and present in no team's repository. That pattern points at admission, because admission is what stands between a submitted request and the stored object. Confirm it from API audit events, which show the object as submitted alongside the object as persisted — the container appears only in the second. Contain by removing or narrowing the mutating webhook registration, which stops the API server sending those requests at all. Then be clear about what that does not do: already-running Pods keep the sidecar until they are recreated, and every object matched during the window was sent to the engine in full, so matched Secrets and ConfigMaps are exposed and need rotating. Do not use the same engine's validating rules to verify the cleanup.

go deeper

for a junior

Know where to look first: the live object versus the manifest that declared it. Be able to say that admission can change an object between submission and storage, so a difference there is not necessarily a team's mistake.

for a middle

Explain how you would attribute the change — audit events showing the submitted object without the container and the persisted object with it — and why the affected set matching the webhook's scope is corroboration.

for a senior

Demonstrate sequencing and scope: stop the source by removing the registration, roll workloads deliberately, treat every object the match selected as disclosed, and verify with something outside the enforcement plane.

for a principal

Own the structural fix: who is accountable for the enforcement plane, and what independent detection exists so the next cluster-wide change is caught by something the plane does not control.

## Reading the symptom An extra container that nobody declared has three tells that separate an enforcement-plane compromise from an ordinary mistake: 1. **It crosses ownership boundaries.** The same container is in Pods belonging to teams with different repositories, different pipelines and different reviewers. No single team's change explains that. 2. **It appears at recreation, not at merge.** Workloads gained it when Pods were rescheduled, drained or rolled — not when anything landed in a repository. Admission acts on API requests, so its effect follows Pod churn. 3. **It is in the cluster and nowhere else.** Diff live objects against the source that produced them — the desired state your delivery tooling reconciles, or the recorded last-applied configuration. The injected container is present only in the stored object. ## Confirming it was admission The stored object itself will not tell you. There is no field on a Pod naming the webhook that patched it; once the patch is applied, the object looks exactly like a legitimately declared one. The place to look is the API server's audit trail: at an audit level that records both the request and the response object, you can see the object as the client submitted it and the object as it was persisted. If the container is absent from the first and present in the second, the change happened inside the admission chain, not in a client. Cross-check that the affected objects are exactly the set your mutating registration matches — that correspondence is strong evidence. ## Containing Containment is removing or narrowing the mutating webhook registration, which stops the API server forwarding those requests to the engine at all. Then be precise about what is and is not fixed: - **Running workloads do not revert.** The patch was applied at admission, so existing Pods carry the sidecar until they are recreated. Rolling them is a deliberate, sequenced step, not a side effect of removing the registration. - **Exfiltration is already done.** Every object the match selected during the window was sent to the engine in full. If the match included Secrets or ConfigMaps, treat the credentials in them as disclosed and rotate them; scoping that rotation is usually the largest part of the incident. - **The engine's own footprint is in scope.** It is a workload with credentials, network egress and, in most designs, a place it ships records of what it decided. All of that is part of the investigation. - **The injected sidecar is evidence.** Its image, the registry it came from, its mounts and its outbound connections tell you what the injection was for. ## The trap: verifying with the thing you are investigating The instinct is to ask the policy plane whether the cluster is clean now. Do not. The validating rules that would flag an unexpected container almost always come from the same engine, and they run after mutation, on the already-patched object. A compromised plane will happily report compliance. Verification has to come from outside it — a diff of live state against source, run by your delivery tooling or a separate reader, is the honest check. ## Afterwards Two structural lessons usually come out of this class of incident. The first is that the match was wider than any rule needed: the engine was receiving — and could inject into — far more than the guardrails required, so narrowing the match shrinks both channels at once. The second is that nothing outside the enforcement plane was comparing live objects to their declared source, which is why a cluster-wide change went unnoticed. That detection has to be owned by something the enforcement plane does not control, because the enforcement plane is the component nobody is watching: it is the thing that watches.

  • You removed the registration and injections stopped. Why is the incident not over?
    Because two things persist. Running Pods keep the sidecar until they are recreated, so a deliberate roll is still needed. And everything the match selected during the window was delivered to the engine in full, so any matched Secrets and ConfigMaps must be treated as disclosed and rotated. The engine workload itself also still needs investigating.
  • Why not just rely on your validating rules to confirm the cluster is clean?
    They come from the same engine and run after mutation, on the object the mutator already patched. A compromised plane reports itself compliant. The trustworthy check is external: compare live objects against the source that declares them, using tooling the engine does not control.
  • What separates this from one team accidentally adding a sidecar?
    Scope and timing. A team's change appears in that team's repository and lands with a deployment; this appears in namespaces owned by unrelated teams, matches no commit anywhere, and arrives as Pods are recreated. The affected set also lines up exactly with the webhook's match rules.

saying these in an interview costs you the question

  • Assumes the stored object records which webhook changed it
  • Thinks removing the registration reverts running Pods
  • Uses the same engine's rules to verify the cleanup
  • Closes the incident without rotating what the match selected
  • Looks only in team repositories for the source of the container

context