skip to content

You must introduce a cluster-wide rule that rejects Pods running as UID 0 across dozens of teams' namespaces in a running cluster. How do you roll that out without breaking production workloads, and how do you handle the teams that legitimately cannot comply?

level: principalimportance: should knowfreq 28%

answer

  1. audit/warn first, deny later — never straight to deny
  2. admission ≠ running fleet: background scan for the real inventory
  3. one policy, many bindings scoped by namespace label
  4. rejections surface on ReplicaSet events, not kubectl apply
  5. exceptions: owner + justification + expiry, never a global Ignore

basics

~20 s

Never start at deny. Ship the rule in audit/warn mode first, measure violations per namespace from the reports, give owners a deadline and help, then enforce namespace by namespace starting with low-risk ones. Handle exceptions as explicit, labelled, time-boxed, individually-reviewed exclusions — not by weakening the global rule.

solid answer

~50 s

Treat it as a migration, not a config change. 1. **Observe first.** Deploy with `validationActions: [Warn, Audit]` (or Gatekeeper `enforcementAction: dryrun`, Kyverno `Audit`). Admission only sees new writes, so also run the engine's background scan to inventory *running* pods. Produce a per-namespace violation list with owners. 2. **Communicate with a deadline.** Publish the rule, the report, the remediation recipe, and the enforcement date. Most fixes are a `securityContext` addition or a base image that no longer needs root. 3. **Enforce progressively.** Bindings are the lever: one policy, many bindings scoped by namespace label. Flip `env: sandbox` first, then non-critical prod, then the rest. Watch rejection metrics after each step. 4. **Design the escape hatch.** Exceptions are labelled namespaces or a PolicyException object with an owner, a justification, and an expiry, reviewed like any other risk acceptance. Never a global `Ignore`. Also remember rejections surface on the ReplicaSet/Deployment status, not in `kubectl apply` output — brief teams, and monitor controller events.

code

bash · 9 lines
bash
kubectl get pods -A -o json \
  | jq -r '
      .items[]
      | select(
          ((.spec.securityContext.runAsNonRoot // false) != true)
          and (any(.spec.containers[]; (.securityContext.runAsNonRoot // false) != true))
        )
      | "\(.metadata.namespace)\t\(.metadata.name)"' \
  | sort | uniq -c | sort -rn

go deeper

for a junior

Know that policy engines have a non-blocking audit or warn mode and that you use it before turning on rejection.

for a middle

Describe the sequence: audit/warn, measure violations, notify owners with a deadline, then enforce — and note that admission does not re-check already-running pods.

for a senior

Add the operational detail: per-namespace binding rollout, rejection and ReplicaSet-event monitoring, remediation recipes, and controlled exceptions with expiry.

for a principal

Own the trade-off explicitly — the coverage gap you accept during the ramp, the compensating background scan, who may grant exceptions, how enforcement is proven still on afterwards, and the date the exception list must be empty.

## Why this is a rollout problem, not a policy problem Writing the rule is trivial: reject a pod whose `securityContext.runAsUser` is 0 or unset with an image that defaults to root. The hard part is that a running cluster is full of workloads that were built before the rule existed, owned by teams with their own release schedules, and the enforcement point — admission — rejects writes at a moment when nobody is watching a terminal. ## Stage 1: measure before you enforce Every serious policy tool has a non-blocking mode, and it exists exactly for this: - Built-in policies: `validationActions: [Warn, Audit]` — `Warn` returns a `Warning:` header that `kubectl` prints to whoever made the change, `Audit` records an annotation in the API audit log. - Gatekeeper: `enforcementAction: dryrun` (or `warn`), with violations also surfaced on constraint status by the audit loop. - Kyverno: `validationFailureAction: Audit`, with PolicyReport objects per namespace. Critically, admission control only sees **new** writes. A pod created last month never passes through your policy again until something recreates it. So a second inventory pass is mandatory: the engine's background scan, or a script over all pods, to find the running fleet's violations. Only that gives you a true denominator. Output of this stage should be a table: namespace, owner, workload, why it violates. That table is the project plan. ## Stage 2: make compliance cheap Enforcement without remediation support just converts a security goal into a political fight. Provide the exact diff teams need: `runAsNonRoot: true`, `runAsUser: <uid>`, and for images that hardcode root, the base-image change or the `USER` directive their image needs. Where a workload binds a privileged port, the modern answer is to bind a high port and let the Service map it, not to keep root. Warn mode helps here: developers see the warning on every apply, which is a far better nudge than an email. ## Stage 3: enforce progressively The policy/binding split exists for this. One policy object, several bindings scoped by namespace label: - `env: sandbox` → `Deny` first, - `tier: non-critical` next, - everything else last. After each step watch admission rejection metrics and controller events for a settling period. If a wave produces surprises, the blast radius is one label's worth of namespaces and rolling back is a one-line binding edit. An important operational subtlety: for workloads managed by controllers, the rejection does not appear where people look. `kubectl apply` of a Deployment succeeds — the Deployment object is fine. The ReplicaSet controller then fails to create pods, and the error appears in the ReplicaSet's status/events and as a stalled rollout. Teams that don't know this report it as "my deploy hangs". Brief them, and alert on ReplicaSet failure events during the rollout. ## Stage 4: institutionalise the exception Some workloads genuinely cannot comply on your timeline — a vendor appliance, a node agent that must write to host paths, a legacy image nobody can rebuild. The wrong responses are: weakening the rule for everyone, switching the webhook to `Ignore`, or an undocumented namespace carve-out that outlives its reason. The right shape is an explicit exception object: which workload, which rule, who owns it, why, and an **expiry date**. Gatekeeper and Kyverno both support scoped exclusions (Kyverno's `PolicyException`, Gatekeeper's `excludedNamespaces`/match exclusions); with built-in policies you express it as a namespace/object label the binding excludes, or an `authorizer` check in the CEL so the exemption is granted through RBAC and therefore audited. Review the exception list on a schedule; an exception with no expiry is a permanently disabled control. ## Stage 5: prove it stays on After enforcement, the failure mode flips from "breaks things" to "silently stops working" — a webhook set to `Ignore` during an incident and never restored, a binding accidentally scoped away, a namespace label that grants exemption applied by someone who could. Defend with: a canary object in CI that the policy must reject (alert if it is admitted), periodic background scans reporting non-conforming *running* pods, and alerting on changes to the policy, binding, and exception objects themselves. Also restrict who may set the exemption label — otherwise the exemption is self-service and the policy is advisory. ## The judgement call The decision worth articulating is where you sit between speed and safety: a hard cutover buys immediate coverage at the risk of an outage and organisational backlash; a slow, per-namespace ramp buys safety but leaves a long window where the control is partial. State the window you are accepting, what compensating control covers it (background scan and remediation SLA), and the date the exception list must be empty.

  • A team says their Deployment stopped rolling out after enforcement, but kubectl apply reported success. Explain what they are seeing.
    The Deployment object itself is valid, so its write is admitted; the policy applies to Pods. The ReplicaSet controller then tries to create pods and each creation is rejected by admission, so the rollout stalls with no new pods. The error lives in the ReplicaSet's status conditions and events (kubectl describe rs / kubectl get events), not in the apply output. This indirection is why controller-created workloads need explicit briefing and event-based alerting during a policy rollout.
  • Why is an exemption implemented as a namespace label risky, and how do you tighten it?
    If teams can edit their own namespace labels, the exemption is self-service and the policy becomes advisory. Tighten it by restricting who may patch namespace objects via RBAC, or by moving the exemption to an object only the security team can create (a PolicyException, or an authorizer-based check in CEL so the exemption is a granted RBAC permission). Either way attach an owner and an expiry and alert on creation or modification of those objects.

It's a building-code retrofit, not a new-build spec: you inspect every existing structure, publish a compliance deadline, retrofit floor by floor, and issue named, dated variances for the few that can't be brought up — you don't repeal the code.

saying these in an interview costs you the question

  • Enabling deny cluster-wide immediately and treating the resulting breakage as the teams' problem.
  • Assuming admission enforcement makes the cluster compliant, forgetting that already-running pods were never re-evaluated.
  • Granting exemptions with no owner and no expiry, which permanently disables the control for those workloads.
  • Relaxing failurePolicy to Ignore during an incident and never restoring it — enforcement then stays silently off.
  • Expecting rejections to appear in kubectl apply output for controller-managed workloads.

context