You are asked to move existing production namespaces to the Kubernetes restricted Pod Security Standard without breaking running workloads. How do you sequence that migration, and how do you diagnose workloads that stop being admitted?
answer
- warn+audit first, enforce last
- observe a full cycle — CronJobs and DaemonSets hide
- kubectl label --dry-run=server = go/no-go
- pin enforce-version so upgrades don't tighten silently
- errors live on the ReplicaSet/Job events, not on apply
basics
~20 sTurn on warn and audit at restricted first, collect violations across a full deploy cycle including CronJobs, fix manifests (non-root UID, drop ALL, no escalation, seccomp, emptyDir for writable paths), verify with server-side dry-run, then set enforce per namespace with a pinned version. Diagnose rejections from ReplicaSet/Job events, which name the violated fields.
solid answer
~60 s**Sequence** 1. Label every namespace `warn` + `audit` = `restricted`, leaving `enforce` alone. Nothing breaks; you now have both a developer-facing signal and a fleet-wide log signal. 2. Harvest violations for at least one full release cycle so CronJobs, Jobs and rarely-restarted Deployments show up. Aggregate the audit annotations by namespace and owner. 3. Fix manifests. The recurring set is: numeric non-root `runAsUser` + `runAsNonRoot`, `capabilities.drop: ["ALL"]`, `allowPrivilegeEscalation: false`, `seccompProfile: RuntimeDefault`, and `emptyDir` mounts for whatever the read-only root broke. Sidecars and init containers too. 4. Verify per namespace with `kubectl label --dry-run=server --overwrite ns X pod-security.kubernetes.io/enforce=restricted` — it reports which live pods would violate. 5. Enforce namespace by namespace, least critical first, pinning `enforce-version` so a cluster upgrade cannot tighten policy under you. Give genuine node-level infrastructure its own namespace at `privileged`, with a named owner. **Diagnosis** — enforce rejects *pods*, so `kubectl apply` looks fine. Look at `kubectl describe rs/<name>` or `kubectl get events`; the message names the exact violated fields. Then check runtime failures separately: `CreateContainerConfigError` is the kubelet's non-root check, `Permission denied` is a read-only root, `Operation not permitted` is a dropped capability or seccomp.
code
bash · 17 lines# 1. observe
kubectl label --overwrite ns orders \
pod-security.kubernetes.io/warn=restricted \
pod-security.kubernetes.io/audit=restricted
# 2. go/no-go check against live pods, without persisting
kubectl label --dry-run=server --overwrite ns orders \
pod-security.kubernetes.io/enforce=restricted
# 3. enforce, pinned
kubectl label --overwrite ns orders \
pod-security.kubernetes.io/enforce=restricted \
pod-security.kubernetes.io/enforce-version=v1.31
# 4. when pods vanish, the reason is on the controller, not the Deployment
kubectl describe rs -n orders -l app=orders | sed -n '/Events/,$p'
kubectl get events -n orders --field-selector reason=FailedCreatego deeper
Know the safe order — warn and audit before enforce — and that the rejection message appears on the ReplicaSet events, not on kubectl apply.
Add the concrete manifest patch and the mapping from symptom to cause: CreateContainerConfigError, Permission denied, Operation not permitted.
Own the whole campaign: measurement window covering CronJobs and DaemonSets, server-side dry-run gates, per-namespace rollout with rollback, version pinning, and a named exception list.
Frame it as platform policy: a cluster-wide default level, an exception process with owners and expiry, upgrade choreography around pinned versions, and where PSA's limits mean you need a broader policy layer.
## What makes this a migration rather than a flag flip The restricted Pod Security Standard differs from baseline in kind: baseline forbids dangerous fields, restricted *requires* hardening fields. That means almost every pre-existing workload fails restricted on day one, not because it does anything dangerous but because it never wrote `allowPrivilegeEscalation: false`. So the migration is a manifest-editing campaign with a policy switch at the end, and the risk is entirely in workloads you did not think to test. ## Phase 1 — observe without breaking Apply `pod-security.kubernetes.io/warn=restricted` and `pod-security.kubernetes.io/audit=restricted` to every namespace. Neither blocks anything. - **warn** returns a warning on the API response and, importantly, evaluates workload resources (Deployment, StatefulSet, DaemonSet, Job, CronJob) — so a developer running `kubectl apply` or a CI pipeline sees it at the moment they change something. - **audit** writes an annotation on the API server audit log entry, giving you a queryable fleet-wide inventory even for workloads nobody has touched in months. Also set the cluster-wide PSA admission-configuration default so newly created namespaces inherit the target level instead of being unrestricted. ## Phase 2 — measure honestly The trap is sampling only what redeploys. Run the observation window long enough to cover: - CronJobs with weekly or monthly schedules - Jobs triggered by incidents or batch runs - DaemonSets that only reschedule on node turnover - Anything installed by a Helm chart you do not own — third-party charts are the most common source of unfixable violations, and you need to know early whether the chart exposes a `securityContext` value or must be forked/replaced An alternative to waiting is to enumerate statically: pull every pod template from the API and evaluate it offline against the restricted rules. ## Phase 3 — fix the manifests The standard patch, applied at pod level where possible and repeated for every container including init and sidecars: - `runAsNonRoot: true` plus a numeric `runAsUser` (a named `USER` in the image is not resolvable by the kubelet's check) - `allowPrivilegeEscalation: false` - `capabilities: { drop: ["ALL"] }`, adding back only `NET_BIND_SERVICE` when a low port is unavoidable - `seccompProfile: { type: RuntimeDefault }` - volume types restricted to the safe list; replace `hostPath` with a PVC, a projected volume or a node-level agent that legitimately lives in a privileged namespace Expect collateral work: images that write into the image layer, apps that bind port 80, JVMs whose temp dir is unset, files in the image owned by root and unreadable by the new UID. `fsGroup` fixes volume ownership; a `chown` in the Dockerfile fixes image paths. ## Phase 4 — verify before enforcing `kubectl label --dry-run=server --overwrite ns <ns> pod-security.kubernetes.io/enforce=restricted` runs the policy against the namespace's existing pods and reports the violators without persisting the label. This is the go/no-go check. Do it per namespace immediately before flipping. Pin the version: `pod-security.kubernetes.io/enforce-version: v1.31`. Without a pin the level tracks `latest`, so a control-plane upgrade can introduce a new requirement that starts rejecting pods during an unrelated maintenance window. ## Phase 5 — enforce progressively Flip namespace by namespace, least critical first, with a rollback (remove the enforce label) rehearsed. Keep `warn` and `audit` set at the same level afterwards so drift is still visible. Infrastructure namespaces that genuinely need host access get an explicit `enforce: privileged` with an owner and a justification recorded — an explicit exception you can review beats an unlabelled namespace you forget. ## Diagnosing failures Split failures into two layers: **Admission-time (API server rejects the pod).** Because enforce acts on pods, the object you applied usually succeeds and the failure hides one level down. Check `kubectl describe replicaset <rs>`, `kubectl describe job`, or `kubectl get events --field-selector reason=FailedCreate`. The message enumerates the violated policy fields verbatim — `allowPrivilegeEscalation != false`, `unrestricted capabilities`, `runAsNonRoot != true`, `seccompProfile` — so it is a checklist, not a mystery. **Runtime (pod admitted, container will not run).** - `CreateContainerConfigError` with a root message → the kubelet's `runAsNonRoot` check; the image's user is UID 0 or an unresolvable name. - `Permission denied` on a write → `readOnlyRootFilesystem` and a missing writable mount. - `Operation not permitted` on a syscall → a dropped capability or the seccomp profile; identify it by running once with a permissive profile and tracing. - CrashLoopBackOff right after the change with a bind error → the app is trying to bind a port below 1024 without `NET_BIND_SERVICE`. ## Signals of success Zero audit annotations at the target level, every namespace explicitly labelled (no implicit defaults), a short reviewed list of privileged exceptions, and enforcement pinned to a version you upgrade deliberately.
- A third-party Helm chart you depend on cannot pass restricted. What are your options?First check whether the chart exposes podSecurityContext and securityContext values — most maintained charts do, and the fix is values-only. If not, you can post-render or patch the rendered manifests, fork the chart, or isolate it in its own namespace enforced at baseline with a recorded owner and an expiry date on the exception. What you should not do is relax the whole cluster to the weakest chart's level.
- How do you stop new namespaces from being created without any Pod Security labels?Set a cluster-wide default in the Pod Security Admission configuration file passed to the API server, which applies a level to namespaces that carry no labels. That way an unlabelled namespace inherits your intended floor rather than being unrestricted. Complement it with a check in whatever provisions namespaces — the GitOps repo or platform tooling — so the labels are part of the namespace template.
- Why pin enforce-version instead of letting it track latest?The policy definitions evolve with Kubernetes minor versions, so 'latest' means the effective rules change when the control plane is upgraded. That can start rejecting previously admitted pods during an upgrade window, mixing a policy incident into an infrastructure change. Pinning makes tightening a deliberate, separately scheduled action; you bump the pin after re-running warn and audit at the new version.
saying these in an interview costs you the question
- Flipping enforce cluster-wide in one step because 'the manifests look fine'
- Only testing Deployments and getting surprised by CronJobs and DaemonSets later
- Looking for the rejection message on the Deployment instead of the ReplicaSet or Job
- Granting a privileged label to a whole application namespace to unblock one sidecar
- Treating restricted as satisfied once pods are non-root, ignoring capabilities and seccomp