Your Kubernetes platform has had several incidents where a bad rollout or a registry problem left large numbers of pods failing to start. What guardrails would you put in place so that pod startup failures stop becoming outages?
answer
- Readiness-gated rollout + maxUnavailable:0 → failure costs latency, not capacity
- progressDeadlineSeconds → detect stalled rollouts, act on it
- Digest pinning + registry mirror = registry off the critical path
- Admission: require limits and probes, ban :latest
- Alert on rates and duration, grouped by cause
basics
~20 sMake failures non-destructive and detectable: readiness-gated rolling updates with a progress deadline and automatic rollback, digest-pinned images served from a registry mirror, admission rules requiring probes and resource limits, and alerts on restart and pull-failure rates rather than on individual pods.
solid answer
~50 sI separate three questions: does a startup failure remove working capacity, how fast do we know, and how fast can we get back. **Containment.** A rolling update only replaces old pods as new ones become *Ready*, so a crash-looping image should never take capacity away — provided readiness probes are real and `maxUnavailable` is conservative. Set `progressDeadlineSeconds` so a stalled rollout is marked Failed instead of hanging, and wire that to automatic rollback. PodDisruptionBudgets protect against voluntary disruption on top. **Removing shared failure modes.** Pin images by digest so a moved tag cannot change what runs; serve pulls from a mirror or pull-through cache so the public registry is not in the critical path of every node start; authenticate pulls to escape rate limits. **Guardrails at admission.** Require resource limits, require probes, forbid `:latest`. **Detection.** Alert on aggregate signals — restart rate per workload, pods not Ready past a threshold, pull-failure count — because a single pod restarting is noise and a hundred is an incident.
code
yaml · 17 linesspec:
replicas: 6
revisionHistoryLimit: 10
progressDeadlineSeconds: 300
strategy:
type: RollingUpdate
rollingUpdate:
maxUnavailable: 0
maxSurge: 2
template:
spec:
containers:
- name: api
image: registry.internal/team/api@sha256:9f2c1b0d4e6a7c8b5d3f21a0c4e7b9d6f8a1c2e3b4d5f60718293a4b5c6d7e8f
readinessProbe:
httpGet: { path: /ready, port: 8080 }
periodSeconds: 5go deeper
Focus on the basics you would ask for: readiness probes, resource limits, specific image tags, and knowing how to roll back.
Explain rolling-update mechanics (maxUnavailable/maxSurge, readiness gating) and why digest pinning and imagePullPolicy defaults matter.
Connect the mechanisms to real incidents: honest readiness, progressDeadlineSeconds with automated rollback, registry mirrors, and aggregate alerting that groups by cause.
Own the tradeoffs — spare capacity for maxUnavailable: 0, running a mirror, tooling for digests, staged admission enforcement — and tier the guardrails by workload criticality rather than mandating one setting everywhere.
## Frame the problem A pod that will not start is, by itself, harmless — that is the whole point of a declarative rollout. It becomes an outage only through one of three failures of the platform around it: replaced capacity that never came back, a shared dependency that failed for everyone at once, or a delay between the failure and anyone noticing. Guardrails should be organised around those three. ## 1. Never trade healthy capacity for unhealthy capacity A Deployment's rolling update is already the main safeguard: new pods must become Ready before old ones are removed, bounded by `maxUnavailable` and `maxSurge`. This only works if readiness is honest. A readiness probe that returns 200 from a static handler while the app cannot serve traffic converts the rollout into a full outage, because Kubernetes believes the replacement succeeded. So the first guardrail is cultural and enforced: readiness must reflect the ability to serve. Then bound the failure in time. `progressDeadlineSeconds` (default 600) marks a Deployment `Progressing=False` with reason `ProgressDeadlineExceeded` when it cannot make progress. On its own it merely stops lying about the state; the value comes from acting on it — a controller or pipeline step that rolls back, or an alert that pages. `maxUnavailable: 0` for critical services means a broken image costs latency in the rollout, never capacity. For StatefulSets, note the different behaviour: ordered updates stop at the first pod that will not become Ready, which contains the blast radius by default but also means a partial, stuck state you must notice. ## 2. Remove the shared dependencies from the pod-start path The registry is the big one. Every pod start on a node without the image is a network call to a system you may not own. Practical mitigations: - **Digest pinning** (`repo@sha256:...`). A tag is a mutable pointer; two nodes can pull different bytes for the same tag on the same day. Digests make what runs unambiguous and make rollback a real, exact operation. - **A registry mirror or pull-through cache** inside the cluster's network. This turns a public-registry outage or rate limit from a cluster-wide inability to start pods into a cache miss. - **Authenticated pulls**, to escape anonymous rate limits, with credentials provisioned per namespace or via node identity so a single expired secret does not break everything at once. - **Warm nodes.** Pre-pulled base images in the node image, or a DaemonSet that pre-pulls the next release, shortens cold starts and reduces exposure during scale-up — which is exactly when you can least afford pull failures. The same reasoning applies to configuration: if pods cannot start because a secret-sync operator is behind, that operator is on your critical path and needs the same treatment (health checks, alerting, and ideally ordering guarantees in delivery). ## 3. Make bad specs fail early, not at container start Admission policy (ValidatingAdmissionPolicy, or an OPA/Kyverno-style engine) is where a class of startup failures can be eliminated: - Require memory and CPU limits, or provide defaults via `LimitRange`, so a missing limit cannot cause node-level memory pressure. - Require readiness probes on anything fronted by a Service, since rollout safety depends on them. - Forbid mutable tags such as `:latest` in production namespaces. - Require images from approved registries — which also happens to be the ones you mirror. Enforce in a staged way: warn/audit first, gather the violations, then enforce. Policy that blocks deploys on day one gets exceptions carved into it and stops meaning anything. ## 4. Detect in aggregate, not per pod Alerting on "a pod restarted" is unusable — pods restart. The signals that matter are rates and durations: - Container restart rate per workload over a window (the classic `kube_pod_container_status_restarts_total` increase). - Pods not Ready for longer than N minutes, grouped by workload and namespace. - Image pull failure count per node and per registry — a spike across many nodes is a registry or credential incident, not an application one. - Deployments with `ProgressDeadlineExceeded`. Group by cause where possible, so the on-call sees "12 workloads cannot pull from registry X" rather than 400 individual pod alerts. ## 5. Make recovery boring Rollback must be one command or one revert, and must be exercised. If images are digest-pinned and manifests are in version control, the previous known-good state is fully described. Keep enough revision history (`revisionHistoryLimit`) to roll back, and make sure the rollback path does not itself depend on the thing that broke — a rollback that requires pulling an image from an unavailable registry is not a rollback. ## The tradeoffs to acknowledge Every guardrail costs something. `maxUnavailable: 0` needs spare capacity. Digest pinning requires tooling to update references. Mirrors are infrastructure to run and to keep in sync. Strict admission policy adds friction and an exception process. The judgement is about which workloads deserve which level: a payment service and a batch report do not need the same protection, and applying the strictest setting everywhere generally produces workarounds rather than reliability.
- Why does a rolling update normally protect you from a crash-looping image, and when does that protection fail?Because new pods must report Ready before old ones are terminated, so replicas that never start simply never replace anything and the old version keeps serving. The protection fails when readiness is not truthful — a probe that returns 200 unconditionally, or no readiness probe at all, so Kubernetes counts a broken pod as available and proceeds to delete healthy ones. It also fails if maxUnavailable is set high, which permits removing capacity before replacements are proven.
- What does pinning images by digest rather than by tag actually buy you operationally?Determinism. A tag is a mutable pointer, so two nodes pulling the same tag at different times can run different bytes, and imagePullPolicy: IfNotPresent will happily keep a stale cached copy. A digest names exact content, so what is deployed is unambiguous, rollback restores exact bits rather than whatever the tag points at now, and an accidental or malicious tag move cannot change running workloads. The cost is tooling to keep digest references updated.
saying these in an interview costs you the question
- Treating a crash-looping rollout as inevitable capacity loss, rather than asking why readiness let old pods be removed
- Alerting on individual pod restarts, producing noise that hides real incidents
- Relying on mutable tags in production and then being unable to say what is actually running
- Assuming the public registry is always available, with no mirror or pre-pulled images
- Applying the strictest admission policy to every workload uniformly and driving teams to exceptions