skip to content

Troubleshooting

Reading a broken cluster from the outside in: describe output and Events, pods that never start or stay Pending, evictions under node pressure, Services that never answer, and an API server that stops replying. Senior interviews are one incident walk-through.

part ofKubernetesoverview, primer and where to startread it →
on this pageshow

questions

page 2 of 2

During a Kubernetes rollout of a 7-replica Deployment on a 210-node cluster, four new pods run but three sit in ImagePullBackOff. How do you find why only some nodes fail?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Map the failing pods to their nodes, then read each pod's Failed event for the actual pull error. The error class (not found, 401/403, 429, network or TLS) combined with what differs about those nodes (cache, egress IP, pool, credentials) explains the split.

open as a page

Where do the numbers reported by kubectl top pods and kubectl top nodes come from, and what should you avoid concluding from them?

level: seniorimportance: should knowfreq 42%

basics

~20 s

They come from metrics-server, an optional add-on that scrapes each kubelet every ~15 seconds and keeps only the latest sample in memory. CPU is an average rate over that window and memory is working-set bytes — so no history, no spikes, and not the same as your app's heap.

open as a page

A Kubernetes pod failed overnight, and by morning kubectl get events shows nothing about it. How long do Kubernetes Events survive, why, and how do you keep that evidence?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Events are ordinary API objects stored in etcd with a lease set by kube-apiserver's --event-ttl, default one hour, renewed whenever the Event is updated. After that they vanish, so keeping them means exporting them continuously to a durable store.

open as a page

Pods on a Kubernetes node are being killed with the message "The node was low on resource: ephemeral-storage". What counts toward a container's ephemeral storage usage, and how would you stop this from recurring?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Ephemeral storage is the container writable layer, emptyDir volumes, and container logs on the node's disk. The kubelet evicts pods when nodefs/imagefs run low, or when a pod exceeds its ephemeral-storage limit. Fix with log rotation, sized emptyDir, ephemeral-storage requests/limits, and image garbage collection.

open as a page

A Kubernetes node goes `NotReady`, and after roughly five minutes its pods are deleted and recreated elsewhere. Explain the mechanism that does this, and how it differs from a kubelet evicting pods because the node is short on memory.

level: seniorimportance: should knowfreq 42%

basics

~20 s

The node controller adds a NoExecute taint (node.kubernetes.io/not-ready or unreachable). Pods carry a default toleration with tolerationSeconds 300, so after 5 minutes the taint manager deletes them and controllers recreate them elsewhere. Kubelet node-pressure eviction is different: local, resource-driven, ranked by QoS.

open as a page

Distinguish two Kubernetes failures: a pod Pending with `pod has unbound immediate PersistentVolumeClaims`, versus a Deployment whose replica count never rises and for which no pod object is ever created. What causes each, and how do you confirm it?

level: seniorimportance: should knowfreq 48%

basics

~20 s

An unbound PVC means storage was never provisioned — check the PVC events and StorageClass. A missing pod entirely means admission rejected creation: describe the ReplicaSet for FailedCreate naming a ResourceQuota, LimitRange, or webhook. One is a scheduling problem; the other happens before scheduling.

open as a page

A Deployment scales from 3 to 6 replicas and the new pods stay Pending with the scheduler message `node(s) didn't match pod topology spread constraints`. Explain what that constraint does and how you would decide between relaxing it and adding capacity.

level: seniorimportance: should knowfreq 45%

basics

~20 s

A topologySpreadConstraint limits how uneven pod placement may be across a topology key (zone, node): the count difference between domains must stay within maxSkew. With whenUnsatisfiable: DoNotSchedule, a pod that would break the skew stays Pending. Relax to ScheduleAnyway, raise maxSkew, or add capacity in the short domains.

open as a page

A container starts normally, serves traffic for about a minute, is then killed and restarted, and after several cycles the pod reports CrashLoopBackOff — yet the application logs show no error before each death. How do you investigate?

level: seniorimportance: should knowfreq 54%

basics

~20 s

A silent kill after a healthy period usually means the kubelet is killing the container because its liveness probe failed. Check kubectl describe pod for Unhealthy events and the probe's path, port, timeout and thresholds, then call the same endpoint from inside the pod.

open as a page

Calls through a Kubernetes Service succeed most of the time, but a small percentage fail with connection resets, and the rate spikes during deployments. What causes this intermittent pattern and how would you fix it?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Endpoint updates are eventually consistent: a terminating pod stops accepting connections before every node's kube-proxy has removed it, so a fraction of requests are sent to a dead backend. Fix with a preStop sleep, graceful shutdown, and readiness flipping before termination.

open as a page

Pods in one Kubernetes namespace suddenly cannot reach a Service in another namespace, and the failure presents as a hang rather than a connection refusal. How do you determine whether a NetworkPolicy is responsible, and what mistakes commonly cause this?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Silent drops point at NetworkPolicy. List policies in both namespaces and see which select the client and server pods: once any policy selects a pod for a direction, everything not explicitly allowed in that direction is denied. Check ingress and egress separately, and confirm the CNI enforces policy at all.

open as a page

A Kubernetes search-autocomplete service shows p99 doubling and 24% of CFS periods throttled during an 18-hour certificate-expiry window of TLS re-handshakes; how do you choose between resizing thread pools, raising the CPU limit, or removing it?

level: seniorimportance: should knowfreq 47%

basics

~20 s

First match the service's parallelism to its CPU limit, since most throttling below the limit is over-parallelism. Raise the limit when real demand exceeds it. Remove it only with an accurate request and an accepted loss of Guaranteed QoS.

open as a page

Your production Kubernetes workloads run distroless images as non-root with read-only root filesystems, and on-call engineers complain they cannot debug incidents. How would you give them a debugging capability without undoing that hardening?

level: principalimportance: should knowfreq 26%

basics

~20 s

Keep images minimal and move debugging out of them: an approved toolbox image attached as an ephemeral container with kubectl debug --target, granted through a time-boxed break-glass role, audited, and paired with in-app diagnostics so most incidents need no shell at all.

open as a page

How does kubectl cp move files between your machine and a container, and why does it sometimes fail with an error mentioning tar?

level: middleimportance: nice to knowfreq 30%

basics

~20 s

kubectl cp is exec plus tar: it runs tar inside the container and streams the archive over the exec channel. If the image has no tar binary — distroless or scratch — the copy fails. Fall back to exec with cat, or a debug container.

open as a page

In Kubernetes, kubectl top pod shows 1870Mi against a 2Gi memory limit with no OOMKills; how do working set and RSS differ?

level: middleimportance: nice to knowfreq 33%

basics

~20 s

Working set is the cgroup's memory usage minus inactive file cache, so it still counts active page cache the kernel can reclaim. RSS counts anonymous memory that cannot be reclaimed without swap. RSS near the limit predicts an OOM kill.

open as a page

After Kubernetes nodes are drained, every pod creation fails with 'failed calling webhook', including the webhook's own pods. Why does this deadlock happen, and how do you break it?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

A webhook with failurePolicy Fail intercepts pod creation, and its backend has no running pods, so every create is rejected, including creates of its replacement pods. Break the loop by patching the webhook configuration to Ignore or narrowing its scope.

open as a page

Your Kubernetes platform has had several incidents where a bad rollout or a registry problem left large numbers of pods failing to start. What guardrails would you put in place so that pod startup failures stop becoming outages?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Make failures non-destructive and detectable: readiness-gated rolling updates with a progress deadline and automatic rollback, digest-pinned images served from a registry mirror, admission rules requiring probes and resource limits, and alerts on restart and pull-failure rates rather than on individual pods.

open as a page

showing 31–46 of 46