Troubleshooting
Reading a broken cluster from the outside in: describe output and Events, pods that never start or stay Pending, evictions under node pressure, Services that never answer, and an API server that stops replying. Senior interviews are one incident walk-through.
part ofKubernetesoverview, primer and where to startread it →on this pageshowhide
explore
- kubectl Triage Workflow5 questions
- CrashLoopBackOff and ImagePullBackOff6 questions
- Pending and Unschedulable Pods5 questions
- Node Pressure and Evictions5 questions
- Debugging Running Containers6 questions
- Service Connectivity Debugging5 questions
- Image Pull Failures3 questions
- Logs and Events3 questions
- CPU Throttling and Latency4 questions
- Control-Plane Failures4 questions
questions
page 2 of 2During a Kubernetes rollout of a 7-replica Deployment on a 210-node cluster, four new pods run but three sit in ImagePullBackOff. How do you find why only some nodes fail?
basics
~20 sMap the failing pods to their nodes, then read each pod's Failed event for the actual pull error. The error class (not found, 401/403, 429, network or TLS) combined with what differs about those nodes (cache, egress IP, pool, credentials) explains the split.
Where do the numbers reported by kubectl top pods and kubectl top nodes come from, and what should you avoid concluding from them?
basics
~20 sThey come from metrics-server, an optional add-on that scrapes each kubelet every ~15 seconds and keeps only the latest sample in memory. CPU is an average rate over that window and memory is working-set bytes — so no history, no spikes, and not the same as your app's heap.
A Kubernetes pod failed overnight, and by morning kubectl get events shows nothing about it. How long do Kubernetes Events survive, why, and how do you keep that evidence?
basics
~20 sEvents are ordinary API objects stored in etcd with a lease set by kube-apiserver's --event-ttl, default one hour, renewed whenever the Event is updated. After that they vanish, so keeping them means exporting them continuously to a durable store.
Pods on a Kubernetes node are being killed with the message "The node was low on resource: ephemeral-storage". What counts toward a container's ephemeral storage usage, and how would you stop this from recurring?
basics
~20 sEphemeral storage is the container writable layer, emptyDir volumes, and container logs on the node's disk. The kubelet evicts pods when nodefs/imagefs run low, or when a pod exceeds its ephemeral-storage limit. Fix with log rotation, sized emptyDir, ephemeral-storage requests/limits, and image garbage collection.
A Kubernetes node goes `NotReady`, and after roughly five minutes its pods are deleted and recreated elsewhere. Explain the mechanism that does this, and how it differs from a kubelet evicting pods because the node is short on memory.
basics
~20 sThe node controller adds a NoExecute taint (node.kubernetes.io/not-ready or unreachable). Pods carry a default toleration with tolerationSeconds 300, so after 5 minutes the taint manager deletes them and controllers recreate them elsewhere. Kubelet node-pressure eviction is different: local, resource-driven, ranked by QoS.
Distinguish two Kubernetes failures: a pod Pending with `pod has unbound immediate PersistentVolumeClaims`, versus a Deployment whose replica count never rises and for which no pod object is ever created. What causes each, and how do you confirm it?
basics
~20 sAn unbound PVC means storage was never provisioned — check the PVC events and StorageClass. A missing pod entirely means admission rejected creation: describe the ReplicaSet for FailedCreate naming a ResourceQuota, LimitRange, or webhook. One is a scheduling problem; the other happens before scheduling.
A Deployment scales from 3 to 6 replicas and the new pods stay Pending with the scheduler message `node(s) didn't match pod topology spread constraints`. Explain what that constraint does and how you would decide between relaxing it and adding capacity.
basics
~20 sA topologySpreadConstraint limits how uneven pod placement may be across a topology key (zone, node): the count difference between domains must stay within maxSkew. With whenUnsatisfiable: DoNotSchedule, a pod that would break the skew stays Pending. Relax to ScheduleAnyway, raise maxSkew, or add capacity in the short domains.
A container starts normally, serves traffic for about a minute, is then killed and restarted, and after several cycles the pod reports CrashLoopBackOff — yet the application logs show no error before each death. How do you investigate?
basics
~20 sA silent kill after a healthy period usually means the kubelet is killing the container because its liveness probe failed. Check kubectl describe pod for Unhealthy events and the probe's path, port, timeout and thresholds, then call the same endpoint from inside the pod.
Calls through a Kubernetes Service succeed most of the time, but a small percentage fail with connection resets, and the rate spikes during deployments. What causes this intermittent pattern and how would you fix it?
basics
~20 sEndpoint updates are eventually consistent: a terminating pod stops accepting connections before every node's kube-proxy has removed it, so a fraction of requests are sent to a dead backend. Fix with a preStop sleep, graceful shutdown, and readiness flipping before termination.
Pods in one Kubernetes namespace suddenly cannot reach a Service in another namespace, and the failure presents as a hang rather than a connection refusal. How do you determine whether a NetworkPolicy is responsible, and what mistakes commonly cause this?
basics
~20 sSilent drops point at NetworkPolicy. List policies in both namespaces and see which select the client and server pods: once any policy selects a pod for a direction, everything not explicitly allowed in that direction is denied. Check ingress and egress separately, and confirm the CNI enforces policy at all.
A Kubernetes search-autocomplete service shows p99 doubling and 24% of CFS periods throttled during an 18-hour certificate-expiry window of TLS re-handshakes; how do you choose between resizing thread pools, raising the CPU limit, or removing it?
basics
~20 sFirst match the service's parallelism to its CPU limit, since most throttling below the limit is over-parallelism. Raise the limit when real demand exceeds it. Remove it only with an accurate request and an accepted loss of Guaranteed QoS.
Your production Kubernetes workloads run distroless images as non-root with read-only root filesystems, and on-call engineers complain they cannot debug incidents. How would you give them a debugging capability without undoing that hardening?
basics
~20 sKeep images minimal and move debugging out of them: an approved toolbox image attached as an ephemeral container with kubectl debug --target, granted through a time-boxed break-glass role, audited, and paired with in-app diagnostics so most incidents need no shell at all.
How does kubectl cp move files between your machine and a container, and why does it sometimes fail with an error mentioning tar?
basics
~20 skubectl cp is exec plus tar: it runs tar inside the container and streams the archive over the exec channel. If the image has no tar binary — distroless or scratch — the copy fails. Fall back to exec with cat, or a debug container.
In Kubernetes, kubectl top pod shows 1870Mi against a 2Gi memory limit with no OOMKills; how do working set and RSS differ?
basics
~20 sWorking set is the cgroup's memory usage minus inactive file cache, so it still counts active page cache the kernel can reclaim. RSS counts anonymous memory that cannot be reclaimed without swap. RSS near the limit predicts an OOM kill.
After Kubernetes nodes are drained, every pod creation fails with 'failed calling webhook', including the webhook's own pods. Why does this deadlock happen, and how do you break it?
basics
~20 sA webhook with failurePolicy Fail intercepts pod creation, and its backend has no running pods, so every create is rejected, including creates of its replacement pods. Break the loop by patching the webhook configuration to Ignore or narrowing its scope.
Your Kubernetes platform has had several incidents where a bad rollout or a registry problem left large numbers of pods failing to start. What guardrails would you put in place so that pod startup failures stop becoming outages?
basics
~20 sMake failures non-destructive and detectable: readiness-gated rolling updates with a progress deadline and automatic rollback, digest-pinned images served from a registry mirror, admission rules requiring probes and resource limits, and alerts on restart and pull-failure rates rather than on individual pods.
showing 31–46 of 46