skip to content

Troubleshooting

Reading a broken cluster from the outside in: describe output and Events, pods that never start or stay Pending, evictions under node pressure, Services that never answer, and an API server that stops replying. Senior interviews are one incident walk-through.

part ofKubernetesoverview, primer and where to startread it →
on this pageshow

questions

page 1 of 2

When a Kubernetes cluster's API server becomes unreachable, what happens to pods that are already running, and what stops working?

level: juniorimportance: must knowfreq 60%

answer

  1. control plane vs data plane
  2. kubelet works from local state
  3. kube-proxy rules stay in the kernel
  4. no scheduling, scaling or endpoint updates
  5. missed CronJob and startingDeadlineSeconds

basics

~20 s

Running pods keep running and existing Service traffic keeps flowing, because kubelets and kube-proxy work from state they already have. Anything that needs the API server stops: scheduling, rescheduling, scaling, rollouts, endpoint updates and CronJob runs.

solid answer

~40 s

The data plane mostly keeps going. Each kubelet keeps running the containers it already has and still restarts crashed ones according to `restartPolicy`. kube-proxy keeps the Service rules it last programmed, so traffic to existing pods keeps flowing. What stops is anything that needs the API server. The scheduler cannot place pods. kube-controller-manager cannot replace pods from a dead node, scale a Deployment or advance a rollout. EndpointSlices stop updating, so a pod that dies can stay in the routing rules. `kubectl` gets no answer at all. A CronJob such as a nightly ledger-reconciliation batch does not start while the control plane is down. So a control-plane outage usually shows up first as "nothing changes", not as "everything is down", and it gets riskier the longer it lasts.

go deeper

for a junior

Remember the split: pods and Service routing already in place keep working, while scheduling, scaling, rollouts and kubectl stop until the API server is back.

for a middle

Explain why: the kubelet and kube-proxy act on state they already hold, and every controller is a loop that needs API reads and writes to act.

for a senior

Name the risks that grow with time: stale EndpointSlices sending traffic to dead pods, failed nodes not replaced, missed CronJobs, and the burst of load when everything reconnects.

for a principal

Treat control-plane availability as a limit on how long the cluster can safely run unattended, and design batch schedules, deadlines and recovery load with that in mind.

## Why the cluster does not stop Kubernetes separates the **control plane** from the **data plane**: - The **control plane** is `kube-apiserver`, etcd (the cluster store), `kube-scheduler` and `kube-controller-manager`. It decides what should happen. - The **data plane** is every node's **kubelet**, container runtime and `kube-proxy`, plus the pods themselves. It carries out the decisions. The kubelet does not ask the API server's permission for each action. It **watches** for pod assignments, keeps a local copy of what it was told to run, and drives the container runtime to match. If the watch breaks, the kubelet keeps the containers it already has. It still runs liveness probes and still restarts a crashed container according to the pod's `restartPolicy`, because both are local decisions. `kube-proxy` has already written iptables, nftables or IPVS rules into the node's kernel, and those rules keep forwarding traffic without any connection to the API server. So in the first minutes of an outage, users of already-running services often notice nothing. ## What stops working Everything that needs a **write to or read from the API server** stops: | Function | Owner | Effect while the API server is down | |---|---|---| | Placing new pods | kube-scheduler | Pending pods stay Pending | | Replacing pods from a failed node | kube-controller-manager | Lost capacity is not replaced | | Scaling and rollouts | Deployment and HPA controllers | Replica counts freeze | | Updating EndpointSlices | EndpointSlice controller | A pod that dies can stay in the routing rules | | Starting CronJobs | CronJob controller | Scheduled runs are missed | | Reporting status | kubelet | Pod and node status in the API goes stale | | Anything a human does | kubectl | No answers and no changes | A few of these need a closer look: 1. **Stale routing.** If a pod crashes for good, or its node dies, the EndpointSlice is not updated. kube-proxy on the other nodes keeps sending that pod a share of the connections. 2. **Missed batch work.** Say a nightly ledger-reconciliation CronJob is due at 02:00 and the control plane is down from 01:40 to 02:25. That run does not start on time. When the control plane comes back, the CronJob controller can still start the missed run, but only if the CronJob's `startingDeadlineSeconds` has not passed (or is unset). The run can therefore start late, or be skipped entirely. 3. **No self-healing beyond the node.** The kubelet restarts containers, but a pod whose node dies is not recreated anywhere else. 4. **Cluster DNS.** CoreDNS typically keeps answering from its last-synced view of Services, so existing names keep resolving. New Services do not appear. ## What happens when the control plane comes back Recovery is its own risk. Kubelets reconnect, list their pods and push a burst of status updates. Controllers re-list everything and reconcile a large backlog. The node lifecycle controller in kube-controller-manager protects against one failure mode: if every zone looks fully unhealthy (for example because no node could renew its heartbeat while the API server was down), it enters a **full-disruption** state and stops evicting pods rather than deleting everything at once. ## How to use this during an incident - Check what users actually see before calling it a total outage. The data plane is often still healthy. - Hold off on node reboots and drains, and on anything else that relies on controllers to recover, until the API server is back. - Make a list of the time-sensitive work (CronJobs, autoscaling and deployments in flight) that has to be checked once the control plane returns. - For long outages, remember the data plane slowly drifts from reality: failed nodes are not replaced and the routing to dead pods is not cleaned up.

  • If a node dies while the Kubernetes API server is down, what happens to its pods?
    Nothing is done about them. The node lifecycle controller can't see heartbeats or taint the Node, and even if it could, no controller can create replacements. The EndpointSlices still list those pod IPs, so kube-proxy on the healthy nodes keeps sending them traffic, and those connections fail. Capacity and routing are only repaired after the API server returns and the controllers reconcile.
  • Why can the return of the Kubernetes API server itself cause trouble?
    Every kubelet and controller reconnects at once, re-lists objects and pushes a backlog of status updates. That burst can overload the API server and etcd. The node lifecycle controller also sees a whole cluster of stale heartbeats. It has a full-disruption state that stops evictions when all zones look unhealthy, so the cluster doesn't evict every pod on the way back.

It is like an airport whose control tower loses power. Planes already in the air keep flying their filed routes, but no new flight gets cleared and nobody gets rerouted around trouble.

saying these in an interview costs you the question

  • All pods stop as soon as the API server goes down
  • The kubelet stops restarting crashed containers without the API server
  • Service traffic stops because kube-proxy needs the API server for every packet
  • Controllers keep replacing failed pods from their local cache
  • A missed CronJob run is always started automatically once the control plane returns
open as a page

How do you run a command or open an interactive shell inside a container of a running Kubernetes pod with kubectl, and what are the limits of that approach?

level: juniorimportance: must knowfreq 72%

basics

~20 s

Use kubectl exec -it POD -c CONTAINER -- sh. It starts an extra process inside an already-running container, streamed through the API server and the kubelet. The binary must exist in the image and the container must be running.

open as a page

What does a Kubernetes container's imagePullPolicy control, what is its default, and how does it interact with images already cached on the node?

level: juniorimportance: must knowfreq 68%

basics

~20 s

imagePullPolicy tells the kubelet when to check the registry: Always on every start, IfNotPresent only when the image is not cached, Never not at all. Unset, it defaults to Always for :latest or images with no tag and no digest, otherwise IfNotPresent.

open as a page

How do you read container logs with kubectl, including output from a container instance that has already exited, and why might kubectl logs return nothing at all?

level: juniorimportance: must knowfreq 70%

basics

~20 s

kubectl logs POD -c CONTAINER shows the current instance; --previous shows the last terminated one. Add -f to follow, --since and --tail to bound output. Nothing appears if the app logs to a file instead of stdout/stderr, the pod never started, or the logs rotated.

open as a page

In Kubernetes, after a pod is deleted and its Deployment replaces it, can you still read the old pod's container logs with kubectl, and why?

level: juniorimportance: must knowfreq 72%

basics

~20 s

No. Kubernetes keeps container logs only as files on the pod's node, tied to that pod; once the pod is deleted the kubelet removes its containers and log directory, and the replacement pod starts with empty logs.

open as a page

A teammate runs `kubectl get pods` and sees several pods with the status `Evicted`. What does that status mean in Kubernetes, what typically causes it, and what would you check first?

level: juniorimportance: must knowfreq 55%

basics

~20 s

Evicted means the kubelet killed the pod to reclaim a scarce node resource (memory or disk). The pod object stays as a dead record. Check kubectl describe pod for the reason message and kubectl describe node for MemoryPressure or DiskPressure conditions.

open as a page

A pod has been stuck in the `Pending` phase for ten minutes. Walk through how you find out why, and explain what a `FailedScheduling` event message such as "0/6 nodes are available: 4 Insufficient cpu, 2 node(s) had untolerated taint" is telling you.

level: juniorimportance: must knowfreq 72%

basics

~20 s

Pending means the pod is accepted but not bound to a node (or images are still pulling). Run kubectl describe pod and read the Events section: the scheduler's FailedScheduling message lists, per rejection reason, how many nodes it disqualified — that count tells you which constraint to fix.

open as a page

A pod is stuck in ImagePullBackOff, and earlier its events showed ErrImagePull. What are the usual causes of a failed container image pull in Kubernetes, and how do you confirm which one you are hitting?

level: juniorimportance: must knowfreq 76%

basics

~20 s

The kubelet cannot fetch the image. Usual causes: wrong repository name, tag or digest; the tag does not exist; a private registry with missing or wrong imagePullSecrets; registry unreachable or rate-limiting. Run kubectl describe pod and read the exact pull error in Events.

open as a page

In a Kubernetes Service and pod spec, explain the difference between the Service's port, its targetPort, a nodePort, and the container's containerPort — and describe what a caller sees when targetPort is set to the wrong value.

level: juniorimportance: must knowfreq 70%

basics

~20 s

port is what the Service listens on at its cluster IP; targetPort is the pod port traffic is forwarded to; nodePort is an extra port opened on every node for NodePort/LoadBalancer Services; containerPort is documentation of what the container serves. A wrong targetPort gives endpoints but connection refused.

open as a page

How do you tell apart kubectl's 'connection refused', 'x509: certificate has expired' and 'TooManyRequests' errors, and where does each one point?

level: middleimportance: must knowfreq 52%

basics

~20 s

Connection refused means nothing is listening, so the API server is down or the address is wrong. An x509 expiry means the TLS handshake failed, usually on an expired serving certificate. TooManyRequests (HTTP 429) means a live API server is rejecting load.

open as a page

A running Kubernetes pod uses a minimal image with no shell and must not be restarted. How do ephemeral containers created with kubectl debug help, and what does the --target flag change?

level: middleimportance: must knowfreq 52%

basics

~20 s

kubectl debug -it POD --image=busybox --target=app adds an ephemeral container to the live pod, giving you tools without a restart. It shares the pod network; --target additionally shares the target container's process namespace so you see its processes and files via /proc.

open as a page

In Kubernetes, how does an imagePullSecrets reference reach the kubelet, and why does a pull secret that works in one namespace fail in another?

level: middleimportance: must knowfreq 62%

basics

~20 s

A pod's imagePullSecrets names Secrets of type kubernetes.io/dockerconfigjson in its own namespace. The kubelet reads them and passes matching registry credentials to the runtime. Because the reference is namespace-local, a Secret in another namespace is never found.

open as a page

What is the difference between kubectl describe pod and kubectl get pod -o yaml, and how do you read Kubernetes Events correctly?

level: middleimportance: must knowfreq 55%

basics

~20 s

describe is a human summary that also joins in Events referencing the object; -o yaml is the exact API object, spec plus status, suitable for scripting. Events are separate namespaced objects, expire after about an hour, are aggregated with counts, and are not sorted by default.

open as a page

A teammate reports that their application running on Kubernetes is broken. Walk through the kubectl commands you run, in what order, and what each one rules in or out.

level: middleimportance: must knowfreq 68%

basics

~20 s

Confirm context and namespace, then kubectl get pods -o wide for phase, readiness and restarts; kubectl describe pod for events and container state; kubectl logs (with --previous if it restarted); kubectl get -o yaml for exact status; then widen to events, the controller, endpoints, nodes and kubectl top.

open as a page

Explain how the kubelet decides that a Kubernetes node is under resource pressure: which node conditions it sets, and the difference between its hard and soft eviction thresholds.

level: middleimportance: must knowfreq 55%

basics

~20 s

The kubelet watches signals like memory.available and nodefs.available. Crossing a hard threshold evicts pods immediately with no grace period; a soft threshold must stay crossed for a configured grace period first, and pods get a bounded termination grace period. Crossing sets node conditions MemoryPressure or DiskPressure.

open as a page

When a Kubernetes kubelet must reclaim memory on a node, in what order does it choose which pods to kill? Explain how a pod's Quality of Service class and its resource requests determine that ranking.

level: middleimportance: must knowfreq 62%

basics

~20 s

The kubelet ranks pods by whether usage exceeds requests, then by pod priority, then by how much usage exceeds requests. In practice BestEffort pods (no requests) go first, then Burstable pods over their request, and Guaranteed pods (limits equal requests) last.

open as a page

The Kubernetes scheduler reports `Insufficient cpu` for a pod even though `kubectl top nodes` shows the cluster is only 30% utilised. Explain why, and what actually determines whether a pod fits on a node.

level: middleimportance: must knowfreq 68%

basics

~20 s

The scheduler sums pods' CPU/memory requests, not actual usage, against the node's allocatable (capacity minus kube-reserved, system-reserved and eviction thresholds). A cluster idle at 30% can be fully reserved if requests are inflated. Fix by right-sizing requests or adding capacity.

open as a page

A pod stays Pending with the scheduler event `node(s) had untolerated taint {dedicated=gpu: NoSchedule}` and, on the remaining nodes, `node(s) didn't match Pod's node affinity/selector`. Explain both mechanisms and how you would resolve each.

level: middleimportance: must knowfreq 58%

basics

~20 s

Taints are set on nodes to repel pods; a pod needs a matching toleration to be allowed there. nodeSelector/nodeAffinity are set on the pod to require node labels. Taints repel, selectors attract — you usually need both: add the toleration and make sure node labels match.

open as a page

A Kubernetes pod shows the status CrashLoopBackOff. What does that status actually mean, and what is your step-by-step method for finding out why the container keeps dying?

level: middleimportance: must knowfreq 82%

basics

~20 s

It means the container keeps exiting and the kubelet is now waiting, with a growing delay, before restarting it again. It is a symptom, not a cause. Use kubectl describe pod for the last exit code and reason, then kubectl logs --previous for the dead container's output.

open as a page

kubectl describe pod shows a container with Last State: Terminated, Reason: OOMKilled, Exit Code: 137, and a rising restart count. Explain exactly what happened and how you would fix it.

level: middleimportance: must knowfreq 68%

basics

~20 s

The container's processes exceeded the memory limit in its cgroup, so the kernel's OOM killer sent SIGKILL — exit 137 is 128+9. Fix by measuring real memory use, then either raising the limit or making the runtime respect it (for example JVM heap sizing).

open as a page

A pod cannot resolve the name of another Kubernetes Service. How do you determine whether cluster DNS is at fault, and what role does the search-domain and ndots configuration in the pod's /etc/resolv.conf play?

level: middleimportance: must knowfreq 62%

basics

~20 s

Test resolution from inside a pod against the cluster DNS service, using the full name svc.namespace.svc.cluster.local. Pods get a search list and ndots:5 in /etc/resolv.conf, so short names are tried with each search domain appended — which is why cross-namespace short names fail and external lookups make extra queries.

open as a page

A Kubernetes Service exists and its backing pods are Running, but requests to the Service's cluster IP hang, and the EndpointSlice for that Service lists no ready addresses. What causes an empty endpoint list, and how do you track it down?

level: middleimportance: must knowfreq 78%

basics

~20 s

A Service only routes to pods whose labels match its selector and that are Ready. Empty endpoints means either the selector matches nothing (label typo, wrong namespace) or the matching pods fail their readiness probe. Compare the Service selector with actual pod labels, then check readiness.

open as a page

In Kubernetes, how can a container with a 1500m CPU limit be throttled while kubectl top pod shows it using only 410m?

level: middleimportance: must knowfreq 68%

basics

~20 s

The kubelet turns a 1500m limit into 150ms of CPU per 100ms period, shared by all threads. Forty-eight busy threads spend it in about 3ms and then stall for 97ms, while the average stays low.

open as a page

A Kubernetes pod running a search-autocomplete service is slow but never restarts or shows OOMKilled; how do you confirm CPU throttling is the cause?

level: juniorimportance: should knowfreq 58%

basics

~10 s

Check that the container has a CPU limit, then compare its throttled CFS periods with its total periods: cAdvisor's container_cpu_cfs_throttled_periods_total against container_cpu_cfs_periods_total, or nr_throttled in cpu.stat. A rising ratio confirms throttling.

open as a page

What does kubectl port-forward actually do, and when is it the right tool compared with exposing a Kubernetes Service?

level: middleimportance: should knowfreq 55%

basics

~20 s

kubectl port-forward opens a local TCP listener and tunnels it through the API server to the kubelet and into one pod's network namespace. It is TCP-only, single-pod, unbalanced and dies with the connection — a debugging and admin tool, never production access.

open as a page

How do you find and inspect a specific subset of objects across a large Kubernetes namespace — for example every pod not in the Running phase on one particular node — using kubectl selectors and output formats?

level: middleimportance: should knowfreq 42%

basics

~20 s

Use -l for label selectors (equality and set-based) and --field-selector for a small whitelist of built-in fields such as status.phase and spec.nodeName; both filter server-side. Shape output with -o wide, custom-columns, jsonpath, and --sort-by, which are client-side.

open as a page

Where does the kubelet keep a Kubernetes container's log files on a node, and how does its containerLogMaxSize rotation change what kubectl logs returns?

level: middleimportance: should knowfreq 46%

basics

~10 s

The runtime writes each container's output to /var/log/pods/<namespace><pod><uid>/<container>/<restart>.log, symlinked from /var/log/containers. The kubelet rotates it at containerLogMaxSize (10Mi), keeps containerLogMaxFiles (5), and kubectl logs reads only the current file.

open as a page

A pod sits in the status CreateContainerConfigError and never produces any application logs. What class of problem does that status indicate, and how is it different from a container that starts and then crashes?

level: middleimportance: should knowfreq 44%

basics

~20 s

The kubelet cannot assemble the container's configuration, almost always because a referenced ConfigMap, Secret or a specific key inside one does not exist. The container is never created, so there is no exit code and no logs — the detail is in the pod's Events.

open as a page

Midway through upgrading a 48-node kubeadm Kubernetes cluster, kubectl stops answering entirely. How do you diagnose the control plane from a control-plane node using crictl and the static-pod manifests?

level: seniorimportance: should knowfreq 45%

basics

~20 s

SSH to a control-plane node. Check the kubelet, then use crictl ps -a and crictl logs on kube-apiserver and etcd, and inspect /etc/kubernetes/manifests. Check etcd quorum first, because an apiserver restart loop is often an etcd symptom.

open as a page

You need a root shell on a Kubernetes worker node itself — to check the kubelet, disk usage, or the container runtime — but SSH to nodes is not available. How do you get one with kubectl, and what does it give you?

level: seniorimportance: should knowfreq 33%

basics

~20 s

kubectl debug node/NODE -it --image=busybox creates a privileged pod pinned to that node in the host network, PID and IPC namespaces, with the node's root filesystem mounted at /host. Run chroot /host to work as if on the node. Delete the pod afterwards.

open as a page

showing 1–30 of 46