A teammate reports that their application running on Kubernetes is broken. Walk through the kubectl commands you run, in what order, and what each one rules in or out.
answer
- context + namespace first
- get -o wide: phase, READY, RESTARTS, AGE, node
- describe: conditions, lastState, Events
- logs, then --previous
- capture before you delete
basics
~20 sConfirm context and namespace, then kubectl get pods -o wide for phase, readiness and restarts; kubectl describe pod for events and container state; kubectl logs (with --previous if it restarted); kubectl get -o yaml for exact status; then widen to events, the controller, endpoints, nodes and kubectl top.
solid answer
~50 sI work from narrow to wide, and from *what the control plane thinks* to *what the app said*. 1. **Am I in the right place**: `kubectl config current-context`, correct `-n`. 2. `kubectl get pods -o wide` — phase, READY count, RESTARTS, AGE, node. This alone splits the problem: Pending means scheduling, 0/1 Ready means probes, high restarts means crashing, Running and Ready means look at traffic or the app. 3. `kubectl describe pod <p>` — conditions, container state and last termination reason, image, probes, and the Events at the bottom, which carry scheduler, kubelet and image-pull messages. 4. `kubectl logs <p> -c <c>`, adding `--previous` if it restarted, `--since=15m` and `-f` when watching. 5. `kubectl get pod <p> -o yaml` when describe is ambiguous — exact `status.containerStatuses[].lastState`, conditions, resolved env and volumes. 6. Widen only if needed: `kubectl get events --sort-by=.lastTimestamp`, the Deployment and ReplicaSet, `kubectl get endpointslices`, `kubectl get nodes`, `kubectl top pod`. And I avoid deleting the pod before capturing logs — that destroys the evidence.
code
bash · 6 lineskubectl config current-context
kubectl -n prod get pods -o wide
kubectl -n prod describe pod api-7d9f-abcde
kubectl -n prod logs api-7d9f-abcde -c api --tail=200 --timestamps
kubectl -n prod logs api-7d9f-abcde -c api --previous
kubectl -n prod get pod api-7d9f-abcde -o yaml > /tmp/pod.yamlgo deeper
Recall the four core commands — get, describe, logs, get -o yaml — and what each shows.
Present them as an elimination sequence and explain what each column and section rules out.
Add evidence preservation, widening to controller, Service and node scope, and reasoning from one bad replica or one bad node.
Turn the sequence into a shared runbook with a defined classification and handoff, and note where kubectl stops and the observability stack begins.
## Why order matters Under time pressure the value of a triage sequence is that each step **eliminates a class of causes**, so you never wander. The spine is: confirm where you are, read the object's summary state, read what the cluster did to it, read what the application said, then read the exact object, then widen to the neighbourhood. ## Step 0 — context and namespace `kubectl config current-context` and `kubectl config view --minify` answer "which cluster", and `-n <ns>` or `kubectl config set-context --current --namespace=<ns>` answer "which namespace". Most "the pod doesn't exist" confusion is a wrong namespace or a wrong cluster, and acting in the wrong cluster is the one triage mistake that causes a second outage. Keep kubectl within one minor version of the API server; skew produces odd missing-field behaviour. ## Step 1 — kubectl get pods -o wide Six columns answer most of the question: - **STATUS** — `Pending` (not scheduled or image not yet pulled), `ContainerCreating`, `Running`, `Completed`, `Error`, `Terminating`, or a backoff reason. - **READY** — `0/1` while `Running` means the readiness probe is failing, so the pod exists but takes no Service traffic; `1/2` in a sidecar pod points at which container is unhealthy. - **RESTARTS** — with the "(x ago)" suffix in modern kubectl, tells you both how often and how recently. - **AGE** — distinguishes "never worked since the deploy 3 minutes ago" from "worked for 12 days and just broke". - **NODE** (from `-o wide`) — if every unhealthy pod is on one node, the problem is that node, not the app. Add `-w` to watch transitions live, `-A` to search across namespaces. ## Step 2 — kubectl describe pod `describe` renders the object plus the Events referring to it. Read bottom-up: Events tell you what the scheduler, kubelet and image puller actually did and when. Above them, check `State` and `Last State` for the current and previous container run (with exit code and reason), the `Conditions` block (`PodScheduled`, `Initialized`, `ContainersReady`, `Ready`), the resolved image and image ID, the probe definitions, and the requests and limits. `describe` also shows `Controlled By`, which is how you find the owning ReplicaSet, Job or StatefulSet. ## Step 3 — kubectl logs What the application itself said. `-c` for a specific container, `--previous` for the instance that died, `--since=15m` and `--tail=200` to bound the volume, `--timestamps` to correlate with the Events you just read. Absence of logs is information too — a pod that never started produces none. ## Step 4 — kubectl get -o yaml When describe's summary is ambiguous, read the API object: `status.containerStatuses[].lastState.terminated` (exitCode, reason, finishedAt), `status.conditions`, the fully resolved `spec` including injected sidecars, env from ConfigMaps and Secrets, volumes, and `metadata.ownerReferences`. This is also the copy you attach to an incident ticket, since it is exact rather than formatted. ## Step 5 — widen If the pod looks healthy, the problem is around it. `kubectl get events --sort-by=.lastTimestamp -n <ns>` for cluster-level activity such as evictions or scaling; `kubectl get deploy,rs` to see whether a rollout is stuck or an old ReplicaSet is still serving; `kubectl get svc,endpointslices` to see whether the pod is actually a backend; `kubectl get nodes` for node conditions; `kubectl top pod` and `kubectl top node` for live resource usage. ## Habits that make triage repeatable Capture before you mutate: `kubectl get pod -o yaml`, `describe`, and `logs --previous` into files or the incident channel *before* deleting or restarting anything, because deleting a pod deletes its logs and its status. Prefer read-only commands until the cause is understood. State findings as eliminations — "the pod is Running and Ready, so this is not a scheduling problem" — because narrowing is the whole point of the sequence.
- Everything shows Running and 1/1 Ready, yet users get errors. Where do you look next with kubectl?Move outward from the pod: check that the Service actually selects those pods and that EndpointSlices list them, compare the pod labels with the Service selector, and confirm the container port matches the Service targetPort. Then look at rollout state with kubectl get deploy and rs, and read logs across all replicas with a label selector, since only one replica may be bad.
- Why is deleting a misbehaving pod a poor first move even though it often 'fixes' the symptom?Deleting the pod destroys its status, its events refer to an object that no longer exists, and its logs become unreachable because kubectl logs reads them per pod from the node. The controller creates a fresh pod that may work, leaving you with no evidence and a recurrence later. Capture yaml, describe output and previous logs first, then delete if you need to restore service.
saying these in an interview costs you the question
- Jumping straight to deleting or restarting pods before capturing logs and status
- Forgetting to check the namespace and cluster context
- Reading only logs and never the Events, so scheduling and image-pull problems are invisible
- Assuming Running means healthy while READY shows 0/1
- Never using -o wide, so a single bad node goes unnoticed