Where do the numbers reported by kubectl top pods and kubectl top nodes come from, and what should you avoid concluding from them?
answer
- metrics-server add-on, metrics.k8s.io
- scrape ~15s, latest sample only, in memory
- CPU = averaged rate; memory = working set
- top node = usage; describe node = requests
- spot check, not trends or alerting
basics
~20 sThey come from metrics-server, an optional add-on that scrapes each kubelet every ~15 seconds and keeps only the latest sample in memory. CPU is an average rate over that window and memory is working-set bytes — so no history, no spikes, and not the same as your app's heap.
solid answer
~50 s`kubectl top` reads the `metrics.k8s.io` API served by **metrics-server**, which is not installed by default. metrics-server scrapes each kubelet's resource-metrics endpoint (fed by cAdvisor) roughly every 15 seconds and keeps only the most recent window in memory — the same data the Horizontal Pod Autoscaler consumes. So CPU is an **average rate** in cores over the last window, not an instantaneous value: a pod that spikes to its limit for two seconds may show a modest number. Memory is **working-set bytes**, which includes some page cache and is not RSS and definitely not JVM heap — it is roughly the figure the kernel's OOM decision tracks, so it is usually higher than what the runtime reports internally. What not to conclude: no trend or leak claim (there is no history), no capacity plan, no alerting. Also note `kubectl top node` shows real usage against allocatable, whereas `kubectl describe node` shows **requests and limits** — a node can be 100% requested and 10% used, which is a different problem.
code
bash · 4 lineskubectl top nodes
kubectl -n prod top pods --sort-by=cpu
kubectl -n prod top pod api-7d9f-abcde --containers
kubectl describe node node-3go deeper
Know the commands and that they need metrics-server, and that they show current usage, not requests.
Explain the metrics-server scrape path and that memory is working set while CPU is an averaged rate.
Contrast usage with requests when diagnosing scheduling versus pressure, and state clearly what conclusions the data cannot support.
Position kubectl top as a spot check inside a broader strategy where retained metrics own trends, autoscaling inputs and alerting.
## The data path `kubectl top pod` and `kubectl top node` query the aggregated API group `metrics.k8s.io`. That API is served by **metrics-server**, an optional cluster add-on. metrics-server discovers nodes, scrapes each kubelet's resource metrics endpoint, and stores **only the latest sample per object in memory**. Nothing is persisted; restarting metrics-server empties it. The same API backs the Horizontal Pod Autoscaler, which is why a missing or unhealthy metrics-server breaks autoscaling and `kubectl top` at the same time — the tell is an error that the Metrics API is not available. The kubelet's numbers come from cAdvisor reading cgroup accounting for each container, so they are the kernel's view of the container's resource use, not the application's self-report. ## What the two numbers actually mean **CPU** is reported in cores or millicores and is a **rate averaged over the scrape window**, computed from the cumulative CPU-time counter. Two consequences: short spikes are smoothed away, and a freshly started pod may show nothing until two samples exist. If a pod is being CPU-throttled by its limit, `top` shows usage pinned near the limit but does not show the throttling itself — that lives in container throttling counters, which is monitoring-stack territory. **Memory** is working-set bytes: memory charged to the cgroup minus inactive file-backed pages. It therefore includes active page cache from files the container read, and is neither RSS nor the language runtime's heap. Practical reading: a JVM with a 512Mi heap can show 700Mi or more of working set once metaspace, thread stacks, direct buffers and cache are counted, and that is normal, not a leak. Working set is close to what the kernel considers when deciding to OOM-kill a container against its memory limit, which makes it the right number for "am I near the limit" and the wrong number for "how big is my heap". ## Useful invocations - `kubectl top pod -n prod --sort-by=cpu` — rank the noisiest pods. - `kubectl top pod mypod --containers` — split by container, identifying whether the sidecar or the app is consuming. - `kubectl top node` — per-node CPU and memory usage plus the percentage of **allocatable** (not capacity: allocatable subtracts kubelet and system reservations). - `kubectl top pod -A --sort-by=memory` — quick cluster-wide offenders. ## The comparison people get wrong `kubectl top node` reports **actual usage**. `kubectl describe node` reports **requests and limits** of the pods scheduled there. These answer different questions: - Pods are Pending and unschedulable while `top node` shows 15% usage: the node is out of **requested** capacity, not out of real capacity. The fix is right-sizing requests, not adding CPU. - `top node` shows 95% memory while requests total 40%: workloads use far more than they request, and the node is at genuine risk of pressure and eviction. Stating which of the two you are looking at is a strong signal of understanding. ## The limits to declare No history, no percentiles, no per-container time series, one sample per object, and a delay of up to a scrape interval or two. That makes `kubectl top` a **spot check**, appropriate for "what is hot right now" during an incident, and unfit for trend analysis, leak confirmation, capacity planning, or alerting — all of which belong to the monitoring stack with real retention. A candidate who tries to prove a memory leak from two `kubectl top` readings is over-reading the tool; the honest use is to spot the outlier pod and then go look at its retained metrics or take a heap dump.
- kubectl top node shows 20% CPU usage, yet new pods stay Pending with insufficient cpu. Explain.Scheduling is based on pod resource requests against the node's allocatable capacity, not on measured usage. If the pods already placed there request far more CPU than they actually consume, the node is fully booked from the scheduler's point of view while sitting nearly idle. The remedy is to right-size requests to reflect real usage, which you see by comparing kubectl top with the requests shown in kubectl describe node.
saying these in an interview costs you the question
- Assuming kubectl top works everywhere — it needs metrics-server installed
- Equating top's memory number with the application's heap or RSS
- Diagnosing a memory leak or capacity need from two spot readings
- Confusing kubectl top node usage with the requests shown by kubectl describe node
- Expecting top to reveal short CPU spikes or throttling