skip to content

Aggregated APIs & Metrics

An APIService hands a whole API group to a second server, which is how metrics-server serves metrics.k8s.io for kubectl top and the HPA. Interviewers ask why kubectl top fails while Prometheus is fine, and what metrics-server deliberately is not.

part ofKubernetesoverview, primer and where to startread it →
on this pageshow

questions

3

In Kubernetes, what does metrics-server provide for `kubectl top` and the HorizontalPodAutoscaler, and what is it deliberately not designed to be?

level: juniorimportance: must knowfreq 72%

answer

  1. narrow, current, in-memory
  2. kubelet /metrics/resource scrape
  3. APIService v1beta1.metrics.k8s.io
  4. two points make a CPU rate
  5. no history, no custom metrics

basics

~20 s

metrics-server scrapes current CPU and memory usage from every kubelet, keeps only the latest points in memory, and serves them as the metrics.k8s.io API. It is not a monitoring system: no history, no custom metrics, no alerting.

solid answer

~40 s

metrics-server is a small cluster add-on that periodically scrapes each kubelet's `/metrics/resource` endpoint for CPU and memory usage of nodes, pods and containers. It keeps only the most recent scrapes in memory and serves them through the `metrics.k8s.io` API, which it plugs into kube-apiserver through an `APIService` object named `v1beta1.metrics.k8s.io`. That API is what `kubectl top` and a CPU- or memory-based HorizontalPodAutoscaler read. It is deliberately *not* a monitoring system: it stores no history, loses everything on restart, serves no application or custom metrics, sends nothing to a backend and does not claim to be an accurate source for billing or capacity reports. For trends, dashboards and alerts you run a real metrics backend alongside it.

code

bash · 3 lines
bash
kubectl top node
kubectl top pod -n checkout --sort-by=memory --containers
kubectl get --raw /apis/metrics.k8s.io/v1beta1/namespaces/checkout/pods | head -c 400

go deeper

for a junior

Remember three facts: it serves current CPU and memory, it is what kubectl top and a CPU-based HPA read, and it keeps no history.

for a middle

Explain the path: kubelet /metrics/resource, periodic scrape, two points in memory, served as metrics.k8s.io through an APIService and the aggregation layer.

for a senior

Show you know when it is enough and when it is not: it covers autoscaling and quick triage, while trends, alerting and chargeback need a separate metrics backend.

for a principal

Frame the cost trade-off: on many small edge clusters, metrics-server alone may be the right footprint, provided the organisation accepts it cannot answer historical questions.

## What metrics-server is **metrics-server** is a Kubernetes SIG-maintained add-on, usually running as a Deployment in `kube-system`, that implements the **resource metrics API**. That API lives in the API group `metrics.k8s.io` and exposes two read-only resources: - `nodes` — current CPU and memory usage per Node - `pods` — current CPU and memory usage per Pod, broken down by container It exists so that core Kubernetes features have a *standard, always-present* answer to one narrow question: "how much CPU and memory is this thing using right now?" The two main consumers are `kubectl top` and the **HorizontalPodAutoscaler** when it targets CPU or memory utilisation. ## How the data flows 1. Every kubelet already measures container usage through the container runtime and publishes it on its HTTPS port (10250 by default) at `/metrics/resource`. 2. metrics-server scrapes that endpoint on every node at a fixed interval, set by `--metric-resolution`. The upstream manifest sets `15s`; the flag rejects anything below `10s`. 3. It keeps the **last two scrapes** per container and node in memory. CPU usage is a rate, so it is computed from the difference between those two points; memory is the latest working-set value. 4. It serves the result as an ordinary-looking Kubernetes API. An `APIService` object named `v1beta1.metrics.k8s.io` tells kube-apiserver to forward every request under `/apis/metrics.k8s.io/v1beta1/` to the `metrics-server` Service. 5. `kubectl top pods` and the HPA controller call kube-apiserver as usual; the **aggregation layer** proxies the call to metrics-server, which answers from memory. Because the answer comes from memory, a request for a 1,180-pod namespace in a ticket-booking checkout is cheap — no query engine, no disk. ## What it deliberately is not The project README states its non-goals plainly: it is not for non-Kubernetes clusters, not "an accurate source of resource usage metrics", and not for autoscaling on anything other than CPU and memory. | Expectation | Reality in metrics-server | |---|---| | History / trends | None — only the latest window; a restart starts from empty | | Application metrics (requests per second, queue depth) | Not served; those come from custom or external metrics adapters | | Dashboards and alerting | Not provided; there is no query language and no rules | | Export to a backend | Nothing is pushed anywhere | | Billing-grade accuracy | Explicitly disclaimed; it is sized for autoscaling decisions | | Object state (desired vs ready replicas) | Not its job; that is kube-state-metrics | A metrics backend such as Prometheus is a **separate, parallel pipeline**. It scrapes its own targets and stores time series; it does not use metrics-server, and metrics-server does not feed it. ## Reading `kubectl top` correctly - The numbers are **recent usage**, not requests or limits. A pod requesting `500m` but idle shows a few millicores. - Right after a pod starts, or right after metrics-server restarts, a pod may be missing: two scrapes are needed before a CPU rate exists. - `kubectl top pod --containers` shows the per-container split; `--sort-by=cpu` or `--sort-by=memory` orders the list. - `kubectl top node` values come from the kubelet's node-level measurement, so they include system daemons and do not equal the sum of the pods. ## Operating it on a small cluster On a 5-node edge cluster in a retail store, metrics-server is often the *only* metrics component, because shipping a full monitoring stack to every store is expensive. That is fine for `kubectl top` and CPU-based scaling of the checkout, but the team must accept that nobody can answer "what was checkout's memory at 14:05 yesterday?" from it. Its footprint grows with the number of pods and nodes it tracks, so its memory request is sized to cluster scale rather than traffic. It also depends on the aggregation layer being configured on kube-apiserver and on kubelets presenting serving certificates it can verify; when either is broken, `kubectl top` fails even though every workload is healthy. ## Common interview traps - **"A metrics backend makes metrics-server redundant"** — they coexist; removing metrics-server breaks `kubectl top` and resource-based HPAs even when a metrics backend is healthy. - **"It reads cAdvisor directly"** — current releases scrape the kubelet's `/metrics/resource` endpoint; the kubelet is the single source. - **"More replicas give more history"** — running several metrics-server replicas is for availability of the API, not for retention; each replica still holds only the latest window. - **"`kubectl top` is authoritative for sizing requests"** — it is a point-in-time glance; sizing needs percentiles over days, which only a stored time series can give.

  • Why does a freshly started pod sometimes not appear in `kubectl top pods` for a short while?
    CPU usage is a rate, so metrics-server needs two scrapes of the same container to compute it. Until the second scrape after the pod starts has landed, which can take one or two `--metric-resolution` intervals, there is no usable point and the pod is left out of the response. The same gap appears for every pod right after metrics-server itself restarts, because its store starts empty.
  • What must be true of a Kubernetes cluster before metrics-server can serve anything?
    kube-apiserver must have the aggregation layer configured, with front-proxy certificates. The control plane must be able to reach the metrics-server pod, and metrics-server must reach every kubelet on its published address and port. Kubelets need webhook authentication and authorization enabled, and their serving certificates must be signed by the cluster CA, unless verification is switched off, which is acceptable only for testing.
  • Can you point a metrics backend at metrics-server to get history instead of running node-level exporters?
    It is the wrong tool. metrics-server serves only CPU and memory, in the Kubernetes API format rather than a scrape format, and is sized for the latest window. A metrics backend should scrape kubelets and node exporters directly, which yields far richer series. The two pipelines are meant to run side by side, each for its own purpose.

metrics-server is a car's speedometer, not its trip log: it tells you how fast you are going right now and forgets the moment you look away.

saying these in an interview costs you the question

  • metrics-server stores a week of usage history for capacity planning
  • kubectl top shows the pod's CPU requests and limits
  • metrics-server is required for a metrics backend such as Prometheus to work
  • metrics-server can serve requests-per-second for autoscaling
  • metrics-server reads usage from etcd
  • kubectl top node equals the sum of all pods on the node
open as a page

Kubernetes defines the metrics.k8s.io, custom.metrics.k8s.io and external.metrics.k8s.io APIs; what serves each one, and where does kube-state-metrics fit?

level: middleimportance: should knowfreq 46%

basics

~20 s

Each group is served by a separate server registered through an APIService: metrics-server for metrics.k8s.io, and adapters for custom and external metrics. kube-state-metrics is not an API at all but a plain exporter of object state.

open as a page

On a 5-node Kubernetes edge cluster, `kubectl top pods` fails with `error: Metrics API not available` while Prometheus dashboards look fine. How do you diagnose it?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Prometheus scrapes targets directly, so it proves nothing about the aggregation path. Check the v1beta1.metrics.k8s.io APIService's Available condition and reason, then fix the failing hop: Service and endpoints, apiserver-to-pod reachability and TLS, or metrics-server-to-kubelet scraping.

open as a page