On a 5-node Kubernetes edge cluster, `kubectl top pods` fails with `error: Metrics API not available` while Prometheus dashboards look fine. How do you diagnose it?
answer
- two pipelines, one untouched
- Available condition reason first
- control plane must reach the pod
- front-proxy CA and client cert
- kubelet serving certs, address types
basics
~20 sPrometheus scrapes targets directly, so it proves nothing about the aggregation path. Check the v1beta1.metrics.k8s.io APIService's Available condition and reason, then fix the failing hop: Service and endpoints, apiserver-to-pod reachability and TLS, or metrics-server-to-kubelet scraping.
solid answer
~40 sA healthy metrics backend is irrelevant here: it scrapes its targets directly, while `kubectl top` goes through kube-apiserver's aggregation layer to metrics-server. I start with `kubectl get apiservice v1beta1.metrics.k8s.io` and read the `Available` condition's reason. `ServiceNotFound` or `EndpointsNotFound` means the Service or its EndpointSlices are missing; `MissingEndpoints` usually means the metrics-server pod is not Ready, often because it cannot scrape kubelets. `FailedDiscoveryCheck` means kube-apiserver cannot reach or trust the pod: no route from the control plane to pod IPs, a firewall on port 10250, or broken front-proxy configuration. Then I read metrics-server's logs for kubelet `x509` or address errors and fix them properly, with CA-signed kubelet serving certificates and the right `--kubelet-preferred-address-types`. Finally I confirm with `kubectl get --raw /apis/metrics.k8s.io/v1beta1/nodes`.
code
bash · 4 lineskubectl get apiservice v1beta1.metrics.k8s.io
kubectl -n kube-system get endpointslices -l kubernetes.io/service-name=metrics-server
kubectl -n kube-system logs deploy/metrics-server --tail=50 | grep -Ei 'x509|dial|forbidden'
kubectl get --raw /apis/metrics.k8s.io/v1beta1/nodes | head -c 300go deeper
Remember that kubectl top depends on metrics-server being reachable through the API server, and that kubectl get apiservice is where to start.
Explain the three hops — client to apiserver, apiserver to metrics-server, metrics-server to kubelets — and which APIService reasons map to which hop.
Diagnose from the condition reason and logs, fix trust properly with CA-signed kubelet certificates, and call out the discovery-wide blast radius.
Treat aggregated API availability as a platform SLO across many clusters, with alerting on the Available condition and a standard, verified install.
## Why the healthy dashboards prove nothing There are two independent pipelines in the cluster: - A **metrics backend** such as Prometheus scrapes kubelets, exporters and applications directly and stores the series itself. - `kubectl top` and CPU-based HorizontalPodAutoscalers call **kube-apiserver**, which proxies `/apis/metrics.k8s.io/...` through the **aggregation layer** to **metrics-server**, which in turn scrapes kubelets. The second path has three hops that the first never touches. `kubectl top` first reads API discovery; if no supported `metrics.k8s.io` version is listed it prints `error: Metrics API not available`. Direct calls to an unavailable group get HTTP 503 from the aggregator. ## Step 0: rule out the trivial Before digging into the aggregation layer, spend thirty seconds on the cheap checks: - Is the metrics-server Deployment present and its pod `Running` and `Ready` in `kube-system`? - Did anything change recently — a node image, a firewall rule on the store network, a certificate rotation on the control plane? - Does `kubectl top node` fail in the same way? If nodes work and pods do not, the problem is narrower than the aggregation path. ## Step 1: read the APIService condition ```bash kubectl get apiservice v1beta1.metrics.k8s.io kubectl get apiservice v1beta1.metrics.k8s.io -o jsonpath='{.status.conditions[?(@.type=="Available")]}' ``` If the object does not exist, metrics-server was never installed on this store's cluster. Otherwise the `reason` narrows the hop: | Reason | What kube-apiserver found | Where to look | |---|---|---| | `ServiceNotFound` | The Service in `spec.service` does not exist | Install drift, wrong namespace | | `ServicePortError` | The Service does not listen on the given port | `spec.service.port` vs the Service | | `EndpointsNotFound` | No EndpointSlices for the Service | Selector mismatch, Deployment missing | | `MissingEndpoints` | EndpointSlices exist but have no usable address | metrics-server pod not Ready | | `FailedDiscoveryCheck` | Endpoints exist but the discovery call failed | Network or TLS between control plane and pod | | `Passed` | All checks succeeded | Problem is elsewhere, e.g. RBAC or stale kubectl | ## Step 2: the kube-apiserver to metrics-server hop `FailedDiscoveryCheck` messages usually contain a timeout or a TLS error. Common causes: 1. **No route from control-plane nodes to pod IPs.** The API server calls the pod's IP on its port (10250 in the upstream manifest). If the control plane is not on the pod network, or a NetworkPolicy or host firewall drops that traffic, the check times out. Options are fixing the routing, allowing the port, or running metrics-server with `hostNetwork`. 2. **Routing to the Service's cluster IP without kube-proxy on the control plane.** `--enable-aggregator-routing` makes kube-apiserver send requests to endpoint IPs rather than the cluster IP. 3. **Front-proxy trust.** The aggregator authenticates to extension servers with the client certificate from `--proxy-client-cert-file` / `--proxy-client-key-file`; the extension server trusts it through `--requestheader-client-ca-file`, published in the `kube-system/extension-apiserver-authentication` ConfigMap. metrics-server needs RBAC to read that ConfigMap; a missing binding or a wrong CA breaks every request. 4. **Serving-certificate trust.** Either `spec.caBundle` matches metrics-server's certificate, or `insecureSkipTLSVerify: true` is set (the upstream manifest does this). ## Step 3: the metrics-server to kubelet hop If the pod is not Ready (`MissingEndpoints`), read its logs: - **`x509` errors** mean kubelet serving certificates are self-signed. The production fix is kubelet `serverTLSBootstrap: true` plus approving the resulting CSRs, so certificates chain to the cluster CA. `--kubelet-insecure-tls` removes verification and belongs only in labs. - **Dial errors or wrong addresses** mean the first address type in `--kubelet-preferred-address-types` is not reachable from the pod; edge nodes often publish a hostname that the store's network cannot resolve. - **403 from the kubelet** means metrics-server's ServiceAccount lacks permission on the `nodes/metrics` subresource, or the kubelet's webhook authorization is not enabled. ## Step 4: confirm and assess blast radius ```bash kubectl get --raw /apis/metrics.k8s.io/v1beta1/nodes | head -c 300 kubectl top pods -n checkout --sort-by=cpu ``` While the group is unavailable: - CPU- and memory-based HPAs for the ticket-booking checkout cannot compute a recommendation and hold their replica count. - Discovery clients report `unable to retrieve the complete list of server APIs`, which can break tooling that treats partial discovery as fatal. - Controllers that must enumerate every resource type, such as the namespace controller, can stall. That is why an unavailable APIService is worth an alert of its own, rather than something noticed when someone types `kubectl top`.
- Why can an unavailable metrics.k8s.io APIService affect things far beyond `kubectl top`?Any client that enumerates every API group sees a discovery error for that group. Tools that treat partial discovery as fatal fail outright, and controllers that must list every resource type, such as the namespace controller when purging a namespace, cannot finish. One unreachable add-on server therefore degrades cluster-wide operations, not just metrics.
- The APIService shows `Available=True`, but `kubectl top pods -n checkout` still returns an error. Where do you look next?The aggregation path is healthy, so check the request itself: whether your user has RBAC to `get`/`list` `pods` in the `metrics.k8s.io` group, whether metrics-server simply has no points yet for new pods, and whether its logs show scrape failures for the specific nodes those pods run on. A partial kubelet outage yields missing pods, not a global failure.
- Is setting `--kubelet-insecure-tls` an acceptable permanent fix on store clusters?No. It disables verification of every kubelet's serving certificate, so metrics-server would accept an impersonated kubelet. The durable fix is having kubelets request serving certificates from the cluster CA with `serverTLSBootstrap`, plus an approval process for those CSRs, so metrics-server can verify them against the cluster CA.
saying these in an interview costs you the question
- Prometheus works, so metrics-server must be fine too
- restart kube-apiserver whenever kubectl top fails
- --kubelet-insecure-tls is the standard production setting
- FailedDiscoveryCheck means the metrics-server Service is missing
- an unavailable APIService only affects kubectl top