A Kubernetes HorizontalPodAutoscaler is not scaling and `kubectl get hpa` prints `<unknown>` in the TARGETS column. How do you diagnose it, and what are the usual root causes?
answer
- describe hpa → ScalingActive / FailedGetResourceMetric
- kubectl top fails too ⇒ metrics-server
- x509 no IP SANs ⇒ kubelet serving certs
- top works, HPA unknown ⇒ missing request
- unavailable metric ⇒ no scale-down
basics
~20 sCheck whether metrics exist at all with kubectl top pods. If that fails, metrics-server is missing or unhealthy (often TLS/kubelet-certificate or APIService issues). If top works but the HPA is unknown, the pods lack a resource request for the metric, the selector matches no ready pods, or a custom-metrics adapter is down. kubectl describe hpa names the failing condition.
solid answer
~60 sWork down the chain, because `<unknown>` just means the controller could not compute a value. 1. **`kubectl describe hpa`** — read the conditions. `ScalingActive: False` with `FailedGetResourceMetric` points at the metrics pipeline; the event message usually names the exact error. 2. **`kubectl top pods`** — if this also fails, the problem is **metrics-server**, not the HPA. Check the `v1beta1.metrics.k8s.io` APIService is `Available`, and check metrics-server logs for the classic kubelet TLS failure (`x509: cannot validate certificate`), which on self-managed clusters is usually solved by serving-certificate rotation rather than the `--kubelet-insecure-tls` shortcut. 3. **If `top` works but the HPA doesn't** — the pods almost certainly have **no resource request** for the targeted resource, since `Utilization` is a percentage of the request. Also check the target's selector actually matches running pods, and that the pods are Ready. 4. **Custom/external metrics** — `<unknown>` means the adapter behind `custom.metrics.k8s.io` or `external.metrics.k8s.io` is unreachable or has no series for that name; query the API directly with `kubectl get --raw`. Remember the safety behavior: while a metric is unavailable the controller will **not scale down**.
code
bash · 5 lineskubectl describe hpa api
kubectl top pods -l app=api
kubectl get apiservices v1beta1.metrics.k8s.io
kubectl -n kube-system logs deploy/metrics-server
kubectl get --raw "/apis/external.metrics.k8s.io/v1beta1" | headgo deeper
Know the two first checks — describe the HPA, then run kubectl top — and that a missing resource request is the classic cause.
Separate the pipeline layers: metrics-server / APIService / requests / selector, and check each in order.
Add the TLS and node-address failure modes, the custom-metrics adapter path via kubectl get --raw, and the no-scale-down-on-unknown safety rule.
Talk about making this undiagnosable-by-accident: alert on ScalingActive, enforce requests via admission policy, and decide how much the platform should depend on a single metrics component.
## What `<unknown>` actually means The TARGETS column is `current/target`. `<unknown>` means the HorizontalPodAutoscaler controller could not produce a `current` value this cycle. It is a *pipeline* symptom, not a scaling decision, so the diagnosis is a walk down the metric supply chain. ## Step 1 — the HPA's own conditions `kubectl describe hpa <name>` prints three conditions: - **AbleToScale** — can it reach and patch the scale subresource at all? - **ScalingActive** — is it computing a recommendation? `False` with reason `FailedGetResourceMetric` or `FailedGetPodsMetric` is the `<unknown>` case, and the message carries the underlying error text. - **ScalingLimited** — is the result being clipped by min/max? Read the message before touching anything; it usually distinguishes "no metrics API registered" from "missing request on container X". ## Step 2 — is the resource metrics API alive? CPU and memory come from `metrics.k8s.io`, served by **metrics-server** through an APIService aggregation registration. Two checks: ``` kubectl top pods # exercises the same API path kubectl get apiservices | grep metrics ``` If the APIService shows `False (MissingEndpoints)` or `ServiceUnavailable`, metrics-server is not running or not reachable from the API server. Common causes: - **metrics-server not installed.** kubeadm, kind and bare clusters do not ship it. Both `kubectl top` and every HPA fail together — a clean signature. - **Kubelet TLS.** metrics-server scrapes kubelets over HTTPS and by default validates their certificates. On clusters whose kubelet serving certs are self-signed and unsigned by the cluster CA, the logs read `x509: cannot validate certificate for <ip> because it doesn't contain any IP SANs`. The correct fix is enabling kubelet serving-certificate rotation and approving the CSRs; `--kubelet-insecure-tls` is a lab workaround, not a production answer. - **Node address resolution.** metrics-server needs a reachable node address; `--kubelet-preferred-address-types=InternalIP` is the usual correction on clouds where hostnames don't resolve. - **API server aggregation disabled** or the aggregator unable to route to the metrics-server service (network policy, missing `--enable-aggregator-routing`). ## Step 3 — metrics exist, HPA still unknown If `kubectl top pods` returns numbers for the workload's pods but the HPA doesn't, the failure is between raw metrics and the computation: - **No resource request on the container.** With `target.type: Utilization` the value is `usage ÷ request`; absent a request there is no denominator and the metric is unknown. This is far and away the most common cause and it is a *per-container* problem — one sidecar without a CPU request is enough, since the pod-level sum needs every container's request. Either add requests everywhere or switch to `ContainerResource` scoped to the container you care about. - **Selector matches nothing.** The HPA uses the scale target's selector; if the Deployment has zero pods, or you targeted the wrong object, there is nothing to average. - **All pods unready or terminating.** Terminating pods are skipped, and unready ones are excluded from the initial computation. A workload stuck in CrashLoopBackOff produces an HPA with no usable sample. - **Fresh pods.** For the first scrape interval after a rollout, samples genuinely do not exist yet. Transient `<unknown>` right after a deploy is normal; persistent is not. ## Step 4 — custom and external metrics For `Pods`, `Object` or `External` metrics the serving API is `custom.metrics.k8s.io` or `external.metrics.k8s.io`, provided by an adapter (Prometheus Adapter, KEDA's metrics adapter, a cloud provider's). Query it directly: ``` kubectl get --raw "/apis/custom.metrics.k8s.io/v1beta1" | jq . kubectl get --raw "/apis/external.metrics.k8s.io/v1beta1/namespaces/prod/queue_depth" | jq . ``` An empty list means the adapter has no rule producing that metric name, or the underlying query returns no series — often because a label selector in the adapter rule doesn't match reality. `<unknown>` here is an adapter configuration bug, not a Kubernetes bug. ## The behavior you must know While a metric is unavailable the controller is deliberately conservative: **it will not scale down**. If some metrics are readable and one is not, scale-up on the readable ones still proceeds, but scale-down is suppressed until every metric is available. So a broken metrics pipeline tends to leave a workload stuck at its current — often elevated — replica count rather than collapsing it. That is the safe failure mode, and pointing it out signals you understand the design intent. ## Prevention Alert on the HPA's `ScalingActive` condition and on metrics-server availability rather than discovering it during an incident; and treat "every container has a request" as an admission-policy rule, since it is the precondition for both utilization-based autoscaling and sane scheduling.
- metrics-server logs show `x509: cannot validate certificate ... doesn't contain any IP SANs`. What is the correct fix?metrics-server scrapes kubelets over HTTPS and validates their serving certificates, which on many self-managed clusters are self-signed without IP SANs. The production fix is to enable kubelet serving-certificate rotation (`--rotate-server-certificates`) so certs are issued and signed by the cluster CA, and approve the resulting CSRs. Passing `--kubelet-insecure-tls` silences validation entirely and is acceptable only in a lab.
- If the metrics pipeline breaks at 3am while a workload sits at 40 replicas, what does the HPA do?Nothing destructive: with the metric unavailable it reports `<unknown>`, sets `ScalingActive: False`, and refuses to scale down. The workload stays at 40 replicas — over-provisioned but serving — until metrics return. The intentional asymmetry is that an unknown metric never justifies removing capacity.
saying these in an interview costs you the question
- Assuming metrics-server is installed by default on every cluster
- Reaching for --kubelet-insecure-tls as the production fix for certificate errors
- Not realizing a single sidecar without a CPU request breaks utilization for the whole pod
- Believing a broken metrics pipeline causes a scale-down to minReplicas
- Debugging the HPA object without ever checking kubectl top or the APIService status