skip to content

What does a Kubernetes HorizontalPodAutoscaler object do, what must be true of the workload for it to work, and what does it deliberately not do?

level: juniorimportance: must knowfreq 70%

answer

  1. scale subresource, not pods directly
  2. metrics-server serves metrics.k8s.io
  3. Utilization = used ÷ requested
  4. min/max clamp, ScalingLimited condition
  5. no DaemonSets, no static replicas field

basics

~20 s

It watches a metric (usually CPU) for a Deployment or StatefulSet and rewrites that workload's replica count, clamped between minReplicas and maxReplicas, to keep the metric near a target. Needs metrics-server plus resource requests on the pods. It adds and removes pods only — it never resizes a pod and never adds nodes.

solid answer

~50 s

A **HorizontalPodAutoscaler (HPA)** is an API object plus a control loop in kube-controller-manager. It points at a scale target — a Deployment, StatefulSet or ReplicaSet, anything exposing the `scale` subresource — reads a metric roughly every 15 seconds, and writes a new `.spec.replicas` so the observed metric moves toward the target, clamped to `minReplicas`/`maxReplicas`. Prerequisites that bite people: - **A metrics source must exist.** CPU/memory targets need **metrics-server** serving the `metrics.k8s.io` API; it is not part of a vanilla cluster. - **Pods must declare resource requests** when the target type is `Utilization`, because utilization is a percentage *of the request*. No request, no denominator, and the HPA reports `<unknown>`. What it does not do: it does not change a pod's CPU/memory sizing, it does not create nodes, and it fights anything else that writes `replicas` on the same object — so drop the static `replicas:` field from a manifest that GitOps keeps reapplying.

go deeper

for a junior

Be able to say what it scales, the metric-target idea, and the two prerequisites: metrics-server and resource requests.

for a middle

Add the mechanics — it patches the scale subresource on a ~15s loop, min/max clamp the result, and Utilization is a percentage of the request.

for a senior

Frame it as one layer of a stack: HPA for replicas, VPA for sizing, cluster autoscaler for nodes, and name the conflicts between them and with GitOps ownership of replicas.

for a principal

Discuss when replica scaling is the right control at all versus queue-based shedding or capacity reservation, and who owns the replica field across the delivery pipeline.

## Horizontal versus vertical "Horizontal" scaling changes the **number of identical replicas** of a workload. "Vertical" scaling changes how much CPU and memory a *single* replica gets. A HorizontalPodAutoscaler only does the first — it never edits a container's `resources` block. ## The object and the loop An HPA is an ordinary Kubernetes resource in the `autoscaling/v2` API group (GA since Kubernetes 1.23; the older `v2beta2` was removed in 1.26). Its key fields are: - `scaleTargetRef` — which workload to resize. It must expose the `scale` subresource: Deployment, StatefulSet, ReplicaSet, ReplicationController, or a custom resource that implements it. **DaemonSets cannot be targeted** — they run one pod per node by design. - `minReplicas` / `maxReplicas` — the hard floor and ceiling. The controller never leaves this range. - `metrics[]` — one or more metric specs. The commonest is a `Resource` metric on `cpu` with `target.type: Utilization` and `averageUtilization: 70`. The HPA controller runs inside kube-controller-manager on a sync loop, by default every 15 seconds (`--horizontal-pod-autoscaler-sync-period`). Each tick it lists the pods matching the target's selector, fetches their metrics, computes a desired replica count, and if it differs from the current one it PATCHes the target's `scale` subresource. The Deployment controller then does the actual pod creation or deletion — the HPA itself never touches pods. ## Why requests are mandatory With `target.type: Utilization`, the metric is not "CPU cores used". It is **used ÷ requested, as a percentage**. A pod requesting 200m and burning 100m is at 50% utilization. If the container has no CPU request at all, there is no denominator, the utilization is undefined, and `kubectl get hpa` shows `<unknown>/70%` while the replica count sits still. This single omission is the most common reason a freshly written HPA appears to do nothing. The alternative `target.type: AverageValue` sidesteps requests — you state an absolute per-pod number such as `200m` — but for CPU the utilization form is the conventional choice because it stays correct when you resize the pods. ## Where the metrics come from CPU and memory arrive through the **resource metrics API**, `metrics.k8s.io`, which is served by the **metrics-server** deployment. metrics-server scrapes each kubelet's summary endpoint every ~15s and keeps only the latest sample in memory — it is a live sensor, not a time-series database, and it is deliberately not suitable for historical queries. Managed distributions usually pre-install it; a kubeadm or kind cluster does not, and without it both `kubectl top` and every HPA fail together, which is a fast way to confirm the diagnosis. ## What it is not - **Not the Vertical Pod Autoscaler.** VPA adjusts requests/limits; running both on CPU for the same workload is a known conflict. - **Not the Cluster Autoscaler.** If the HPA asks for 20 replicas and the nodes are full, the extra pods sit `Pending` until node-level autoscaling provides capacity. The two are complementary layers, not substitutes. - **Not a scheduler.** The HPA has no say in *where* pods land. - **Not compatible with a hardcoded replica count.** Leaving `replicas: 3` in a manifest that Argo CD or Flux reconciles produces a tug-of-war: the HPA scales to 9, the sync scales back to 3. Remove the field (or mark it ignored) and let the HPA own it. ## Reading its state `kubectl get hpa` shows `TARGETS` as `current/target`, plus min, max and current replicas. `kubectl describe hpa` prints the conditions — `AbleToScale`, `ScalingActive`, `ScalingLimited` — and recent events explaining each decision. `ScalingLimited: True` means the desired count was clipped by min or max, which is usually the answer to "why did it stop at 10?".

  • Your HPA and your GitOps tool disagree about the replica count. What is happening and how do you fix it?
    Both are writing `.spec.replicas` on the same Deployment: the HPA scales up, the next reconcile resets it to the committed value, and the workload oscillates. The fix is to make the HPA the sole owner — delete the `replicas` field from the manifest entirely, or tell the GitOps controller to ignore that path (for example Argo CD's `ignoreDifferences` on `/spec/replicas`).
  • Can you attach a HorizontalPodAutoscaler to a DaemonSet?
    No. A DaemonSet has no `scale` subresource because its replica count is defined by the set of matching nodes, not by a number you choose. If you need per-node capacity to grow, that is a node-count or resource-sizing problem, not a horizontal pod scaling one.

A thermostat wired to the number of radiators in a room, not to how hot each one runs: it can switch radiators on and off, but it cannot make one hotter, and it cannot build you a bigger house.

saying these in an interview costs you the question

  • Believing the HPA changes CPU/memory limits on pods
  • Assuming metrics-server ships with every cluster
  • Writing a Utilization target on containers that declare no resource requests
  • Thinking the HPA creates nodes when pods go Pending
  • Leaving a hardcoded replicas value in a manifest that a GitOps controller reapplies

context