What does a Kubernetes HorizontalPodAutoscaler object do, what must be true of the workload for it to work, and what does it deliberately not do?
answer
- scale subresource, not pods directly
- metrics-server serves metrics.k8s.io
- Utilization = used ÷ requested
- min/max clamp, ScalingLimited condition
- no DaemonSets, no static replicas field
basics
~20 sIt watches a metric (usually CPU) for a Deployment or StatefulSet and rewrites that workload's replica count, clamped between minReplicas and maxReplicas, to keep the metric near a target. Needs metrics-server plus resource requests on the pods. It adds and removes pods only — it never resizes a pod and never adds nodes.
solid answer
~50 sA **HorizontalPodAutoscaler (HPA)** is an API object plus a control loop in kube-controller-manager. It points at a scale target — a Deployment, StatefulSet or ReplicaSet, anything exposing the `scale` subresource — reads a metric roughly every 15 seconds, and writes a new `.spec.replicas` so the observed metric moves toward the target, clamped to `minReplicas`/`maxReplicas`. Prerequisites that bite people: - **A metrics source must exist.** CPU/memory targets need **metrics-server** serving the `metrics.k8s.io` API; it is not part of a vanilla cluster. - **Pods must declare resource requests** when the target type is `Utilization`, because utilization is a percentage *of the request*. No request, no denominator, and the HPA reports `<unknown>`. What it does not do: it does not change a pod's CPU/memory sizing, it does not create nodes, and it fights anything else that writes `replicas` on the same object — so drop the static `replicas:` field from a manifest that GitOps keeps reapplying.
go deeper
Be able to say what it scales, the metric-target idea, and the two prerequisites: metrics-server and resource requests.
Add the mechanics — it patches the scale subresource on a ~15s loop, min/max clamp the result, and Utilization is a percentage of the request.
Frame it as one layer of a stack: HPA for replicas, VPA for sizing, cluster autoscaler for nodes, and name the conflicts between them and with GitOps ownership of replicas.
Discuss when replica scaling is the right control at all versus queue-based shedding or capacity reservation, and who owns the replica field across the delivery pipeline.
## Horizontal versus vertical "Horizontal" scaling changes the **number of identical replicas** of a workload. "Vertical" scaling changes how much CPU and memory a *single* replica gets. A HorizontalPodAutoscaler only does the first — it never edits a container's `resources` block. ## The object and the loop An HPA is an ordinary Kubernetes resource in the `autoscaling/v2` API group (GA since Kubernetes 1.23; the older `v2beta2` was removed in 1.26). Its key fields are: - `scaleTargetRef` — which workload to resize. It must expose the `scale` subresource: Deployment, StatefulSet, ReplicaSet, ReplicationController, or a custom resource that implements it. **DaemonSets cannot be targeted** — they run one pod per node by design. - `minReplicas` / `maxReplicas` — the hard floor and ceiling. The controller never leaves this range. - `metrics[]` — one or more metric specs. The commonest is a `Resource` metric on `cpu` with `target.type: Utilization` and `averageUtilization: 70`. The HPA controller runs inside kube-controller-manager on a sync loop, by default every 15 seconds (`--horizontal-pod-autoscaler-sync-period`). Each tick it lists the pods matching the target's selector, fetches their metrics, computes a desired replica count, and if it differs from the current one it PATCHes the target's `scale` subresource. The Deployment controller then does the actual pod creation or deletion — the HPA itself never touches pods. ## Why requests are mandatory With `target.type: Utilization`, the metric is not "CPU cores used". It is **used ÷ requested, as a percentage**. A pod requesting 200m and burning 100m is at 50% utilization. If the container has no CPU request at all, there is no denominator, the utilization is undefined, and `kubectl get hpa` shows `<unknown>/70%` while the replica count sits still. This single omission is the most common reason a freshly written HPA appears to do nothing. The alternative `target.type: AverageValue` sidesteps requests — you state an absolute per-pod number such as `200m` — but for CPU the utilization form is the conventional choice because it stays correct when you resize the pods. ## Where the metrics come from CPU and memory arrive through the **resource metrics API**, `metrics.k8s.io`, which is served by the **metrics-server** deployment. metrics-server scrapes each kubelet's summary endpoint every ~15s and keeps only the latest sample in memory — it is a live sensor, not a time-series database, and it is deliberately not suitable for historical queries. Managed distributions usually pre-install it; a kubeadm or kind cluster does not, and without it both `kubectl top` and every HPA fail together, which is a fast way to confirm the diagnosis. ## What it is not - **Not the Vertical Pod Autoscaler.** VPA adjusts requests/limits; running both on CPU for the same workload is a known conflict. - **Not the Cluster Autoscaler.** If the HPA asks for 20 replicas and the nodes are full, the extra pods sit `Pending` until node-level autoscaling provides capacity. The two are complementary layers, not substitutes. - **Not a scheduler.** The HPA has no say in *where* pods land. - **Not compatible with a hardcoded replica count.** Leaving `replicas: 3` in a manifest that Argo CD or Flux reconciles produces a tug-of-war: the HPA scales to 9, the sync scales back to 3. Remove the field (or mark it ignored) and let the HPA own it. ## Reading its state `kubectl get hpa` shows `TARGETS` as `current/target`, plus min, max and current replicas. `kubectl describe hpa` prints the conditions — `AbleToScale`, `ScalingActive`, `ScalingLimited` — and recent events explaining each decision. `ScalingLimited: True` means the desired count was clipped by min or max, which is usually the answer to "why did it stop at 10?".
- Your HPA and your GitOps tool disagree about the replica count. What is happening and how do you fix it?Both are writing `.spec.replicas` on the same Deployment: the HPA scales up, the next reconcile resets it to the committed value, and the workload oscillates. The fix is to make the HPA the sole owner — delete the `replicas` field from the manifest entirely, or tell the GitOps controller to ignore that path (for example Argo CD's `ignoreDifferences` on `/spec/replicas`).
- Can you attach a HorizontalPodAutoscaler to a DaemonSet?No. A DaemonSet has no `scale` subresource because its replica count is defined by the set of matching nodes, not by a number you choose. If you need per-node capacity to grow, that is a node-count or resource-sizing problem, not a horizontal pod scaling one.
A thermostat wired to the number of radiators in a room, not to how hot each one runs: it can switch radiators on and off, but it cannot make one hotter, and it cannot build you a bigger house.
saying these in an interview costs you the question
- Believing the HPA changes CPU/memory limits on pods
- Assuming metrics-server ships with every cluster
- Writing a Utilization target on containers that declare no resource requests
- Thinking the HPA creates nodes when pods go Pending
- Leaving a hardcoded replicas value in a manifest that a GitOps controller reapplies