skip to content

Horizontal Pod Autoscaler

The Horizontal Pod Autoscaler recomputes a replica count from a metric against a target, inside min/max bounds, with stabilization windows so it stops flapping. Interviewers ask for the arithmetic and why it is meaningless unless requests are set.

part ofKubernetesoverview, primer and where to startread it →
on this pageshow

questions

5

What does a Kubernetes HorizontalPodAutoscaler object do, what must be true of the workload for it to work, and what does it deliberately not do?

level: juniorimportance: must knowfreq 70%

answer

  1. scale subresource, not pods directly
  2. metrics-server serves metrics.k8s.io
  3. Utilization = used ÷ requested
  4. min/max clamp, ScalingLimited condition
  5. no DaemonSets, no static replicas field

basics

~20 s

It watches a metric (usually CPU) for a Deployment or StatefulSet and rewrites that workload's replica count, clamped between minReplicas and maxReplicas, to keep the metric near a target. Needs metrics-server plus resource requests on the pods. It adds and removes pods only — it never resizes a pod and never adds nodes.

solid answer

~50 s

A **HorizontalPodAutoscaler (HPA)** is an API object plus a control loop in kube-controller-manager. It points at a scale target — a Deployment, StatefulSet or ReplicaSet, anything exposing the `scale` subresource — reads a metric roughly every 15 seconds, and writes a new `.spec.replicas` so the observed metric moves toward the target, clamped to `minReplicas`/`maxReplicas`. Prerequisites that bite people: - **A metrics source must exist.** CPU/memory targets need **metrics-server** serving the `metrics.k8s.io` API; it is not part of a vanilla cluster. - **Pods must declare resource requests** when the target type is `Utilization`, because utilization is a percentage *of the request*. No request, no denominator, and the HPA reports `<unknown>`. What it does not do: it does not change a pod's CPU/memory sizing, it does not create nodes, and it fights anything else that writes `replicas` on the same object — so drop the static `replicas:` field from a manifest that GitOps keeps reapplying.

go deeper

for a junior

Be able to say what it scales, the metric-target idea, and the two prerequisites: metrics-server and resource requests.

for a middle

Add the mechanics — it patches the scale subresource on a ~15s loop, min/max clamp the result, and Utilization is a percentage of the request.

for a senior

Frame it as one layer of a stack: HPA for replicas, VPA for sizing, cluster autoscaler for nodes, and name the conflicts between them and with GitOps ownership of replicas.

for a principal

Discuss when replica scaling is the right control at all versus queue-based shedding or capacity reservation, and who owns the replica field across the delivery pipeline.

## Horizontal versus vertical "Horizontal" scaling changes the **number of identical replicas** of a workload. "Vertical" scaling changes how much CPU and memory a *single* replica gets. A HorizontalPodAutoscaler only does the first — it never edits a container's `resources` block. ## The object and the loop An HPA is an ordinary Kubernetes resource in the `autoscaling/v2` API group (GA since Kubernetes 1.23; the older `v2beta2` was removed in 1.26). Its key fields are: - `scaleTargetRef` — which workload to resize. It must expose the `scale` subresource: Deployment, StatefulSet, ReplicaSet, ReplicationController, or a custom resource that implements it. **DaemonSets cannot be targeted** — they run one pod per node by design. - `minReplicas` / `maxReplicas` — the hard floor and ceiling. The controller never leaves this range. - `metrics[]` — one or more metric specs. The commonest is a `Resource` metric on `cpu` with `target.type: Utilization` and `averageUtilization: 70`. The HPA controller runs inside kube-controller-manager on a sync loop, by default every 15 seconds (`--horizontal-pod-autoscaler-sync-period`). Each tick it lists the pods matching the target's selector, fetches their metrics, computes a desired replica count, and if it differs from the current one it PATCHes the target's `scale` subresource. The Deployment controller then does the actual pod creation or deletion — the HPA itself never touches pods. ## Why requests are mandatory With `target.type: Utilization`, the metric is not "CPU cores used". It is **used ÷ requested, as a percentage**. A pod requesting 200m and burning 100m is at 50% utilization. If the container has no CPU request at all, there is no denominator, the utilization is undefined, and `kubectl get hpa` shows `<unknown>/70%` while the replica count sits still. This single omission is the most common reason a freshly written HPA appears to do nothing. The alternative `target.type: AverageValue` sidesteps requests — you state an absolute per-pod number such as `200m` — but for CPU the utilization form is the conventional choice because it stays correct when you resize the pods. ## Where the metrics come from CPU and memory arrive through the **resource metrics API**, `metrics.k8s.io`, which is served by the **metrics-server** deployment. metrics-server scrapes each kubelet's summary endpoint every ~15s and keeps only the latest sample in memory — it is a live sensor, not a time-series database, and it is deliberately not suitable for historical queries. Managed distributions usually pre-install it; a kubeadm or kind cluster does not, and without it both `kubectl top` and every HPA fail together, which is a fast way to confirm the diagnosis. ## What it is not - **Not the Vertical Pod Autoscaler.** VPA adjusts requests/limits; running both on CPU for the same workload is a known conflict. - **Not the Cluster Autoscaler.** If the HPA asks for 20 replicas and the nodes are full, the extra pods sit `Pending` until node-level autoscaling provides capacity. The two are complementary layers, not substitutes. - **Not a scheduler.** The HPA has no say in *where* pods land. - **Not compatible with a hardcoded replica count.** Leaving `replicas: 3` in a manifest that Argo CD or Flux reconciles produces a tug-of-war: the HPA scales to 9, the sync scales back to 3. Remove the field (or mark it ignored) and let the HPA own it. ## Reading its state `kubectl get hpa` shows `TARGETS` as `current/target`, plus min, max and current replicas. `kubectl describe hpa` prints the conditions — `AbleToScale`, `ScalingActive`, `ScalingLimited` — and recent events explaining each decision. `ScalingLimited: True` means the desired count was clipped by min or max, which is usually the answer to "why did it stop at 10?".

  • Your HPA and your GitOps tool disagree about the replica count. What is happening and how do you fix it?
    Both are writing `.spec.replicas` on the same Deployment: the HPA scales up, the next reconcile resets it to the committed value, and the workload oscillates. The fix is to make the HPA the sole owner — delete the `replicas` field from the manifest entirely, or tell the GitOps controller to ignore that path (for example Argo CD's `ignoreDifferences` on `/spec/replicas`).
  • Can you attach a HorizontalPodAutoscaler to a DaemonSet?
    No. A DaemonSet has no `scale` subresource because its replica count is defined by the set of matching nodes, not by a number you choose. If you need per-node capacity to grow, that is a node-count or resource-sizing problem, not a horizontal pod scaling one.

A thermostat wired to the number of radiators in a room, not to how hot each one runs: it can switch radiators on and off, but it cannot make one hotter, and it cannot build you a bigger house.

saying these in an interview costs you the question

  • Believing the HPA changes CPU/memory limits on pods
  • Assuming metrics-server ships with every cluster
  • Writing a Utilization target on containers that declare no resource requests
  • Thinking the HPA creates nodes when pods go Pending
  • Leaving a hardcoded replicas value in a manifest that a GitOps controller reapplies

context

open as a page

A Kubernetes HorizontalPodAutoscaler is set to an average CPU utilization target of 50%. Walk through exactly how the controller turns the observed metric into a new replica count.

level: middleimportance: must knowfreq 72%

basics

~20 s

desiredReplicas = ceil(currentReplicas × currentMetric ÷ targetMetric). currentMetric is the mean across ready pods of (usage ÷ request) as a percentage. If the ratio is within a tolerance of 1 (10% by default) nothing happens, and the result is clamped to minReplicas/maxReplicas.

open as a page

A Kubernetes HorizontalPodAutoscaler is not scaling and `kubectl get hpa` prints `<unknown>` in the TARGETS column. How do you diagnose it, and what are the usual root causes?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Check whether metrics exist at all with kubectl top pods. If that fails, metrics-server is missing or unhealthy (often TLS/kubelet-certificate or APIService issues). If top works but the HPA is unknown, the pods lack a resource request for the metric, the selector matches no ready pods, or a custom-metrics adapter is down. kubectl describe hpa names the failing condition.

open as a page

A Kubernetes deployment's replica count oscillates all day — up 20, down 8, up 15. Which HorizontalPodAutoscaler settings control that, and how do the scaleUp and scaleDown stabilization windows actually work?

level: seniorimportance: should knowfreq 44%

basics

~20 s

The behavior block. behavior.scaleDown.stabilizationWindowSeconds (default 300) makes the controller take the highest recommendation from the last window before shrinking; the scaleUp window (default 0) takes the lowest. Policies under each direction cap how many pods or what percentage may change per period. Widen the scale-down window and add a percent/pod policy to damp flapping.

open as a page

CPU is a poor proxy for load on a queue-consuming service. How do you drive a Kubernetes HorizontalPodAutoscaler from a non-CPU signal such as queue depth or requests per second, and what does that require in the cluster?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Use the autoscaling/v2 metric types Pods, Object or External instead of Resource. Those are served by custom.metrics.k8s.io and external.metrics.k8s.io, which core Kubernetes does not implement — you install an adapter (Prometheus Adapter, KEDA, a cloud provider's) that translates a query into that API. Use External with AverageValue for queue depth so the target means backlog per pod.

open as a page