A Kubernetes HorizontalPodAutoscaler is set to an average CPU utilization target of 50%. Walk through exactly how the controller turns the observed metric into a new replica count.
answer
- ceil(current × current/target)
- denominator is the request, not the limit
- 10% tolerance dead-band
- missing/unready pods biased against over-reacting
- multiple metrics → take the max
basics
~20 sdesiredReplicas = ceil(currentReplicas × currentMetric ÷ targetMetric). currentMetric is the mean across ready pods of (usage ÷ request) as a percentage. If the ratio is within a tolerance of 1 (10% by default) nothing happens, and the result is clamped to minReplicas/maxReplicas.
solid answer
~60 sThe core formula is: ``` desiredReplicas = ceil( currentReplicas × ( currentMetricValue / desiredMetricValue ) ) ``` For a `Resource` metric with `type: Utilization`, `currentMetricValue` is the **average across ready pods of usage divided by that pod's request**, expressed as a percentage — so 6 pods requesting 200m and each burning 150m is 75%. Against a 50% target with 6 replicas: `ceil(6 × 75/50) = 9`. Three guards wrap that arithmetic: 1. **Tolerance.** If `|ratio − 1|` is under the tolerance (default 0.1, i.e. 10%) the controller does nothing. This kills constant one-pod jitter. 2. **Missing and not-ready pods.** Pods with no metrics or not yet Ready are set aside, then reintroduced conservatively: missing pods count as 0% when the decision is to scale up and 100% when it is to scale down, and unready pods count as 0% for scale-up. The bias is always toward *not* over-reacting. 3. **Clamping and behavior.** The result is bounded by `minReplicas`/`maxReplicas`, then filtered through the `behavior` stabilization windows and rate policies. With multiple metrics the controller computes a desired count per metric and takes the **maximum**.
code
yaml · 18 linesapiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: api
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: api
minReplicas: 3
maxReplicas: 30
metrics:
- type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 50go deeper
Recall the formula and that utilization is a percentage of the request; a worked two-line example is enough.
Add the tolerance dead-band, the ceiling, min/max clamping, and the max-across-metrics rule.
Explain the conservative handling of unready/missing pods and why request sizing and the target are a single coupled tuning problem.
Reason about the control loop's stability — proportional gain, sample latency, the dead-band — and about which signal actually correlates with user-visible degradation.
## The formula Every sync (default 15 seconds) the HorizontalPodAutoscaler controller computes: ``` desiredReplicas = ceil( currentReplicas × ( currentMetricValue / desiredMetricValue ) ) ``` The ceiling matters: fractional pods do not exist, and rounding up biases toward availability. `currentReplicas` is read from the target's `scale` subresource, not counted from live pods. ## What currentMetricValue is, per metric type - **Resource + Utilization** — for each pod, `usage ÷ request` for that resource; the controller averages those ratios over the counted pods and expresses it as a percentage. Crucially the denominator is the **request**, not the limit and not node capacity. Raising a request without changing actual usage *lowers* reported utilization and quietly scales the workload down. - **Resource + AverageValue** — the raw average usage per pod, e.g. `300m`, with no request involved. - **ContainerResource** — same as Resource but scoped to one named container in the pod, which is how you stop a sidecar's idle CPU from diluting the app container's signal. - **Pods** — an average across pods of a custom metric such as `http_requests_per_second`. - **Object / External** — a single value describing one object (an Ingress's request rate) or something outside the cluster (a queue depth). For `External` with `type: Value` the value is compared as a whole; with `AverageValue` it is divided by the replica count first — that division is what makes queue-depth scaling work. ## Worked example 4 replicas, each container requests 500m CPU, each pod is using 400m. Utilization = 400/500 = 80%. Target 50%. `ratio = 80/50 = 1.6` → outside the 10% tolerance → `ceil(4 × 1.6) = 7` replicas. Next tick, with load unchanged and spread over 7 pods, each uses ~229m → 46% → `ratio = 0.92` → **inside** the tolerance → no change. The system settles. ## Tolerance The default tolerance is 0.1: no action while the ratio sits between 0.9 and 1.1. Historically this was a single cluster-wide flag (`--horizontal-pod-autoscaler-tolerance`) that nobody could tune per workload; recent Kubernetes releases add per-HPA `behavior.scaleUp/scaleDown.tolerance` fields, so state your version if you claim you can set it per object. Tolerance is why a metric hovering just above target produces no scaling and why very small replica counts move in visible jumps — at 2 replicas the smallest possible step is a 50% capacity change. ## Pods that are missing, unready, or terminating The controller does not blindly average whatever it received: - Pods that are **terminating** are ignored outright. - Pods with **no metric sample** (just started, metrics-server hasn't scraped them) are excluded from the first computation, then added back with an assumption chosen to be conservative: 0% usage if the tentative decision was to scale up, 100% if it was to scale down. Then the result is recomputed. This prevents a burst of new pods from immediately triggering another scale-up. - Pods that are **not Ready** count as 0% when scaling up, so a slow-starting replica cannot inflate the apparent shortage. The net effect is a deliberate asymmetry: the HPA is eager to have enough capacity but skeptical about removing it. ## Multiple metrics List several entries under `metrics` and the controller evaluates each independently, then takes the **largest** desired count. It is an OR of pressure signals — any one metric can drive a scale-up, but a scale-down requires *all* of them to agree. If one metric is unavailable, scale-up on the others still proceeds while scale-down is blocked, again erring toward capacity. ## After the arithmetic The number is clamped into `[minReplicas, maxReplicas]` — hitting a bound sets the `ScalingLimited` condition, visible in `kubectl describe hpa`. Then the `behavior` block applies stabilization windows and rate policies, which can delay or shrink the move. Only then is the `scale` subresource patched. ## Consequences worth stating in an interview Because the denominator is the request, **request sizing and autoscaling are one problem, not two**. Oversized requests mean utilization never reaches target and the workload never scales out even while latency degrades; undersized requests mean it scales out constantly and each pod is throttled. And because the formula is proportional to *current* replicas, a workload pinned at `minReplicas: 1` under a sudden spike can only reach `ceil(1 × ratio)` per step, which is why cold-start-sensitive services set a higher floor rather than relying on fast reaction.
- A team doubles the CPU request on their pods without changing the code. What happens to a utilization-based HorizontalPodAutoscaler on that workload?Reported utilization halves, because utilization is usage divided by request. The ratio to target drops, so the HPA scales the workload down — possibly to minReplicas — even though real CPU consumption is identical. Request sizing and the utilization target must be tuned together; changing one silently re-tunes the other.
- An HPA lists both a CPU metric and a custom requests-per-second metric. How is the final replica count chosen?The controller computes a desired count from each metric independently and takes the maximum. So any single metric under pressure can scale the workload up, while scaling down requires every metric to be below its target. If one metric cannot be fetched, scale-down is suppressed while scale-up on the remaining metrics still works.
saying these in an interview costs you the question
- Saying utilization is measured against the CPU limit or against node capacity
- Expecting immediate scaling for any deviation, forgetting the 10% tolerance dead-band
- Thinking multiple metrics are averaged rather than maxed
- Forgetting the ceiling, and predicting fractional or rounded-down replica counts
- Ignoring that unready and metric-less pods are handled specially rather than averaged in