A Kubernetes deployment's replica count oscillates all day — up 20, down 8, up 15. Which HorizontalPodAutoscaler settings control that, and how do the scaleUp and scaleDown stabilization windows actually work?
answer
- behavior.scaleUp / scaleDown
- down window default 300s, up window default 0
- max over window for down, min for up
- policies: Pods / Percent per periodSeconds
- selectPolicy Max | Min | Disabled
basics
~20 sThe behavior block. behavior.scaleDown.stabilizationWindowSeconds (default 300) makes the controller take the highest recommendation from the last window before shrinking; the scaleUp window (default 0) takes the lowest. Policies under each direction cap how many pods or what percentage may change per period. Widen the scale-down window and add a percent/pod policy to damp flapping.
solid answer
~50 sThrashing is controlled by `spec.behavior`, which has independent `scaleUp` and `scaleDown` sections. **Stabilization window.** The controller keeps a history of its recommendations. For **scale-down** it takes the **maximum** recommendation within `scaleDown.stabilizationWindowSeconds` (default **300s**) — so it only shrinks once the whole window agrees fewer pods are needed. For **scale-up** it takes the **minimum** within `scaleUp.stabilizationWindowSeconds`, which defaults to **0**, meaning immediate response. That asymmetry is intentional: react fast to load, retreat slowly. **Policies.** Each direction takes a list of rate limits, e.g. `type: Percent, value: 100, periodSeconds: 15` or `type: Pods, value: 4, periodSeconds: 15`, combined by `selectPolicy: Max` (default) or `Min`. Defaults let scale-up double every 15s or add 4 pods, whichever is larger; scale-down may drop 100% in one step *after* its 300s window. For a flapping workload: widen `scaleDown.stabilizationWindowSeconds` to 600–900, add `scaleDown` policy `Pods: 1 per 60s` with `selectPolicy: Min`, and check the metric itself isn't just noisy or sampled too finely.
code
yaml · 36 linesapiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: api
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: api
minReplicas: 4
maxReplicas: 60
metrics:
- type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 60
behavior:
scaleUp:
stabilizationWindowSeconds: 0
policies:
- type: Percent
value: 100
periodSeconds: 15
- type: Pods
value: 8
periodSeconds: 15
selectPolicy: Max
scaleDown:
stabilizationWindowSeconds: 900
policies:
- type: Pods
value: 1
periodSeconds: 120
selectPolicy: Mingo deeper
Know that behavior exists and that scale-down is deliberately slower than scale-up by default.
State the defaults — 300s down, 0s up — and how Pods/Percent policies cap the rate of change.
Explain max-over-window for down and min-over-window for up, pick concrete values for a flapping service, and identify slow warm-up or noisy metrics as the real cause.
Frame it as control-loop tuning against cost: how much idle capacity you buy to avoid oscillation, and where scale-down should be disabled outright.
## Why autoscalers oscillate An HPA is a proportional controller with a delay. Load rises, replicas are added, load per pod falls, replicas are removed, load per pod rises again — a classic feedback loop. It gets worse when the metric is noisy, when the workload has a slow start (new pods burn CPU warming caches or JIT-compiling, which *raises* the average and triggers more scale-up), or when the signal lags the action. ## The behavior block `spec.behavior` (autoscaling/v2) has two independent sub-objects, `scaleUp` and `scaleDown`, each with a stabilization window and a list of policies. ```yaml behavior: scaleUp: stabilizationWindowSeconds: 0 policies: - type: Percent value: 100 periodSeconds: 15 - type: Pods value: 4 periodSeconds: 15 selectPolicy: Max scaleDown: stabilizationWindowSeconds: 300 policies: - type: Percent value: 100 periodSeconds: 15 ``` That is the **default** behavior, worth memorizing because most questions are really "what changes if I leave it out". ## How the stabilization window works — precisely The controller records its raw recommendation every sync. When deciding a new replica count it does **not** use the latest recommendation directly: - **Scaling down**: it takes the **maximum** recommendation observed over the past `scaleDown.stabilizationWindowSeconds`. A single low reading cannot shrink the fleet; every reading in the window must be low before the maximum itself drops. - **Scaling up**: it takes the **minimum** recommendation over `scaleUp.stabilizationWindowSeconds`. With the default of 0 the window is empty and the current recommendation applies immediately. So the window is not a cooldown timer and not an average — it is a directional envelope. Understanding "max for down, min for up" is the crisp answer interviewers listen for. ## How policies work Each policy caps *change per period*: - `type: Pods, value: 4, periodSeconds: 60` — at most 4 pods added (or removed) in any 60-second window. - `type: Percent, value: 10, periodSeconds: 60` — at most 10% of the current replica count per 60 seconds. `selectPolicy` decides how multiple policies combine: `Max` (default) takes the most permissive, `Min` the most restrictive, and `selectPolicy: Disabled` on a direction **forbids scaling that way entirely** — the standard way to build a scale-up-only autoscaler where a human handles shrinking. Note that percent policies interact badly with tiny fleets: 10% of 3 replicas rounds to a change of 1 at best, so combine a percent policy with a pods policy for predictable behavior at both ends of the range. ## Recipes **Damp a flapping web service:** keep scale-up aggressive, widen scale-down. ```yaml behavior: scaleDown: stabilizationWindowSeconds: 900 policies: - type: Pods value: 1 periodSeconds: 120 selectPolicy: Min ``` This shrinks by one pod every two minutes and only after 15 minutes of consistently low demand — expensive in idle capacity, cheap in incidents. **Absorb a known traffic spike:** raise `scaleUp` percent to 200–300% per 15s and raise `minReplicas` ahead of the event, because proportional math from a low base is slow: from 2 replicas even a 100% policy needs several syncs to reach 30. **Freeze scale-down during business hours:** `scaleDown.selectPolicy: Disabled`, flipped by a scheduled process, is more predictable than a very long window. ## What behavior cannot fix Tuning windows treats the symptom. Also check: - **Slow-starting pods.** If a new replica spends 90 seconds at 100% CPU warming up, it inflates the average and drives further scale-up. Fix with readiness probes that gate traffic until warm, `ContainerResource` metrics scoped to the app container, and a startup-aware request size. - **Noisy or high-resolution metrics.** Autoscaling on a raw 15-second gauge invites oscillation; smooth it in the adapter query instead. - **Tolerance too tight** for the fleet size. At 2–3 replicas each step is a huge relative capacity change; a higher `minReplicas` makes the control loop better behaved. - **Disruption cost.** Rapid scale-down repeatedly terminates pods; combine with a PodDisruptionBudget and proper `terminationGracePeriodSeconds` so shrinking never drops in-flight work. ## Observing it `kubectl describe hpa` events log each change with the reason. If replica counts move but events are sparse, something other than the HPA is writing `spec.replicas`.
- What exactly does the scale-down stabilization window compute — an average, a cooldown, or something else?Neither an average nor a timer: the controller takes the maximum of all recommendations made within the window and scales down only to that value. A single dip in load therefore has no effect, because the older, higher recommendations still dominate the maximum until they age out of the window.
- How would you configure an HPA that may grow automatically but must never shrink on its own?Set `behavior.scaleDown.selectPolicy: Disabled`. That disables the scale-down direction entirely, so the controller only ever increases replicas; reductions then require a human or a separate scheduled process to lower the count or the max. It is a common pattern for workloads where terminating a pod is expensive or risky.
- New pods spike CPU while warming up, which triggers more scale-up. How do you stop the runaway?Treat it at the source: use readiness probes so warming pods take no traffic, and consider `ContainerResource` metrics so a busy sidecar or init-time work doesn't dominate. Then add a modest `scaleUp` rate policy so the fleet cannot double repeatedly within the warm-up period, and raise minReplicas so fewer cold starts are needed at all.
saying these in an interview costs you the question
- Describing the stabilization window as a simple cooldown timer or moving average
- Thinking scaleUp has a 300s window by default — it is 0
- Setting a long scale-up stabilization window to fight flapping, which delays reaction to real load
- Relying only on percent policies with 2–3 replicas, where rounding makes them meaningless
- Tuning behavior instead of fixing a noisy metric or a slow-warming pod