Why is running the Kubernetes Vertical Pod Autoscaler in Auto mode alongside a Horizontal Pod Autoscaler that targets CPU utilization considered unsafe, and what can you do instead?
answer
- HPA Utilization = usage / request
- VPA moves the denominator
- loop: smaller request -> higher utilization -> scale out -> lower usage
- allowed: HPA on custom/external metrics
- standard split: VPA memory, HPA CPU
basics
~20 sCPU-based horizontal scaling measures usage divided by the request. VPA moves the request, so the same real load changes the measured utilization and makes the horizontal controller scale in or out for no workload reason. The two form a feedback loop. Split them: VPA on memory only, horizontal scaling on CPU or on custom metrics.
solid answer
~60 sAn HPA targeting `Resource: cpu` with `type: Utilization` computes **usage ÷ request** per Pod and compares it to the target. That denominator is exactly what VPA rewrites. Suppose real usage is 300m against a 1000m request: utilization is 30%, well under a 70% target. VPA observes the same usage, decides 400m is enough, and lowers the request. Utilization is now 75% at unchanged load, so the HPA scales out. More replicas means less load per Pod, so usage per Pod drops, so VPA recommends smaller requests again — and round it goes. The two controllers are steering the same ratio from opposite ends with no shared state, so the system oscillates and each VPA cycle also restarts Pods. Options: (1) run VPA in `Off` mode and right-size by hand; (2) restrict VPA with `controlledResources: ["memory"]` while the HPA owns CPU — the standard split; (3) have the HPA target a custom or external metric such as requests per second or queue depth, which does not involve the request value at all.
code
yaml · 34 linesapiVersion: autoscaling.k8s.io/v1
kind: VerticalPodAutoscaler
metadata:
name: api-vpa
spec:
targetRef:
apiVersion: apps/v1
kind: Deployment
name: api
updatePolicy:
updateMode: "Auto"
resourcePolicy:
containerPolicies:
- containerName: app
controlledResources: ["memory"]
---
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: api-hpa
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: api
minReplicas: 3
maxReplicas: 30
metrics:
- type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 70go deeper
Know that VPA changes a Pod's requests, that CPU-based horizontal scaling is measured against those requests, and that combining them is discouraged.
Derive the loop concretely with numbers and name the standard fix of restricting VPA to memory while the HPA owns CPU.
Add the restart cost of Auto mode, the CPU-shares consequence of shrinking requests, the custom-metrics escape route, and how to recognise the oscillation from HPA events and Pod ages.
Decide the estate-wide policy: recommendation-only VPA with manual right-sizing versus a demand-based horizontal signal, and who owns requests as a platform contract versus a team choice.
## The mechanism, precisely A Horizontal Pod Autoscaler with a `Resource` metric of type `Utilization` does not read raw CPU. For each Pod it computes `usage / request` for the containers it tracks, averages that across ready Pods, and applies the ratio formula: `desiredReplicas = ceil(currentReplicas × currentValue / targetValue)`. The request in that denominator is a **declared** value, not a measured one. VPA's entire job is to change declared requests. So the HPA's input signal moves whenever VPA acts, even if the workload is perfectly steady. ## Walking the loop 1. Steady traffic. Each Pod uses 300m CPU, requests 1000m. Utilization 30%, HPA target 70% — no action. 2. VPA sees consistent 300m usage against a 1000m request and recommends roughly 400m. 3. In Auto mode the updater evicts the Pods; the webhook recreates them with `requests.cpu: 400m`. 4. Same 300m usage, new denominator: utilization is now 75%. Over target, so the HPA scales out. 5. The extra replicas share the same total traffic, so per-Pod usage falls to, say, 200m. 6. VPA now sees 200m and recommends smaller still. Utilization climbs relative to the new request, HPA adds more replicas. 7. Eventually traffic drops, per-Pod usage collapses, the HPA scales in hard, the surviving Pods spike over their now-tiny requests, get throttled, and VPA reacts again. Second-order damage matters as much as the oscillation. Every VPA action in Auto mode is an **eviction**, so this loop produces continuous rolling restarts. Restarts churn connections, cold caches and JVM warm-up, which itself perturbs CPU usage and feeds the loop further. And a request driven very low makes the Pod a poor citizen: cgroup CPU shares are proportional to the request, so it loses arbitration under contention exactly when it needs CPU. ## Official position The VPA documentation states directly that VPA should not be used with an HPA on CPU or memory. It is explicitly compatible with an HPA driven by **custom or external metrics**, because those never reference the resource request. ## The three practical resolutions **1. VPA in `Off` mode.** The recommender still publishes `status.recommendation`; nobody mutates anything. Engineers, or a periodic job that opens a pull request, update the manifests. The HPA keeps a stable denominator. This is the lowest-risk arrangement and is what most mature platforms actually run. **2. Split the resources.** Set `resourcePolicy.containerPolicies[].controlledResources: ["memory"]` so VPA owns memory only, and let the HPA own CPU. This is the widely used split and it is sound, because memory requests are not an HPA input when the HPA targets CPU. Two cautions: memory changes still require a Pod restart in the classic implementation, and if the HPA also has a memory-utilization metric the same conflict returns for memory. Some managed platforms package this split as a supported product feature. **3. Move the HPA off resource utilization.** Target requests per second, queue depth, or another custom or external metric. This is the cleanest fix because horizontal scaling then follows demand directly rather than a proxy that happens to be defined in terms of the request. It costs a metrics pipeline — an adapter serving `custom.metrics.k8s.io` or `external.metrics.k8s.io`. A fourth, weaker option is to use `type: AverageValue` instead of `Utilization` on the CPU metric, which compares absolute millicores against a fixed number and therefore does not divide by the request. It removes the direct coupling, though you then have to keep that absolute target in step with the Pod size VPA is choosing, which reintroduces the coupling in human form. ## Diagnosing it in the wild Symptoms: replica count sawtooths with no matching change in traffic; `kubectl describe hpa` shows repeated SuccessfulRescale events in both directions; Pod ages are uniformly short; the VPA object shows recommendations moving each cycle. Correlate the HPA's `currentCPUUtilizationPercentage` against actual request rate — if utilization moves while request rate does not, something is moving the denominator. ## How to answer Do not just say "they conflict". Say **why**: utilization is a ratio whose denominator VPA owns. Then give the split — VPA on memory, HPA on CPU — or custom metrics, and mention that `Off` mode plus human right-sizing is a perfectly respectable production answer.
- Would the conflict disappear if the HPA used an absolute CPU target instead of a utilization percentage?The direct coupling would, because an AverageValue target compares measured millicores against a fixed number and never divides by the request. But you have then hard-coded an absolute figure that only makes sense for a particular Pod size, and VPA is changing exactly that size. It removes the automated feedback loop and replaces it with a manual one, so it is a mitigation rather than a clean fix.
- Is it safe to combine VPA in Auto mode with an HPA driven by requests per second?Yes, and this is the officially compatible combination. A custom or external metric such as requests per second never references resource requests, so VPA changing them does not move the horizontal controller's signal. You still carry the ordinary cost of Auto mode, namely Pod evictions and restarts whenever recommendations drift, so PodDisruptionBudgets and a sensible minimum replica count still matter.
It is a thermostat reading temperature as a percentage of a setpoint while a second device keeps quietly moving the setpoint. Nothing about the room changed, yet the heating keeps switching on and off.
saying these in an interview costs you the question
- Saying they conflict without being able to name the mechanism — the request is the denominator of the utilization ratio.
- Claiming VPA and HPA are always incompatible; VPA is explicitly fine with an HPA on custom or external metrics.
- Proposing to just lower the HPA target to compensate, which does not stop the feedback loop.
- Forgetting that Auto mode evicts Pods, so the conflict also produces continuous restarts, not only replica churn.
- Applying the memory-only split while the HPA also scales on memory utilization, which recreates the same conflict.