skip to content

What are the update modes of the Kubernetes Vertical Pod Autoscaler, and what does each one actually do to a running Pod?

level: middleimportance: should knowfreq 44%

answer

  1. recommender + updater + admission webhook
  2. Off = advice only; Initial = at creation; Auto/Recreate = evict
  3. VPA sizes Pods, HPA counts Pods
  4. minAllowed / maxAllowed / controlledResources / controlledValues
  5. min-replicas 2 before the updater will evict

basics

~20 s

Off only publishes recommendations for humans to read. Initial applies them when a Pod is created and never again. Auto (and Recreate) additionally evicts running Pods so they are recreated with new resource requests. VPA changes requests per Pod; it never changes replica count.

solid answer

~60 s

VPA sets **CPU and memory requests** on containers, based on observed usage. It ships as three components: a **recommender** that computes target, lower and upper bounds from usage history; an **updater** that acts on Pods outside those bounds; and an **admission controller** webhook that rewrites requests as Pods are created. `spec.updatePolicy.updateMode`: - **Off** — recommender only. Recommendations appear in `status.recommendation`; nothing mutates. This is the safe way to run VPA in production and the mode most teams should start with. - **Initial** — the webhook applies the recommendation at Pod creation only. Existing Pods are untouched, so changes land on your normal rollout cadence. - **Recreate** — the updater evicts Pods whose requests drift far from the recommendation; the controller recreates them and the webhook stamps in the new values. - **Auto** — currently behaves as Recreate; it is defined as "use the best available mechanism", so it is expected to move to in-place resizing as that matures. The disruptive point: for a long time the only way to change a Pod's requests was to replace the Pod.

code

yaml · 24 lines
yaml
apiVersion: autoscaling.k8s.io/v1
kind: VerticalPodAutoscaler
metadata:
  name: checkout-vpa
spec:
  targetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: checkout
  updatePolicy:
    updateMode: "Off"
  resourcePolicy:
    containerPolicies:
    - containerName: app
      minAllowed:
        cpu: "100m"
        memory: "256Mi"
      maxAllowed:
        cpu: "2"
        memory: "4Gi"
      controlledResources: ["cpu", "memory"]
      controlledValues: RequestsOnly
    - containerName: istio-proxy
      mode: "Off"

go deeper

for a junior

Name the modes and state that VPA adjusts CPU and memory requests rather than replica count.

for a middle

Explain the three components, exactly what each mode does to a running Pod, and the resourcePolicy knobs such as minAllowed, maxAllowed and controlledValues.

for a senior

Discuss the disruption cost of Auto, peak-driven memory recommendations, the min-replicas guard, PDB interaction, and starting in Off mode to build trust.

for a principal

Position it as a fleet efficiency programme: recommendation-only across the estate, guardrails through resource policies, and the link between right-sized requests and node scale-down as the actual source of savings.

## The problem VPA solves Container `resources.requests` drive everything in Kubernetes: which node a Pod fits on, its QoS class, and whether node autoscaling considers a node full. Humans set them badly — usually far too high "to be safe", which wastes money and blocks node scale-down, or too low, which causes CPU throttling and OOM kills. VPA sets them from measured usage instead. It is orthogonal to horizontal autoscaling: horizontal scaling changes how many Pods there are, VPA changes how big each one is. ## The three components **Recommender.** Reads current usage from the metrics API and historical usage from Prometheus or its own checkpoints, and maintains a decaying histogram per container. From that it derives a `target` (roughly a high percentile of CPU, a peak-oriented figure for memory plus a safety margin), plus `lowerBound` and `upperBound` — the band outside which the updater is allowed to act. Recommendations appear on the VPA object's `status`, so you can read them without applying anything. **Updater.** Watches running Pods. If a Pod's requests are outside the recommended bounds for long enough, it evicts the Pod through the Eviction API. It does not create the replacement — the owning Deployment or StatefulSet does. It respects PodDisruptionBudgets and by default refuses to act on a controller with fewer than two replicas (`--min-replicas`), because evicting the only replica means downtime. **Admission controller.** A mutating webhook on Pod creation that overwrites the requests in the incoming Pod spec with the current recommendation. This is the component that actually applies values; without it, an eviction would just recreate the Pod with its original requests. ## The modes, precisely - **`Off`** — recommender runs, nothing else. No webhook mutation, no eviction. Use it to gather data, to feed a dashboard, or to drive an offline right-sizing process where a human edits the manifests. It is completely safe and is the correct first deployment. - **`Initial`** — the webhook applies recommendations to newly created Pods; the updater never evicts. Values therefore change only when Pods are recreated for other reasons: a rollout, a node drain, a crash. Predictable, non-disruptive, but slow to converge. - **`Recreate`** — updater evicts drifted Pods, webhook stamps new values on recreation. Converges quickly at the cost of restarts. - **`Auto`** — documented as "whichever mechanism is best available"; today that is eviction and recreation, identical to Recreate. Do not assume it is non-disruptive because of the name. Newer VPA releases, paired with in-place Pod resize support in recent Kubernetes versions, add a mode that tries to patch a running Pod's resources without restarting it and falls back to recreation when the change cannot be applied in place — for example when memory must shrink. Treat this as an emerging capability and check what your cluster version actually supports rather than assuming it. ## The resource policy `spec.resourcePolicy.containerPolicies` constrains the recommender per container: - `minAllowed` / `maxAllowed` — clamp the recommendation. `maxAllowed` matters because an unclamped recommendation can exceed any node's capacity, producing a Pod that cannot be scheduled at all. - `controlledResources` — `["cpu"]`, `["memory"]`, or both. Restricting VPA to memory only is a common way to coexist with CPU-based horizontal scaling. - `controlledValues` — `RequestsAndLimits` (default) scales limits proportionally with requests, preserving the original ratio; `RequestsOnly` leaves limits untouched. - `mode: "Off"` per container excludes a sidecar you do not want touched. `spec.targetRef` points at the controller (Deployment, StatefulSet, DaemonSet), not at Pods. ## Operational cautions - **Restart cost is real.** Auto and Recreate mean rolling restarts triggered by a controller you do not directly drive. Long-lived connections, warm caches, and slow starts all pay. - **Memory recommendations are peak-driven.** A once-a-month batch job can raise a service's memory request permanently, since the histogram decays slowly. - **New workloads have no history.** Early recommendations after deployment are unreliable; VPA needs a meaningful observation window. - **It requires metrics-server** (or an equivalent metrics API) to be running. - **It changes scheduling, not just cost.** Larger requests can make Pods unschedulable on existing nodes and trigger node scale-up; smaller requests free nodes for scale-down, which is often the real saving.

  • Why is updateMode Off still useful if it never changes anything?
    It gives you measured target, lower-bound and upper-bound requests per container with no risk of disruption. Teams use it to right-size manifests by hand, to feed cost dashboards, and to build confidence before enabling an applying mode. It is also the only mode that is safe on single-replica or restart-sensitive workloads.
  • What does controlledValues RequestsOnly change compared with the default?
    By default VPA scales limits together with requests, keeping the original request-to-limit ratio, which can silently raise or lower a container's memory limit. RequestsOnly leaves the limits exactly as authored and moves only the requests. That is usually what you want when limits encode a deliberate ceiling, for example an OOM guard tuned to the application's heap settings.

Off is a tailor measuring you and handing you the numbers. Initial is the tailor cutting your next suit to those numbers. Auto is the tailor taking the suit off you today and giving you a new one.

saying these in an interview costs you the question

  • Thinking VPA changes the number of replicas — that is horizontal scaling; VPA changes per-Pod requests.
  • Assuming Auto applies changes without restarting Pods; in the classic implementation it evicts and recreates them.
  • Believing Off mode still mutates Pods — it only publishes recommendations to the VPA status.
  • Forgetting that the admission webhook is what actually writes the values, so eviction alone would change nothing.
  • Leaving maxAllowed unset and ending up with a recommendation larger than any node can satisfy.

context