skip to content

What does the `cluster-autoscaler.kubernetes.io/safe-to-evict` annotation on a Kubernetes pod do, and when would you set it to `true` or `false`?

level: middleimportance: should knowfreq 44%

answer

  1. read by one component only
  2. drainability rules in the simulation
  3. local storage, kube-system, bare pods
  4. Eviction API still checks budgets
  5. a third value waits for completion

basics

~10 s

It is a per-pod override for Cluster Autoscaler's scale-down check: true lets the autoscaler remove the pod's node even if the pod would normally block it, and false pins the node against scale-down.

solid answer

~40 s

The annotation `cluster-autoscaler.kubernetes.io/safe-to-evict` is read only by Cluster Autoscaler when it decides whether a node can be emptied. With `"true"`, the pod counts as movable even if a default rule would block it: a bare pod with no controller, a pod with `hostPath` or disk-backed `emptyDir` storage, or an unbudgeted non-DaemonSet pod in `kube-system`. With `"false"`, the pod blocks scale-down of its node, which suits work that cannot be interrupted. `"on-completion"` makes the autoscaler wait for the pod to finish. It is not a PDB bypass: removal still goes through the Eviction API, which enforces budgets. I set `true` on stateless pods with disposable scratch data, and `false` sparingly, because each pinned pod keeps a whole node alive.

code

yaml · 15 lines
yaml
apiVersion: batch/v1
kind: Job
metadata:
  name: webhook-replay-2026-09-16
spec:
  backoffLimit: 0
  template:
    metadata:
      annotations:
        cluster-autoscaler.kubernetes.io/safe-to-evict: "false"
    spec:
      restartPolicy: Never
      containers:
        - name: replay
          image: registry.example.com/webhook-replay:1.7.0

go deeper

for a junior

Recall that this annotation is read only by Cluster Autoscaler, and that true means the pod may be moved while false pins its node.

for a middle

Explain which default blockers true overrides (local storage, unbudgeted kube-system pods, pods with no controller) and why removal still respects PodDisruptionBudgets.

for a senior

Show you would reach first for the local-volumes variant, audit every false in the cluster, and tie true to proof that the pod's data really is disposable.

for a principal

Treat false as a cost decision that strands whole nodes, and decide who may set it and where those pods are allowed to run.

## What the annotation is for **Cluster Autoscaler** removes a node only after it has simulated moving every pod off it. During that simulation it runs a fixed list of **drainability rules** over each pod on the node: some rules say "this pod can move", some say "this pod blocks the node", and most say nothing. The pod annotation `cluster-autoscaler.kubernetes.io/safe-to-evict` lets the workload owner answer that question directly instead of leaving it to the default rules. Only Cluster Autoscaler reads it. The kubelet, `kubectl drain`, kube-scheduler and the Eviction API ignore it, so it changes nothing about where the pod is scheduled or how the kubelet treats it under node pressure. ## The values | Value | Effect on Cluster Autoscaler scale-down | |---|---| | `"true"` | The pod counts as movable, even when a default rule would block it | | `"false"` | The pod blocks removal of the node it runs on | | `"on-completion"` | The autoscaler waits for the pod to finish rather than evicting it | | absent | The default drainability rules decide | A narrower companion annotation, `cluster-autoscaler.kubernetes.io/safe-to-evict-local-volumes`, takes a comma-separated list of volume names. The listed volumes stop counting as blocking local storage, while every other blocker still applies. Prefer it when local storage is the only reason a pod blocks. ## What `"true"` overrides The default rules that `"true"` overrides are the classic scale-down blockers: - a pod with **local storage**, meaning a `hostPath` volume or an `emptyDir` that is not memory-backed, while `--skip-nodes-with-local-storage` is `true` (the default); - a non-DaemonSet, non-mirror pod in **`kube-system`** that no PodDisruptionBudget covers, while `--skip-nodes-with-system-pods` is `true` (the default); - a pod with **no controller at all**, or one whose owner is a kind the autoscaler does not recognise while `--skip-nodes-with-custom-controller-pods` is `true` (the default). ## What `"true"` does not do 1. **It does not bypass a PodDisruptionBudget when the pod is actually removed.** Cluster Autoscaler removes pods through the Eviction API, and the API server refuses an eviction the budget does not allow. A budget with zero allowed disruptions can still stop the drain. 2. **It does not make the pod fit somewhere else.** The simulation still has to place the pod on another node that satisfies its requests, affinity and topology rules. If no node fits, the node stays. 3. **It does not preserve data.** An evicted pod's `emptyDir` contents are deleted with the pod, and a bare pod with no controller is simply gone, because nothing recreates it. ## When to set each value - **`"true"`** on stateless replicas whose local scratch data is disposable: caches, spool directories that are flushed on `SIGTERM`, build scratch space. - **`"false"`** on work that cannot be interrupted and cannot resume, such as a long single-pass migration. Use it sparingly: every pinned pod keeps a whole node alive, however empty that node is. - **`"on-completion"`** on run-to-finish pods where waiting is acceptable and killing is not. - **Nothing** when the defaults already describe the pod correctly. Adding `"true"` by reflex hides real data loss. To find where the annotation is already in use, list pods with their annotations and search for the key. Every `"false"` in the output is a node that scale-down will never touch while that pod runs, so treat the list as a cost report as well as a configuration check. ## Worked example: a webhook-delivery dispatcher A webhook-delivery dispatcher runs in a 1,180-pod namespace. Each pod mounts a disk-backed `emptyDir` as a retry spool and drains that spool to a durable queue in its `preStop` hook. After a delivery burst, Cluster Autoscaler leaves 23 lightly loaded nodes in place, and its logs name local storage as the blocker on each one. The spool is disposable once flushed, so the team marks that one volume as non-blocking in the pod template: ```yaml apiVersion: apps/v1 kind: Deployment metadata: name: webhook-dispatcher spec: replicas: 1120 selector: matchLabels: app: webhook-dispatcher template: metadata: labels: app: webhook-dispatcher annotations: cluster-autoscaler.kubernetes.io/safe-to-evict-local-volumes: "retry-spool" spec: terminationGracePeriodSeconds: 45 containers: - name: dispatcher image: registry.example.com/webhook-dispatcher:4.12.3 volumeMounts: - name: retry-spool mountPath: /var/spool/dispatch volumes: - name: retry-spool emptyDir: sizeLimit: 2Gi ``` On the next scale-down pass, those nodes become candidates again. The dispatcher's PodDisruptionBudget still paces how many of its pods are evicted at once. The one-off reconciliation Job that replays a day of failed deliveries, which cannot resume midway, gets `"false"` instead, so its node stays until the Job ends.

  • Why is `cluster-autoscaler.kubernetes.io/safe-to-evict-local-volumes` often a better choice than `safe-to-evict: "true"`?
    It is narrower. It lists specific volume names whose data is disposable, and only those volumes stop counting as blocking local storage. Every other rule still applies, so a bare pod or an unbudgeted `kube-system` pod still blocks. `safe-to-evict: "true"` overrides all of those default blockers at once, which can hide a pod that really should not be moved.
  • Does a memory-backed `emptyDir` block Cluster Autoscaler scale-down?
    No. The local-storage rule counts `hostPath` volumes and `emptyDir` volumes whose `medium` is not `Memory`. A `medium: Memory` volume is tmpfs and does not trip that rule, although its contents are still lost when the pod is evicted.

saying these in an interview costs you the question

  • The kubelet honours safe-to-evict during memory-pressure eviction
  • Setting it to true lets the pod ignore its PodDisruptionBudget
  • safe-to-evict false is a harmless default to add to every pod
  • The annotation also stops kubectl drain from evicting the pod
  • Memory-backed emptyDir volumes block scale-down just like disk-backed ones