skip to content

In an Argo CD Application's spec.syncPolicy.automated block, what do prune: true and selfHeal: true each change, and what happens when each of them is left off?

level: middleimportance: must knowfreq 80%

answer

  1. two independent switches, both off by default
  2. one is about absence, one about divergence
  3. deleting from Git is not deleting from the cluster
  4. kubectl edit survives unless told otherwise
  5. allowEmpty guards the accidental empty render

basics

~20 s

prune lets automated sync delete live resources that were removed from Git; selfHeal lets it re-apply Git over changes made directly in the cluster. With both off, automated sync only applies new Git revisions and reports everything else as OutOfSync without touching it.

solid answer

~50 s

`syncPolicy.automated` means Argo CD syncs by itself when the rendered desired state changes, but by default it is deliberately timid. Deleting a manifest from Git does **not** delete the live object unless `prune: true` — otherwise the resource is reported as requiring pruning and the app stays `OutOfSync` forever. And a change made straight in the cluster with `kubectl edit` does **not** get reverted unless `selfHeal: true` — without it, Argo CD notices the drift, marks the app `OutOfSync`, and waits for a human. The two are independent: prune is about *absence* in Git, self-heal is about *divergence* in the cluster. Self-heal also changes what triggers a sync: normally automated sync fires on a new Git revision, whereas self-heal fires on the live-state diff, debounced by a few seconds so a fight with another controller does not become a hot loop. `automated.allowEmpty` (default false) is the guard that stops a bad commit that empties the path from pruning every resource.

code

yaml · 22 lines
yaml
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
  name: checkout
  namespace: argocd
spec:
  project: payments
  source:
    repoURL: https://github.com/acme/manifests.git
    path: apps/checkout/overlays/prod
    targetRevision: main
  destination:
    server: https://kubernetes.default.svc
    namespace: checkout
  syncPolicy:
    automated:
      prune: true
      selfHeal: true
      allowEmpty: false
    syncOptions:
      - CreateNamespace=true
      - PruneLast=true

go deeper

for a junior

Know that enabling automated sync alone does not delete anything and does not undo manual cluster edits, and that prune and selfHeal are the two separate opt-ins that add those behaviours.

for a middle

Explain the different triggers — new source revision versus live-state drift — the debounce behind self-heal, and why allowEmpty exists as a guard on a bad render.

for a senior

Show how you would roll these on safely: resource tracking and its adoption hazards, PruneLast and per-resource Prune=false for stateful objects, and how you diagnose a self-heal fight with another controller.

for a principal

Own the policy: which classes of workload may run fully automated, what must stay gated, and how you keep an irreversible prune from ever being one bad merge away across an estate of applications.

## The three switches ```yaml spec: syncPolicy: automated: prune: true selfHeal: true allowEmpty: false ``` Adding `automated` at all is the first decision: it tells the application controller to run a sync on its own when the desired state changes, instead of waiting for `argocd app sync` or a click in the UI. What surprises people is how narrow that promise is by default. ## prune — acting on what is missing from Git When a manifest disappears from the source (someone deleted the file, or an overlay stopped generating it), Argo CD can see that a live resource has no counterpart in the desired state. It refuses to delete it unless `prune: true`. This is a safety default: deletion is the one irreversible operation in the model, and an accidental delete of a `PersistentVolumeClaim` or a namespace is not a Git revert away. With `prune: false` the app sits `OutOfSync` with the resource marked as requiring pruning, and the UI offers a manual prune. Teams often discover this after the fact — an app that looks broken for weeks because a removed resource keeps it out of sync. How does Argo CD know a live object belongs to it at all? By a tracking marker it stamps on everything it applies — historically the `app.kubernetes.io/instance` label, and with `application.resourceTrackingMethod: annotation` the `argocd.argoproj.io/tracking-id` annotation. This matters in practice: if a resource carries a stale instance label from another app, prune can delete something you did not mean to give it. Label-based tracking is also why an object hand-labelled with an app name can be adopted (and then pruned) unexpectedly. Two related controls are worth knowing. `PruneLast=true` as a sync option defers pruning until everything else has been applied and is healthy, so you do not delete the old thing before the new thing works. `PrunePropagationPolicy=foreground|background|orphan` chooses the Kubernetes deletion propagation used when pruning. And a per-resource annotation `argocd.argoproj.io/sync-options: Prune=false` exempts one object from pruning entirely — the usual way to protect a database or a PVC that lives in the same app. ## selfHeal — acting on drift in the cluster `selfHeal: true` says: whenever live state diverges from the rendered desired state, re-apply. That is what makes Argo CD an actual guard against manual changes. Turn it on and `kubectl scale deployment/api --replicas=10` is undone within seconds; turn it off and Argo CD merely reports the drift. The subtlety is the *trigger*. Ordinary automated sync is driven by the source: a new commit produces new desired state, Argo CD syncs. Self-heal is driven by the live state, so it must react to cluster events, and it is debounced (a short timeout on the application controller, a few seconds by default, tunable) precisely so that a resource being fought over by two controllers does not produce a continuous sync loop. That fight is the classic self-heal incident: an HPA writes `spec.replicas`, Argo CD writes it back from Git, the HPA writes it again. The fix is not to disable self-heal but to stop claiming ownership of the field — remove `replicas` from the manifest, or list it under `ignoreDifferences` together with the `RespectIgnoreDifferences=true` sync option. Self-heal also has a hard limit: it repairs *modifications and deletions of tracked resources*. It does not stop the change from happening — Argo CD is not an admission controller. If you need the edit rejected rather than reverted, that is cluster policy, not GitOps. ## allowEmpty — the blast-radius guard With `prune: true`, a commit that accidentally empties the source path (a bad Kustomize edit, a chart that renders nothing) would prune every resource in the app. `automated.allowEmpty` defaults to `false`, which makes Argo CD refuse to sync to an empty desired state. Turning it on is rare and deliberate. ## Choosing the combination - **No automated policy** — Argo CD is a diff tool plus a big Sync button; suitable for environments where deploys are gated by a human. - **automated, no prune, no self-heal** — commits roll out, but nothing is ever deleted or reverted automatically. A common first step, and a common source of permanently out-of-sync apps. - **automated + prune + selfHeal** — the full statement that Git is authoritative. Sensible for stateless workloads in non-critical or well-tested environments, and increasingly for production once the team trusts its manifests. A good answer names the asymmetry: prune concerns absence, self-heal concerns divergence, and both are off by default because both can destroy something no commit can bring back.

  • With selfHeal enabled, what happens if an HPA and the Git manifest both set spec.replicas?
    They fight: the HPA scales, Argo CD reverts to the Git value, the HPA scales again, and the debounce only slows the loop. The fix is to stop owning the field — drop `replicas` from the manifest, or add `/spec/replicas` to `ignoreDifferences` plus the `RespectIgnoreDifferences=true` sync option so a sync does not re-apply it.
  • Does selfHeal prevent someone from running kubectl edit on a managed resource?
    No. Argo CD is a controller, not an admission plugin, so the edit is accepted by the API server and takes effect immediately; self-heal only reverts it moments later. If the change must be rejected outright you need cluster-side policy — RBAC that denies the write, or an admission policy — with self-heal as the backstop.
  • How would you protect a PersistentVolumeClaim inside an app that has prune enabled?
    Annotate the resource with `argocd.argoproj.io/sync-options: Prune=false`, which exempts that one object while the rest of the app still prunes normally. Adding `Delete=false` also keeps it when the Application itself is deleted with cascade. Many teams go further and manage stateful resources in a separate Application with no automated prune at all.
  • What does the PruneLast sync option change about the order of a sync?
    It defers all pruning until every other resource has been applied and reported healthy, so removals happen after the replacements are actually running. Without it, pruning is interleaved with the rest of the sync by wave and kind ordering, which can delete the old workload before the new one is serving.

saying these in an interview costs you the question

  • Assuming automated sync deletes removed manifests by default
  • Thinking selfHeal blocks the kubectl edit rather than reverting it
  • Believing prune and selfHeal are one setting
  • Claiming automated sync reverts drift without selfHeal
  • Saying a Git revert always undoes a prune

context