skip to content

In Argo CD, what does the argocd.argoproj.io/sync-wave annotation change about a sync operation, and what does Argo CD do between one wave and the next?

level: middleimportance: should knowfreq 40%

answer

  1. an integer ordering key per resource
  2. default is zero, negatives go first
  3. the gate between waves is health
  4. one stuck wave blocks everything after
  5. only applies during a sync operation

basics

~20 s

The sync-wave annotation gives a resource an integer ordering key within a sync. Argo CD applies resources wave by wave, from the lowest number upward, and waits for the resources in each wave to report healthy before applying the next.

solid answer

~50 s

By default Argo CD applies everything in one pass, ordered by a built-in kind precedence (namespaces and CRDs before the workloads that use them) and then by name. The annotation `argocd.argoproj.io/sync-wave: "-1"` overrides that with an explicit integer — default `0`, negatives allowed and applied first. Argo CD groups the sync into waves, applies all resources of a wave, then **waits until they are healthy** before starting the next; a short delay between waves is also configurable. That waiting is the whole point: it is how you express "the database must be up before the API starts", or "install the CRD, then the custom resource that needs it". The cost is that a wave whose resources never become healthy stalls the entire sync — the app sits `Progressing` and later resources are never applied at all, which is a confusing failure until you know to look at the earliest unhealthy wave. Waves apply only during a sync operation; they say nothing about steady-state ordering.

code

yaml · 34 lines
yaml
apiVersion: batch/v1
kind: Job
metadata:
  name: db-migrate
  annotations:
    argocd.argoproj.io/sync-wave: "-1"
spec:
  template:
    spec:
      restartPolicy: Never
      containers:
        - name: migrate
          image: ghcr.io/acme/checkout:1.8.3
          command: ["/app/migrate", "up"]
---
apiVersion: apps/v1
kind: Deployment
metadata:
  name: checkout
  annotations:
    argocd.argoproj.io/sync-wave: "0"
spec:
  replicas: 3
  selector:
    matchLabels:
      app: checkout
  template:
    metadata:
      labels:
        app: checkout
    spec:
      containers:
        - name: api
          image: ghcr.io/acme/checkout:1.8.3

go deeper

for a junior

Know that Argo CD applies resources in an order, that the sync-wave annotation takes an integer with default 0, and that lower numbers are applied first.

for a middle

Explain the health gate between waves, how waves sit inside the sync phases, and give a real ordering need such as a migration Job before the Deployment or a CRD before its custom resource.

for a senior

Diagnose a stalled sync from the lowest unhealthy wave, and argue when application-level tolerance beats an explicit ordering; know the sync options that change apply and prune behaviour.

for a principal

Own the standard: how much ordering the platform encourages, since every wave boundary is a serialisation point and a place a fleet-wide sync can stall, and what teams should do instead.

## The default ordering A sync is not a single `kubectl apply` of a directory. Argo CD sorts the resources it is about to apply, using a built-in ordering by kind: namespaces, resource quotas and CRDs go before the objects that depend on them, workloads later, and within a kind by name. That default is enough for most applications and is why simple apps never need to think about ordering at all. ## When the default is not enough Some dependencies are semantic, not structural. A schema-migration Job must finish before the new Deployment starts. A secret-generating operator must be running before the custom resources that ask it for secrets. An external-DNS or cert-manager `Issuer` must exist before a `Certificate` referencing it. Kind ordering cannot know any of that. ## The annotation ```yaml metadata: annotations: argocd.argoproj.io/sync-wave: "-1" ``` Any integer, on any resource. The default is `0`, so a negative wave is the idiomatic way to say "before everything else" without renumbering the rest. Resources are grouped by wave and processed in ascending order; within a wave, the normal kind-and-name ordering still applies. The sequence Argo CD runs for each wave is: apply every resource in the wave, then wait until they are all assessed as healthy, then move on. There is also a short configurable pause between waves so that dependent controllers have a moment to react. That health gate is what makes waves useful — without it, waves would only reorder the applies, which rarely solves anything. Waves interact with the sync *phases* (the pre-sync, sync and post-sync stages that resource hooks run in): each phase has its own wave ordering. The hooks themselves are a separate mechanism; what matters here is that a wave number orders resources within the phase it belongs to. ## The failure mode If a resource in wave 1 never becomes healthy — a Job that keeps failing, a Deployment whose Pods crash-loop — Argo CD never starts wave 2. The Application reports `Progressing`, and the resources in later waves simply do not exist yet. Engineers who do not know about waves waste time asking why a manifest that is definitely in Git produced no object at all. The diagnostic instinct is: find the lowest wave containing an unhealthy resource; everything above it is blocked, not broken. That is also the reason to use waves sparingly. Every wave boundary is a place where the sync can stall, and heavy wave use turns a declarative apply into a fragile ordered script. Prefer to make workloads tolerate their dependencies being briefly absent — retry on startup, use an init container to wait — and reserve waves for orderings that genuinely cannot be expressed any other way. ## Sync options, the other half of sync behaviour Ordering is one of several knobs on how a sync executes. `spec.syncPolicy.syncOptions` (or the per-resource annotation `argocd.argoproj.io/sync-options`) carries flags such as: - `CreateNamespace=true` — create the destination namespace if it does not exist. - `PruneLast=true` — do all pruning after everything else is applied and healthy. - `PrunePropagationPolicy=foreground` — choose the Kubernetes deletion propagation used when pruning. - `ApplyOutOfSyncOnly=true` — apply only the resources that actually differ, which shortens syncs on very large apps. - `ServerSideApply=true` — apply with server-side apply, cooperating with other field managers and avoiding the client-side last-applied annotation size limit. - `Replace=true` — use replace instead of apply for resources that cannot be patched. - `Validate=false` — skip client-side schema validation. Set at the Application level they apply to the whole sync; set as a resource annotation they apply to one object, which is how you exempt a single PVC from pruning or force-replace one stubborn CRD. ## What to say in an interview Name the annotation and its default, say that the gate between waves is *health*, give one concrete ordering you have needed (migrations before the app, CRD before CR), and volunteer the failure mode — a stalled wave silently prevents everything after it. That last point is what separates someone who has read the docs from someone who has debugged it at two in the morning.

  • An Argo CD sync has been Progressing for twenty minutes and later resources were never created. Where do you look first?
    At the lowest wave that contains an unhealthy resource. Argo CD will not start the next wave until the current one is healthy, so a crash-looping Deployment or a failing Job in an early wave blocks everything after it. Fix or remove that resource, or reconsider whether the ordering needed a wave at all.
  • Why do teams over-using sync waves end up with fragile deployments?
    Because each wave boundary is a serialisation point that can stall, turning a declarative apply into an ordered script with no rollback. Slow syncs and mysterious partial states follow. Most orderings are better solved by making the workload tolerant — retries on startup, an init container that waits — and reserving waves for hard dependencies like a CRD before its custom resource.
  • What does the ApplyOutOfSyncOnly sync option do, and when does it matter?
    It restricts the sync to resources that actually differ from the desired state, instead of re-applying every object in the app. On applications with hundreds or thousands of resources this cuts sync duration and API-server load substantially. The tradeoff is that resources Argo CD believes are in sync are not re-applied, so a diff-invisible drift is not corrected.

saying these in an interview costs you the question

  • Thinking waves order steady state, not just the sync
  • Assuming Argo CD proceeds to the next wave immediately
  • Believing the default wave is 1
  • Using waves in place of application-level retry logic
  • Confusing a wave number with a priority for scheduling

context