An Argo CD sync applies every manifest successfully, then the PostSync smoke-test Job fails. What does Argo CD do with the already-applied resources, and how do you get an actual rollback?
answer
- a sync is not a transaction
- the change is already live
- operation Failed, resources untouched
- the only undo is the repository
- the database revert is not covered
basics
~20 sNothing is reverted. Argo CD marks the sync operation Failed, runs any SyncFail hooks, and leaves the new manifests live — the bad version keeps serving. Rolling back means changing Git or delegating the abort to a progressive-delivery controller.
solid answer
~50 sArgo CD has no transaction semantics. Once the `Sync` phase applied the manifests, they are live; a failing `PostSync` hook only marks the operation `Failed` and triggers any `SyncFail` hooks, which are your chance to alert or run a compensating action. The Application typically ends up `Synced` — the cluster does match Git — with health depending on whether the workload itself reports Degraded. To actually roll back you change the declared state: revert the commit, or sync the Application to the previous revision (with `syncPolicy.automated` enabled, a manual revert will be overwritten again unless the Git change is the one you make). The structurally better answer is to stop relying on a post-hoc test: put the verification *inside* the rollout with a progressive-delivery controller such as Argo Rollouts, so failing analysis aborts and shifts traffic back to the stable version before all users see it.
go deeper
Know that Argo CD applies what Git declares and does not undo it on its own, so a failed check after deployment leaves the new version running.
Explain the phase ordering that makes this inevitable: PostSync runs after the manifests are already applied, so the only outcome is a failed operation plus SyncFail hooks, not a revert.
Demonstrate the production judgment: revert in Git as the real rollback, know that manual sync-to-old-revision is defeated by automated sync, and flag that the database migration is outside what Git can undo.
Own the design conclusion — a post-hoc smoke test detects rather than prevents. Argue for where verification belongs (readiness, progressive delivery with automated abort) and what recovery-time target justifies that extra machinery.
## What actually happens, step by step 1. `PreSync` hooks run and succeed. 2. The `Sync` phase applies every manifest. The new pods start; the old ones go away as the workload controller replaces them. 3. Argo CD waits for the applied resources to report Healthy. 4. `PostSync` hooks run. The smoke-test Job exits non-zero, so Argo CD assesses it as Degraded. 5. The **operation** is marked `Failed`. `SyncFail` hooks, if any exist, run now. 6. Nothing is undone. That last line is the whole lesson. A sync is not a transaction. There is no snapshot to restore, no "previous manifest set" that Argo CD holds in reserve and re-applies. By the time a `PostSync` hook can observe a problem, the change is already in production. ## Why the Application often still shows Synced The sync status is a diff verdict. The cluster genuinely does match the target revision — Argo CD applied it faithfully. What is broken is the *content* of that revision, which the diff cannot see. Health status may or may not catch it: if the new version crash-loops, the Deployment goes Degraded and the Application does too; if it starts cleanly and merely returns wrong answers, the Application sits at `Synced` and `Healthy` while the smoke test says otherwise. The operation state (`Failed`, with the hook's logs) is the only place the failure is recorded, which is why alerting on Application health alone is insufficient. ## Getting back to a working version The declared state is the only lever: - **Revert in Git.** `git revert` the offending commit and let reconciliation converge. This is the GitOps-correct answer: the repository stays the source of truth and the rollback is auditable like any other change. - **Sync to a previous revision** from the CLI or UI for speed. Understand the caveat: if the Application has automated sync enabled, the reconciler will pull it forward to the tracked branch head again shortly, so this buys minutes, not a fix. It is a stopgap while the revert PR merges. - **Pause automation** first if you need the cluster to hold still while you investigate. Note also that reverting the *application* manifests does not revert whatever a `PreSync` migration did to your database. Schema changes are not covered by GitOps rollback at all — this is the argument for expand/contract migrations that are safe to leave in place while the code goes back. ## Using SyncFail properly `SyncFail` hooks run only when the operation fails, which makes them a reasonable place for: ```yaml metadata: annotations: argocd.argoproj.io/hook: SyncFail argocd.argoproj.io/hook-delete-policy: BeforeHookCreation ``` …a Job that posts to an incident channel, captures diagnostics, or triggers a compensating action. Do not model them as an automatic rollback: they are ordinary manifests with no privileged knowledge of the previous state, and one that tries to re-apply an old revision immediately fights the reconciler. ## The structural fix A smoke test that runs *after* the whole fleet has been replaced is a detector, not a guard. Two changes move the verification earlier: - **Make the workload's own readiness the gate.** If the new version cannot serve, it should fail to become ready, the controller stops replacing old pods, and the blast radius is limited to the surge. - **Put verification inside the rollout.** A progressive-delivery controller shifts a fraction of traffic first and evaluates metrics before continuing, so a bad version is aborted while most users are still on the stable one. That is a different mechanism from a `PostSync` hook, and it is the one to reach for when "we found out afterwards" is not acceptable. ## What interviewers are testing Whether you understand that a declarative reconciler guarantees *convergence*, not *safety*. Candidates who assume the tool rolls back on failure have usually only used it on a happy path. The strong answer states the no-rollback fact plainly, names Git as the only real undo, flags the database as the part Git cannot undo, and then proposes moving the check earlier rather than adding more post-hoc hooks.
- With automated sync enabled, why does syncing to an earlier revision from the CLI not stick?Because automated sync continuously reconciles the Application toward the head of its tracked `targetRevision`. Pinning it to an old revision by hand leaves the live state differing from that branch, so the reconciler pulls it forward again on the next pass. The durable fix is to change what the branch points at — revert the commit — or to disable automation while you investigate.
- Does reverting the commit also undo what your PreSync migration did to the database?No. Git reconciliation covers Kubernetes objects only; a migration that already ran altered state Argo CD does not manage. That is the practical argument for expand/contract migrations — additive, backward-compatible steps that the previous version can still run against — so that rolling application code back does not require rolling schema back.
- Is a SyncFail hook a reasonable place to trigger a rollback?It is a reasonable place to alert, capture diagnostics or run a compensating action, but not to roll back. A SyncFail hook is an ordinary manifest with no memory of the previous state, and one that re-applies an old revision immediately fights the reconciler that is converging toward Git. Rollback belongs in the repository or in a controller that owns traffic shifting.
saying these in an interview costs you the question
- Assumes Argo CD reverts the sync when a hook fails
- Thinks a failed operation means the manifests were never applied
- Expects Synced to turn OutOfSync because the smoke test failed
- Believes a Git revert also undoes the database migration
- Treats SyncFail hooks as an automatic rollback mechanism