A GitOps agent is configured to delete cluster resources that no longer appear in Git. An engineer renames the manifests directory and the next reconcile deletes a live namespace. Why does deletion behave differently from ordinary drift correction, and how would you bound its blast radius?
answer
- absence is the trigger, not a value
- a rename can render nothing at all
- render failure must not mean delete everything
- scope decides what can even be a candidate
- revert restores manifests, not data
basics
~20 sDeletion (prune) acts on absence, so anything that makes the rendered source look empty — a rename, a wrong path, a branch checkout — reads as an intentional removal. Bound it by scoping ownership narrowly, protecting stateful objects, and never pruning on a render failure.
solid answer
~50 sDrift correction acts on what Git *says*: it rewrites fields on objects the source still declares, and the worst case is that the wrong value gets reasserted. Prune acts on what Git *no longer says*, so its input is an absence — and absence is exactly what a rename, a typo in the path, an empty branch or a partial render produces. That asymmetry is why deletion is a separate opt-in policy with a much larger blast radius. To bound it: keep each agent's ownership scoped to a narrow set of resources and a namespace, so nothing outside that scope is even a deletion candidate; make sure the agent treats a fetch or render failure as an error and skips the pass rather than reconciling zero objects; keep stateful resources out of the automatically pruned set or protect them explicitly; and alert on the size of the destructive diff, because a pass proposing forty deletions is almost never a real intent.
go deeper
Know that removing a manifest from Git can cause the agent to delete the live resource, and that this behaviour is a separate opt-in from ordinary drift correction.
Explain why deletion is triggered by absence rather than by a declared value, and list the ways a source can render empty by accident — renamed paths, wrong branch, failed templating, changed ownership labels.
Show how you would bound it in production: scoped ownership, failing closed on render errors, keeping stateful resources out of the automatically managed set, and alerting on destructive diff size. Be honest that reverting restores manifests, not data.
Own the estate-level position: which resource classes may ever be deleted by automation, who approves changes to that boundary, how deletion events are reviewed, and why blanket disabling trades an acute risk for a chronic one — an estate full of orphans nobody can attribute.
## Three policies, three blast radii It helps to separate what a reconciler can be told to do, because candidates routinely collapse them into one switch: 1. **Detect** — compute the diff and report it. Nothing in the cluster changes. Blast radius: none, beyond a noisy dashboard. 2. **Correct (self-heal)** — re-apply declared values to objects Git still describes. Blast radius: bounded by what is declared. The worst outcome is that the cluster is forced to a state a human wrote down and reviewed. 3. **Prune** — delete objects that the agent previously created and that the source no longer contains. Blast radius: unbounded by the source, because the trigger is absence. That third row is the whole question. Correction is driven by a positive statement; prune is driven by a missing one, and missing is the failure mode of every step upstream in the pipeline. ## How the accident actually happens The directory rename is the classic, but the family is larger: - The path the agent watches no longer resolves, so the rendered set is empty. - A templating step succeeds but produces nothing because a values file was renamed or a condition flipped. - The agent is pointed at a branch or tag that does not contain the manifests. - A refactor moves resources between applications, and for one pass neither owns them. - Labels or annotations used to identify what the agent owns are changed, so previously adopted objects are no longer recognised — or, worse, objects it never created suddenly look owned. In each case reconciliation is working exactly as designed: the desired state is "these zero objects", and the live state has forty, so it removes forty. This is why the interesting engineering is in refusing to reconcile a source you do not trust. ## Bounding the blast radius **Fail closed on render errors.** Mature agents distinguish "the source failed to fetch or build" from "the source built and contains nothing". The first must abort the pass and leave live state untouched. Verify this behaviour rather than assuming it, and alarm on repeated source failures — a paused reconciliation is a silent outage of your delivery path. **Scope ownership narrowly.** One agent configuration owning one namespace and a known set of resource kinds means a bad render can only delete inside that box. A single cluster-wide configuration owning everything turns one typo into an estate-wide event. **Keep destructive candidates small.** Stateful resources — persistent volume claims, databases created by operators, anything holding data — are the ones you cannot recover with a `git revert`. Either exclude them from the automatically managed set, protect them with a retain/deletion-protection mechanism at the resource or storage level, or manage them under a separate configuration that requires human confirmation. **Gate on diff size.** A reconcile that proposes deleting one object is routine; one proposing to delete a namespace and everything inside it is not. Alerting on destructive diff volume, or requiring approval above a threshold, catches the rename class before it lands. **Rehearse in a lower environment first.** The same manifests should reconcile into a non-production cluster ahead of production, on a promotion path, so a structural mistake destroys something disposable. ## Why 'just revert the commit' is only half a fix Candidates reach for this immediately and it is worth pushing back. Reverting restores the *manifests*, and the agent recreates the objects. It does not restore: - **Data.** A recreated PersistentVolumeClaim is empty unless the underlying volume had a retain policy and you rebind it deliberately. - **Externally allocated identity.** Cloud load balancers, static IPs, DNS records and certificates issued to the old object may come back with different values, so the outage continues after the resources exist again. - **Anything a controller had reconciled downstream** of the deleted object, which now has to converge again from scratch. So the recovery story is "restore the source, then check what state was destroyed", not "revert and relax". ## The judgement being tested Senior candidates are expected to say that deletion is the one place where the pull-based, continuously-reconciling model can act destructively without a human in the loop, and that the mitigation is not to disable it — an estate that never removes anything accumulates orphaned resources nobody dares touch — but to make the destructive path narrow, observable, and unable to trigger on an error condition.
- If everything is in Git, why can't you simply revert the commit and move on?Reverting brings the manifests back and the agent recreates the objects, but data is not in Git. A recreated volume claim is empty unless the underlying volume was retained, and externally allocated things — load balancer addresses, issued certificates, DNS — may come back different. Recovery means restoring the source and then auditing what state was actually destroyed.
- How should an agent distinguish a failed render from a genuinely empty source?Fetch and build errors must be treated as errors: skip the pass, keep live state, and surface the failure loudly. Only a successful render that legitimately produces zero objects should ever be reconciled as an empty desired state — and even then, a diff proposing mass deletion deserves a gate rather than immediate execution.
- Why not just disable deletion everywhere?Because then nothing is ever removed, and the cluster fills with orphaned resources that no source of truth describes — the exact condition GitOps exists to prevent. Removing a manifest silently doing nothing is its own trap. The answer is narrow scope, protected stateful resources, and observability on destructive diffs, not blanket avoidance.
saying these in an interview costs you the question
- Thinks drift correction and deletion are one setting
- Assumes an empty render means there is nothing to do
- Enables cluster-wide deletion to keep the estate tidy
- Says reverting the commit fully restores a deleted namespace
- Believes recreated volume claims come back with their data