skip to content

In a Kubernetes cluster managed by a GitOps agent, an engineer edits a live Deployment by hand instead of changing Git. What does the agent do, and what decides whether that edit survives?

level: juniorimportance: should knowfreq 62%

answer

  1. Git is compared, not consulted once
  2. the next loop notices the difference
  3. reporting and correcting are separate settings
  4. self-heal reverts, detection only alarms
  5. pause the agent before touching production

basics

~20 s

The agent's next reconcile compares live state against the revision in Git, sees the difference and reports the resource as out of sync. Whether the edit is reverted depends on whether automatic correction is enabled; with detection only, it is reported but left alone.

solid answer

~40 s

Every pass, the agent renders the desired state from Git and diffs it against what is live. The hand edit shows up as a difference and the resource is marked out of sync — that part happens regardless of configuration. What happens next is a separate policy: with drift correction (self-heal) on, the agent re-applies the declared values and the edit disappears within a sync interval; with detection only, the difference is reported and alarmed, and a human decides. Either way the edit is doomed, because Git never learned about it — even in detect-only mode the next legitimate apply of that manifest overwrites it. The correct emergency procedure is to pause reconciliation for that application, make the change, then land it in Git and resume, so the fix is both immediate and recorded.

go deeper

for a junior

Know that the agent re-compares Git against the cluster on a loop and will flag the hand edit as out of sync. Say clearly that whether it gets reverted is a configuration choice, not a law.

for a middle

Explain the mechanics of the diff — rendered desired objects versus live objects, per declared field — and distinguish detect-only from self-heal. Be ready to say why the edit is lost even with self-heal off.

for a senior

Demonstrate the operational side: the break-glass sequence, why a paused application needs its own alert, and how you decide which applications get self-heal enabled versus detection only while an estate is being onboarded.

for a principal

Own the policy question. Decide where automatic correction is mandatory versus advisory across an estate, how manual-change events feed incident review, and how you keep the guardrail from being routed around by teams under delivery pressure.

## What the agent is actually comparing On each pass the agent does three things: fetch the source at some revision, render it into a concrete set of objects, and compare that set to what the cluster reports. The comparison is per-object and, in practice, per-field over the fields you declared. A hand edit that changes a declared field — replicas, an image tag, a resource limit — lands squarely in that diff. The result is a status: this application is *out of sync*, and here is the difference. That reporting is the part people underrate. Before GitOps, the answer to "is production what we think it is?" was inference from pipeline logs. With a reconciler it is a live, queryable status you can alert on. ## Detection and correction are two different settings The question interviewers are really probing is whether you know that noticing drift and fixing drift are separate decisions: - **Detect only.** The agent computes the diff, marks the resource out of sync, and stops. Nothing in the cluster changes. This is what you want while you are onboarding an estate you do not fully trust yet, or where operators legitimately mutate objects. - **Self-heal.** The agent re-applies the declared state whenever the diff is non-empty, so hand edits are reverted within roughly one interval. This is what makes the phrase "Git is the source of truth" literally true rather than aspirational. Neither of these is the same as **deleting** objects that vanished from Git — that is a third policy with a much larger blast radius, and it is deliberately opt-in separately. ## Why the edit is doomed either way Even with self-heal off, the hand edit is living on borrowed time. The moment anyone merges a change that re-renders that object, the declared values are applied over it. So the failure mode of "quick manual fix, tell nobody" is not that it gets reverted instantly — it is that it gets reverted at an unpredictable time, usually while somebody else's unrelated change is being investigated as the cause. ## The break-glass procedure A good answer names the operational path rather than moralising about it. Emergencies happen and Git round-trips take minutes you may not have: 1. Suspend or pause reconciliation for the affected application, so the agent does not fight you mid-incident. 2. Make the change directly. 3. Immediately land the equivalent change in Git. 4. Resume reconciliation and confirm the application returns to synced. Step 3 is the one teams skip, and skipping it is how a cluster quietly accumulates state nobody can reproduce. Some teams add a rule that any paused application raises an alert after a short window, precisely so the pause cannot become permanent. ## What reconciliation does not do It is one-directional. The agent never writes the live state back into Git; there is no "capture what is running" mode, because that would let an unreviewed change become the source of truth. Automation that does commit into the repo — for example writing a newly built image reference — is a separate mechanism producing a reviewable commit, not drift capture. It also does not police fields you never declared. If the manifest says nothing about an annotation that a mutating webhook adds at admission time, most agents will not report it as drift, because the desired state made no claim about it. That distinction — Git owns what it declares, not the whole object — is worth stating explicitly, since it is the reason well-configured clusters are not permanently out of sync. ## Idempotence in this picture Re-applying the declared manifest is idempotent: doing it on a cluster that already matches changes nothing, which is what makes it safe to run the loop every few minutes forever. If re-applying had side effects, continuous reconciliation would be unaffordable.

  • How should an engineer make an emergency change during an incident, then?
    Suspend reconciliation for that application first, make the change, then land the same change in Git and resume. That way the fix is immediate, the agent does not fight it, and the repository still ends up describing what is running. Leaving the application paused is the real hazard — many teams alert on any pause older than a short window.
  • Does the agent ever commit the live state back to Git?
    No. Reconciliation runs one way: Git to cluster. Writing observed state back would let an unreviewed hand edit become the source of truth, which defeats the point. Automation that does commit to the repo — writing an updated image reference, for instance — produces a normal reviewable commit; it is not capturing drift.

saying these in an interview costs you the question

  • Thinks the cluster blocks manual edits outright
  • Assumes every GitOps agent reverts drift automatically
  • Believes the agent writes the manual change back to Git
  • Says drift is only noticed when someone merges
  • Treats pausing the agent as unnecessary during an incident

context