skip to content

Compare infrastructure that is reconciled only when a person triggers a run with infrastructure reconciled by an agent that runs the same loop continuously. What does the continuous model buy you, and what new failure modes does it introduce?

level: seniorimportance: should knowfreq 50%

answer

  1. how long is the estate wrong?
  2. minutes versus the next release
  3. merge becomes deploy
  4. your console fix gets reverted
  5. one writer per field; know the pause switch

basics

~20 s

Continuous reconciliation shrinks the time infrastructure spends wrong from days to minutes and removes the human from the critical path. In exchange, a bad commit reaches production unreviewed, emergency console fixes get reverted automatically, and the loop itself becomes a system you must observe, rate-limit and be able to pause.

solid answer

~50 s

A triggered run reconciles once, then stops caring. Whatever happens to that infrastructure afterwards persists until the next run, so drift lives for however long sits between runs — often weeks. A continuous agent re-evaluates the whole desired state on a short interval, so the window collapses to minutes and correction needs no human. The costs are real. The human gate disappears, so a merged mistake is live before anyone reads a diff. An engineer's emergency console fix is reverted by the loop unless they pause it or land the change in code, which is a genuine incident hazard. Two loops, or a loop and a runtime component, can fight over the same field and flap forever. And the loop is now production software: it needs credentials that stand permanently, its API calls count against rate limits, and it needs metrics, alerts and a documented way to stop it.

go deeper

for a junior

Know the basic difference: a triggered run fixes things once and stops, while an agent keeps checking and re-fixing on its own. Be able to say that the second one corrects a console change without anyone asking it to.

for a middle

Explain the mechanics of the trade — where the review gate sits in each model, how long drift survives, and why merging becomes deploying once a loop is enforcing the declaration continuously.

for a senior

Demonstrate that you have run one: the pause procedure for incidents, the two-writers flapping problem and how you diagnosed it from audit events, standing credentials versus per-run credentials, and what you alert on when the loop stops converging.

for a principal

Own the policy: which classes of resource are placed under a continuous loop at all, how a change is still staged across environments when merge means deploy, and how the loop's own availability and blast radius are governed.

## Two shapes of the same loop The reconciliation loop — read desired, observe actual, compute the gap, close it — does not care what invokes it. What differs is the trigger. In the **one-shot** model a person or a pipeline invokes the tool. It refreshes, diffs, applies, exits. Between invocations nothing watches the infrastructure. The tool is a command, not a service. In the **continuous** model an agent holds the desired state and re-runs the comparison on an interval, forever. Nobody invokes anything; the current state of the world is compared to the declaration again and again. Cluster controllers work this way natively, and continuous delivery agents that watch a repository apply the same shape to infrastructure. ## What continuous buys **Drift lifetime.** This is the headline number. With triggered runs, the time between a console edit and its correction is bounded by your release cadence — for a rarely-touched estate, that is measured in weeks, and often the correction is discovered as a surprise inside an unrelated change. With a loop running every few minutes, the same edit is reverted before the person who made it has closed the browser tab. **No human in the recovery path.** If a resource is deleted by accident, the loop recreates it. Nobody has to notice, find the right repository, and run the right command with the right credentials at 3am. **The repository becomes true.** When a loop continuously enforces the declaration, "what is in the code" and "what is running" stop being two separate questions. That is a real change in how a team reasons about production, and it is why the model spread. ## What continuous costs **The review gate moves.** With triggered applies, the plan is a checkpoint: a human reads the proposed actions before anything happens. Under continuous reconciliation, merging *is* deploying. Whatever review discipline you had must move entirely to the change proposal, because after merge there is no second look. **Emergency changes get reverted.** The most common way this model bites is that an engineer mitigates an incident by hand and the loop undoes it within the minute, sometimes repeatedly, while they are still trying to work out what is happening. Every continuous setup needs a documented, fast, well-known way to suspend reconciliation for a scope, and every responder needs to know it exists. **Fights between writers.** A loop asserts a value; something else asserts a different one; each overwrites the other and the resource flaps. The other writer is often legitimate — a runtime component that adjusts a value as part of its job, or a second automation someone else added. The fix is always the same shape: exactly one writer per field. This is easy to say and hard to enforce, because the second writer is usually invisible until the flapping starts. **A bad declaration propagates fast.** The property that makes recovery automatic makes mistakes automatic too. A merged change with a wrong value is applied everywhere the loop reaches, immediately, with no staged rollout unless you built one. **The loop is production infrastructure.** It needs standing credentials with broad authority — a permanent, high-value target, versus short-lived credentials minted per pipeline run. It polls provider APIs constantly and can hit rate limits, especially over a large estate. It can wedge: stuck retrying a change the provider will never accept, hammering an API and filling logs. So it needs the same treatment as any service: metrics on loop duration and failures, an alert when it has not converged for N minutes, and a runbook. ## Choosing The two models are not exclusive and mature estates mix them. Continuous suits things that are numerous, cheap to recreate, and where automatic correction is unambiguously right — cluster workloads, network rules, configuration. Triggered suits the small set of resources where being wrong is cheaper than being reverted at the wrong moment: databases and anything whose replacement loses data, and anything a human should think about for thirty seconds first. The mistake is treating this as a religious question rather than a per-resource one. ## Answering well Interviewers ask this to see whether you have operated a loop, not whether you like GitOps. The tell is that you volunteer the pause mechanism and the two-writers problem without prompting — those are the two things nobody thinks about until they have been on the wrong end of them.

  • An engineer needs to change something by hand during an incident on infrastructure a loop owns. What should the process be?
    Suspend reconciliation for that scope first, make the change, then land the same change in the declaration and resume. The suspend mechanism has to be documented in the runbook and fast enough to use under pressure, because the alternative — fighting the loop while it reverts you mid-incident — turns one incident into two. Resuming without codifying the fix simply schedules the revert for later.
  • How would you detect that a continuous loop and something else are fighting over the same resource?
    Look for oscillation rather than a one-time diff: the same attribute changing back and forth on a regular cadence, a resource that never reports converged, and provider audit events alternating between two principals. The audit trail is the decisive evidence, because it names the other writer. The remedy is to make one side authoritative — stop declaring the field, or stop the other component writing it.
  • Does continuous reconciliation remove the need for staged rollout across environments?
    No, it makes staging more important. Automatic application means a wrong declaration reaches everything the loop covers as soon as it merges. You still promote a change through environments by pointing each environment's loop at a different revision or a different declaration, so production picks up the change only after the earlier stage has been observed to be healthy.

saying these in an interview costs you the question

  • Continuous reconciliation removes the need for change review
  • A loop means drift can no longer happen
  • You can safely fix things in the console; the loop will notice
  • Two systems managing one attribute just settles on the last write
  • The agent is infrastructure glue, not something to monitor

context