You inherit an estate where some infrastructure is reconciled continuously by an agent, some by a nightly job, and some only when an engineer opens a pull request. How do you decide which loop each resource should belong to?
answer
- classify, do not standardise
- cost of wrong vs cost of correction
- one writer per field, always
- reversible and numerous go fast
- no fast loop without a pause switch
basics
~20 sDecide per resource by comparing the cost of being wrong with the cost of being corrected automatically. Reversible, numerous, cheap-to-recreate resources belong on the fast loop; destructive or data-bearing ones stay human-gated. Above all, give every field exactly one owning loop.
solid answer
~50 sI would not pick one cadence for the estate; I would classify resources. The first question for each is what an automatic correction costs versus what living with the deviation costs. Stateless, numerous, easily recreated things — network rules, cluster workloads, host configuration — belong on the fast loop, because being wrong is worse than being reverted. Anything whose correction can destroy data or interrupt service stays on a human-gated run, and I would want an explicit guard so the loop can never replace it silently. The second question is ownership: each attribute needs exactly one authoritative writer, so anything a runtime component legitimately adjusts must not also be asserted by a loop, or the two will flap. Then I would look at how often each resource legitimately changes, whether the fast loop can be paused per-scope during an incident, and whether we can actually observe the loops. The mixture is fine; the undocumented mixture is what I would fix first.
go deeper
Know that not everything should be corrected automatically, and be able to give the contrast: a firewall rule can be reverted freely, a production database cannot. That instinct is what the tiering formalises.
Explain the trade in both directions — the cost of a deviation surviving versus the cost of an automatic correction — and be ready to place a few concrete resource types on the right side of it.
Demonstrate the operational preconditions you would insist on before letting a loop own production: a scoped pause procedure, convergence metrics and alerts, staged promotion between environments, and guards on destructive actions.
Own it as written policy rather than case-by-case judgment: a rule that assigns new resources to a tier, single-writer-per-field enforced across teams, and a measured time-to-convergence you use as evidence to move classes between tiers.
## Treat it as a classification problem The instinct is to standardise — move everything to the fast loop, or pull everything back behind review. Both are wrong, because the resources genuinely differ. The useful framing is a single trade made per resource class: **what does it cost to be wrong, versus what does it cost to be corrected without a human?** For a firewall rule, being wrong is a security exposure and being reverted is nothing. Fast loop. For a production database, being wrong is usually a tolerable inconsistency and being "corrected" may mean replacement and data loss. Human gate, with an explicit protection so the loop cannot destroy it even if the declaration says so. ## The axes that actually decide it **Reversibility and blast radius.** Can the correcting action destroy or interrupt something? Anything that carries data, holds a durable identity, or is the single instance of something goes behind a gate. Anything that can be thrown away and recreated in seconds does not need one. **Multiplicity.** Ten thousand similar objects cannot be human-reviewed individually and must be enforced automatically. A handful of foundational, rarely-changed accounts and networks can be, and benefit from it. **Legitimate rate of external change.** If something is *supposed* to be adjusted at runtime — by a scaling component, by an operational tool — then either it is not declared at all, or its value is deliberately not asserted. Which brings the decisive rule. **Single writer per field.** This is the rule that outranks cadence. Two systems asserting the same attribute produce oscillation that is expensive to diagnose and destabilising while it runs. Before choosing intervals, I would inventory who writes what, and remove the second writer for every contested field. A loop that reverts a runtime component every ninety seconds is worse than no loop. **Change velocity and lead time.** Something changed several times a day gains a lot from automatic application; something changed twice a year gains almost nothing, and the standing credentials the loop needs are then pure risk for very little benefit. ## What must be true before anything goes on a fast loop Three preconditions, none negotiable. *A pause switch, scoped and documented.* Responders must be able to suspend reconciliation for one system in seconds, without a merge, and must know it exists before the incident. A loop that cannot be stopped will eventually be stopped by revoking its credentials, which is worse. *Observability of the loop itself.* Time since last successful convergence, per-scope failure counts, and an alert on "not converged for N minutes". A loop that silently stopped looks identical to a healthy one that has nothing to do. *Staged promotion.* If merging applies immediately, the change still has to reach environments in order. Each environment's loop must track a different revision, or the fast loop simply distributes mistakes faster. ## The nightly job in the middle The nightly tier is often the least examined and deserves a decision, not inheritance. Ask what it is for. If it is *enforcing*, then a resource is being corrected up to twenty-four hours after it goes wrong, which is either acceptable (in which case say so explicitly) or not (in which case it belongs on the fast loop). If it is *reporting* — a scheduled comparison that alerts rather than acts — that is a legitimate and underrated third mode, and it is the right first step for a class of resources you are not yet willing to let a loop touch. Making the mode explicit is more valuable than moving the resource. ## How I would sequence the work 1. Inventory: for every resource class, which loop owns it, and who else writes to it. The second column is the one that surprises people. 2. Eliminate contested fields — one writer each. 3. Classify by the cost-of-wrong versus cost-of-correction trade, and write the classification down as policy so new resources land in a tier by rule rather than by whoever created them. 4. Fix the preconditions before widening the fast loop: pause, observability, staged promotion. 5. Measure the outcome — how long deviations survive per tier — and move classes between tiers on evidence. ## What the interviewer listens for That you refuse the false binary, that you name single-writer as a hard rule rather than a nicety, that you demand a pause mechanism before granting a loop authority over production, and that you make the tiering an explicit written policy instead of an accident of history. The wrong answer is a preference for one cadence; the right answer is a rule that assigns cadence.
- How do you protect a data-bearing resource that lives in a repository the fast loop reconciles?Layer it. Put the resource behind an explicit deletion guard so any destructive action fails rather than proceeding, add a policy check that fails the change when a destroy or replace touches a protected class, and require a human approval for that class. Relying on reviewers to spot a replacement in a long diff is not a control — the guard has to be mechanical and the loop must be unable to override it.
- What metric tells you whether the current tiering is right?Time-to-convergence per tier — how long a deviation actually survives before it is corrected — measured, not assumed. Pair it with the count of changes reverted by a loop that a human had made deliberately. The first says whether a tier is fast enough; the second says whether a resource is on a loop that should not own it. Both move classes between tiers on evidence rather than opinion.
- A team asks to put their resources on the fast loop. What do you require before agreeing?That nothing else writes the fields they declare, that their resources are recreatable or explicitly guarded, that they have staged promotion so production is not the first environment to see a change, and that they know and have rehearsed the pause procedure. If any of those is missing, the scheduled reporting tier is the right place until it is fixed.
saying these in an interview costs you the question
- Everything should be continuously reconciled for consistency
- Databases are fine on the fast loop if reviewers are careful
- Two systems writing one attribute is fine; last write wins
- The cadence matters more than who owns the field
- A loop needs no pause mechanism if the code is correct