skip to content

A GitOps agent reconciles a cluster from Git on a fixed interval of a few minutes. What determines how long after a merge the change is actually running, and why do teams keep the periodic reconcile after wiring a push notification from the repository?

level: middleimportance: nice to knowfreq 35%

answer

  1. latency is a sum, not one number
  2. half a period on average, full period worst case
  3. a webhook is one delivery, no retry
  4. cluster drift emits no repository event
  5. fast comes from notifications, correct from the timer

basics

~20 s

End-to-end latency is the build and publish time, plus the commit landing in the config repository, plus up to one poll interval before the agent notices, plus apply and rollout. A push notification cuts the detection wait but is lossy, and cluster-side drift raises no repository event at all.

solid answer

~50 s

Break the latency into stages: CI builds and publishes the artifact, the manifest change lands in the config repo, the agent notices it, applies it, and the workload finishes rolling out. Only the third stage is the interval, and its cost averages half the period and worst-cases at the full period. A push notification from the repo host collapses that to seconds, which is why teams add one. They keep the timer anyway for two reasons. First, notifications are edge-triggered and lossy — a delivery can fail, the agent can be restarting, a network partition can swallow it — and nothing retries a missed event. Second, and more fundamentally, drift inside the cluster produces no repository event whatsoever, so only the periodic pass ever finds it. The timer is what keeps the system level-triggered and self-correcting; the notification is a latency optimisation layered on top.

go deeper

for a junior

Be able to say the agent checks Git on a schedule, so a merged change waits up to one interval before it is picked up, and that the interval is only one part of the total wait.

for a middle

Decompose the latency into build, commit, detection, apply and rollout, and explain that a push notification shortens only the detection stage. State clearly that the timer remains because notifications can be missed and drift produces no event.

for a senior

Set the interval as a deliberate target for how long you tolerate being wrong about live state, and account for the aggregate load of many applications polling one repository host. Mention staggering and separating source-fetch cadence from re-apply cadence.

for a principal

Own the estate-wide numbers: how many reconcilers point at one repository host, what that costs in requests and rate limits, and how delivery latency targets are met by notifications rather than by driving every interval toward zero.

## Where the time actually goes When someone asks "why did my change take eight minutes to go live", the interval is usually the smallest term. The stages: 1. **Build and publish.** CI compiles, tests, and pushes an immutable artifact. Typically minutes. 2. **Manifest change lands.** The new artifact reference is committed to the configuration source — sometimes automatically, sometimes as a pull request that a human merges. This stage has the widest variance, because a review gate is unbounded. 3. **The agent notices.** Zero to one interval if polling; seconds if a push notification arrives. 4. **Apply.** The agent submits the changed objects to the API. Fast. 5. **The workload converges.** The rollout of the new version proceeds under the workload controller's own rules, which is where readiness and surge settings dominate. Shortening step 3 from minutes to seconds is a real improvement, but a team complaining about delivery latency usually has its time in steps 1, 2 or 5. ## Why the timer stays **Notifications are edge-triggered and lossy.** A webhook is one HTTP request. If the agent is restarting, unreachable, or the delivery fails, there is no second chance for that event, and the change simply waits — until the timer fires. The interval is the recovery mechanism that makes a missed notification a delay rather than a loss. **Drift raises no event.** This is the deeper reason and the one an interviewer is listening for. If someone edits an object in the cluster, the repository is untouched, so no push notification exists to send. The only thing that ever discovers cluster-side divergence is a pass that compares. Remove the timer and you have rebuilt an event-driven pipeline with extra steps. **Convergence needs repetition.** A pass that fails — an unavailable dependency, a transient API error, a webhook rejecting the request — must be retried. The loop is the retry. A useful way to say it: the notification makes the system *fast*, the interval makes it *correct*. ## Choosing the interval Shorter is not free. Each pass fetches the source, renders it, and reads live state, which costs requests against the repository host and the cluster API. Multiply by hundreds of applications and you meet rate limits, noisy status churn, and a control plane doing measurable work just to confirm nothing changed. Longer intervals reduce that load but stretch the window in which drift goes unnoticed and failed passes go unretried. Practical shape of the decision: - Use a notification for delivery latency, so the interval no longer has to be short for speed. - Set the interval by how long you are willing to be wrong about the cluster — this is a drift-detection SLO, not a delivery one. Minutes for production, longer for large low-risk sets. - Stagger intervals across many applications so passes do not synchronise into a thundering herd against the repo host and the API. - Note that many agents separate the cadence of *fetching the source* from the cadence of *re-evaluating and applying*, so you can poll a busy repository less often than you re-check the cluster. ## The interval is what makes it not a pipeline Come back to the definitional point. A pipeline runs once per event and its duration is the whole of its existence. A reconciler has no end; the interval is simply how often it takes its next look. That is why the answer to "what happens if the reconcile fails" is "it runs again shortly" rather than "someone re-triggers the job", and why the answer to "how do you know production matches the repo" is a live status rather than a log from the last deploy. ## Common mistakes to avoid saying Claiming the change is live the moment CI goes green skips both the commit and the reconcile entirely. Claiming a webhook makes the interval unnecessary misses that drift generates no event. And attributing the whole wait to the interval when the rollout itself takes minutes points at the wrong stage — measure before tuning.

  • Why does drift correction still depend on the interval even with a webhook configured?
    Because the webhook fires on repository activity, and drift happens in the cluster. Someone scaling a Deployment or an operator patching a field changes nothing in Git, so no event exists to deliver. Only a scheduled pass that compares live state against the declared revision can find it, which makes the timer the sole drift-detection mechanism.
  • What breaks if you drop the interval to a few seconds across hundreds of applications?
    Load and noise. Every pass fetches and renders the source and reads live state, so you multiply requests against the repository host — where you will meet rate limits — and against the cluster API. Status churn grows too. The fix is a notification for latency plus staggered, longer intervals for correctness.
  • How would you choose the interval for a production application?
    Treat it as a drift-detection target rather than a delivery one: how long are you willing to be wrong about what is running, and how long may a failed pass go unretried? Answer that, then get delivery speed from a push notification instead of from a short period. A few minutes is a common production answer.

saying these in an interview costs you the question

  • Thinks a webhook removes the need for periodic reconciliation
  • Assumes the change is live the moment CI turns green
  • Sets the interval to seconds without considering rate limits
  • Believes drift in the cluster triggers a repository event
  • Confuses the reconcile interval with the rollout duration

context