skip to content

In a production cluster a Deployment's replica count flips between two values every few seconds and no human is running kubectl. How do you diagnose and fix that kind of fight over a single field?

level: seniorimportance: should knowfreq 34%

answer

  1. flapping = two owners, both working correctly
  2. managedFields + audit log = who is writing
  3. resourceVersion churn on an idle object
  4. drop the field from the wrong source; ignore-paths in GitOps
  5. server-side apply turns silence into 409

basics

~20 s

Two writers each treat that field as theirs and keep correcting each other's drift. Find them via metadata.managedFields, the audit log and controller logs, then give the field exactly one owner — usually by removing it from the declarative manifest and letting the autoscaler or operator own it.

solid answer

~60 s

Flapping means two independent reconcile loops disagree about desired state for one field and each is doing its job: writer A sets 3, writer B observes drift from its own intent and sets 5, repeat. Diagnose by identity, not by guessing. `metadata.managedFields` names the field managers touching each path; the API server audit log shows the actual requests, their user agents and service accounts; churning resourceVersion plus controller logs confirms the pair. Common culprits: a GitOps controller applying a pinned spec.replicas against an autoscaler; two overlapping Helm releases or Kustomize overlays; an operator and a human-owned manifest managing the same annotation; a mutating webhook rewriting a value that an applier then reverts. The fix is ownership, not force. Remove the contested field from whichever source should not own it — drop replicas from the git manifest, or configure the GitOps tool to ignore that path. Convert appliers to server-side apply with named field managers so the next collision surfaces as a 409 Conflict instead of silent flapping. Add an alert on rapid resourceVersion churn so drift wars are detected rather than discovered.

code

bash · 3 lines
bash
kubectl get deploy web -w -o custom-columns=RV:.metadata.resourceVersion,REPL:.spec.replicas
kubectl get deploy web -o yaml --show-managed-fields | grep -E 'manager:|f:replicas|time:'
kubectl get events --field-selector involvedObject.name=web --sort-by=.lastTimestamp

go deeper

for a junior

Recognize the symptom and know that two things are writing the same field; escalate with the object name and observed churn.

for a middle

Run the diagnosis: managedFields, events, controller logs; name the common GitOps-versus-autoscaler pairing and the omit-the-field fix.

for a senior

Own the full path — attribute writers via audit logs, decide ownership, encode ignore rules, migrate appliers to server-side apply, and add detection.

for a principal

Set cluster-wide ownership policy for contested fields and require new controllers to declare the fields they claim, so overlaps are caught at design review rather than in production.

## Why flapping happens at all Every controller is a level-triggered loop closing the gap between its own notion of desired state and what it observes. That is safe while each field has one authority. When two authorities disagree, each one's correction is the other's drift, and the loop never converges — you get a stable oscillation at the speed of the faster reconcile. Crucially, nothing is malfunctioning. Both writers are behaving exactly as designed. The defect is organizational: undefined field ownership. ## Diagnosis, in order 1. **Confirm the churn.** Watch `metadata.resourceVersion` and the field itself: `kubectl get deploy web -w -o custom-columns=RV:.metadata.resourceVersion,REPL:.spec.replicas`. A resourceVersion advancing many times a minute on an otherwise idle object is the signature. 2. **Read managedFields.** `kubectl get deploy web -o yaml --show-managed-fields` lists each field manager and the paths it owns, with timestamps. Two managers alternating on the same path is the answer in most cases. 3. **Check the audit log.** The API server audit log records the verb, resource, user or service account, and user agent of every write. This attributes writes to a real identity when managedFields is ambiguous (client-side apply and plain updates record less). 4. **Read both controllers' logs and events.** Confirm what each believes desired state is and why. 5. **Consider webhooks.** A mutating admission webhook that rewrites a value on every write creates the same symptom against a single applier: the applier sees its value changed and reapplies forever. ## Classic pairs - **GitOps versus autoscaler.** The manifest in git pins `spec.replicas`, an autoscaler scales the same field. Most common cause by far. - **Two deployment tools.** Overlapping Helm releases, a Helm release plus a Kustomize apply, or two pipelines applying the same namespace. - **Operator versus human manifest.** An operator owns part of a derived object while someone also templates it directly. - **Defaulting versus explicit value.** A tool sends a field the API server defaults differently, so the object never matches what the tool sent. ## Fixing it properly The correct fix is to establish a single owner per field. - **Remove the field from the wrong source.** If the autoscaler should own replicas, the git manifest must not mention it. Most GitOps tools also support ignoring specific field paths for diffing, which stops them reverting the autoscaler. - **Use server-side apply with explicit field managers.** Then the next collision is a 409 Conflict at apply time — loud, attributable, and blocking — rather than silent oscillation in production. Do not paper over it with a blanket `--force-conflicts`. - **Split the objects** if two systems truly need independent control of overlapping concerns; overlapping label selectors between workload controllers cause the same class of fight over pods. - **Prevent recurrence.** Alert on high write rate per object, review new controllers for which fields they claim, and write down ownership rules for the fields that always attract conflict: replicas, resource requests, image tags, and common annotations. ## Why not just stop one writer Disabling the autoscaler or pausing the GitOps sync stops the symptom and is a fine emergency action, but it leaves the cluster with the wrong desired state and reintroduces the fight the moment it is re-enabled. Ownership must be decided, encoded in configuration, and made visible — the flapping is a design gap surfacing as noise, and the cure is a boundary, not a bigger hammer.

  • Your GitOps tool must keep managing the Deployment but the autoscaler has to own scaling. How do you configure that?
    Remove spec.replicas from the manifest in git so the applier never claims the field, and configure the GitOps tool to ignore that path when computing drift so an absent field is not reported as out of sync. The tool continues to own the pod template, labels and everything else, while the autoscaler owns replicas alone.
  • How would you catch this class of problem before an incident?
    Alert on write rate per object — a Deployment updated dozens of times a minute with no deploy in progress is almost always a drift war. Adopting server-side apply with named field managers converts future overlaps into 409 Conflicts at apply time, and reviewing which fields a new controller claims before rollout catches the rest.

Two thermostats wired to the same furnace, set to different temperatures. Neither is broken; the building just cycles forever until you decide which one is in charge.

saying these in an interview costs you the question

  • Assuming one of the controllers is buggy rather than that ownership is undefined
  • Fixing it by force-applying on a schedule, or by adding a cron job to reset the value
  • Deleting entries from managedFields as a remedy — the server rewrites them on the next apply
  • Permanently disabling the autoscaler instead of deciding ownership
  • Ignoring mutating admission webhooks as a possible second writer

context