skip to content

You pushed a commit to the repository Flux watches and nothing changed on the cluster. How do you work out where delivery stopped, using the flux CLI and the resource statuses?

level: seniorimportance: must knowfreq 55%

answer

  1. find the first stale link, not the logs
  2. source first, then the consumer
  3. compare revisions at each stage
  4. suspended objects fail silently
  5. force a run to separate broken from slow

basics

~20 s

Work outward from the source. Check the GitRepository's Ready condition and the revision it fetched, then the Kustomization's condition, revision and suspend state. A failed fetch, a stale artifact, a build error, an unmet dependency or a long interval each stop delivery at a different, visible point.

solid answer

~50 s

Flux's pipeline has distinct stages, and each one reports its own status, so the job is to find the first red or stale link rather than to read logs. Start with `flux get sources git` — is the `GitRepository` Ready, and is its revision the commit you pushed? If not, the failure is fetch-side: wrong branch or ref, credentials, an unreachable host, or simply that the source's `interval` has not elapsed. If the source has your commit, run `flux get kustomizations` — is it suspended, is it Ready, and does its applied revision match the source's? A `Kustomization` can be stuck on a build error, blocked on a `dependsOn` that is not ready, failing a health check, or waiting for its own interval. `flux reconcile kustomization <name> --with-source` forces the whole chain immediately and is the fastest way to separate "broken" from "not yet". Then `flux logs`, `flux diff` and the object's conditions give the specific cause.

code

bash · 12 lines
bash
# 1. did the source see the commit?
flux get sources git

# 2. did the consumer act on that revision?
flux get kustomizations

# 3. force the whole chain now, to separate broken from slow
flux reconcile kustomization apps --with-source

# 4. specific cause
flux logs --level=error --all-namespaces
flux diff kustomization apps --path=./apps/production

go deeper

for a junior

Know the two commands that show the state — flux get sources git and flux get kustomizations — and that each resource reports its own Ready condition and revision.

for a middle

Explain the chain: fetch produces an artifact with a revision, then the Kustomization builds and applies it. Compare revisions between the stages to find where the pipeline stopped.

for a senior

Diagnose by stage under pressure: distinguish fetch failure, suspension, build error, unmet dependency, failed health check and plain interval latency, then use a forced reconcile with source to separate broken from merely slow.

for a principal

Own the detectability problem — a stale-but-Ready source silently freezes delivery, so alerting on source and Kustomization conditions, plus a push webhook with polling as backstop, is a platform requirement rather than a per-team choice.

## Think in stages, not in one agent Because Flux is several controllers, "nothing happened" is never one thing. The chain is: commit exists in the remote → source-controller fetches it and publishes an artifact with that revision → kustomize-controller builds the configured `path` from that artifact → it applies the result → the applied objects become healthy. Every arrow can stall independently, and every stage writes its outcome into a status condition. A structured walk finds the break in about two minutes; guessing at logs takes an hour. ## Stage 1 — did the source see the commit? ```bash flux get sources git kubectl -n flux-system describe gitrepository fleet ``` You want two things: `Ready=True`, and a revision that contains your commit SHA (as of Flux 2.x the revision reads like `main@sha1:<commit>`). Common findings here: - **Not Ready.** The message names it: authentication failure (the referenced Secret is wrong or the deploy key was rotated), host unreachable, unknown ref. Credentials are the number-one cause after a key rotation nobody logged. - **Ready but an older revision.** Either the source's `interval` has not elapsed, or you pushed to a different branch than `spec.ref.branch`, or the repo you pushed to is not the URL in `spec.url` — a fork, a mirror, a second remote. Confirm with `git rev-parse HEAD` against the reported revision instead of trusting the branch name. - **Ready with the right revision.** Fetching is fine; move on. Note that a fetch failure leaves the *previous* artifact in place, so the cluster stays at the last good state rather than degrading — good behaviour, but it means the cluster looks fine while being silently stale. That is exactly what alerting on source readiness is for. ## Stage 2 — did the Kustomization act on it? ```bash flux get kustomizations kubectl -n flux-system describe kustomization apps ``` Things that show up here: - **Suspended.** `flux suspend` sets `spec.suspend: true` and reconciliation stops silently. Somebody suspended it during an incident and never resumed. The CLI marks it, so look before you theorise. - **Blocked on a dependency.** The condition says it is waiting for another Kustomization; the real problem is one layer down. - **Build failure.** A kustomize error — a missing file referenced in `resources:`, a bad patch, a `postBuild` substitution with no value. Nothing is applied and, importantly, nothing is pruned. - **Apply failure.** Server-side apply conflicts, admission webhook rejection, or RBAC denial when the Kustomization runs under a restricted `serviceAccountName`. - **Health-check failure.** It applied, but a named object in `healthChecks` never became healthy within `timeout`. The manifests are live; the workload is broken. Now you are debugging the app, not Flux. - **Wrong path.** Everything is green and the revision matches — because the Kustomization builds a `path` that does not contain the file you changed. Green status, no effect. - **Just early.** The Kustomization's `interval` has not elapsed since the artifact updated. ## Stage 3 — force it, to separate broken from slow ```bash flux reconcile kustomization apps --with-source ``` The CLI does this by annotating the object with `reconcile.fluxcd.io/requestedAt`, which the controller watches; `--with-source` reconciles the referenced source first, so you exercise fetch and apply in one command. If everything converges instantly, your problem was latency, and the fix is a `Receiver` webhook from the Git host so pushes trigger reconciliation instead of waiting on polling. If it fails, you now have a fresh, timestamped error instead of a stale one. ## Stage 4 — the specific cause ```bash flux logs --level=error --all-namespaces flux diff kustomization apps --path=./apps/production flux tree kustomization apps ``` `flux logs` aggregates controller logs so you do not have to know which controller to tail. `flux diff` builds the manifests locally and diffs them against the live cluster — the direct answer to "did my change even alter the rendered output?", which catches wrong-path and patch-not-applied cases immediately. `flux tree` shows the inventory of objects a Kustomization owns, which answers "is this object even managed by Flux?". And `flux check` verifies the controllers themselves are running and version-compatible — worth one run when *everything* is stale rather than one app. ## The one non-Flux cause worth keeping in mind If the object changed and then changed back, you are not looking at a delivery failure but at a fight over ownership: something else in the cluster is modifying a field Flux owns, and each side reverts the other on its own loop. The tell is a value that flips rather than a value that never arrives, and the resolution is to decide who owns the field — either stop the other actor or stop managing that field in Git. ## The habit to demonstrate Name the stages, check them in order, and say what each status *means* rather than what command you would type. Interviewers are listening for whether you know that a healthy Kustomization can sit on a stale artifact, and that green Flux resources do not imply a healthy application unless you asked Flux to check.

  • The GitRepository is Ready but its revision is two days old. What are you looking at?
    A stale artifact that is being reported as healthy for the *last successful* fetch — Flux does not fail the resource retroactively. Check whether fetches are actually succeeding now (the last-transition time and any recent error), whether the ref still exists after a branch rename, and whether the repository URL still points at the remote you push to.
  • Everything reports Ready and the revision matches, but your change is not live. What is the likely cause?
    The change is outside what that Kustomization builds — a different `path`, a directory not referenced by the kustomization, or a file excluded by the source's `ignore` rules. `flux diff kustomization <name>` renders the manifests and diffs them against the cluster, which shows immediately that the rendered output never changed.
  • How do you get pushes applied in seconds rather than at the polling interval?
    Create a `Receiver` in notification-controller, expose its webhook endpoint, and configure the Git host to call it on push with the shared token. The Receiver then triggers reconciliation of the named sources immediately. Keep the polling `interval` configured as the backstop for missed or failed webhook deliveries.

saying these in an interview costs you the question

  • Goes straight to controller logs before reading statuses
  • Assumes a green Kustomization means the app is healthy
  • Forgets that a resource can be suspended
  • Restarts the controllers as a first response
  • Thinks a fetch failure rolls the cluster back

context