A new image tag was pushed to the registry 20 minutes ago, Flux reports nothing unhealthy, and the cluster still runs the old image. How do you work out where the image automation chain stopped?
answer
- four stages, check them in order
- scanned, selected, committed, applied
- stop at the first stage that disagrees
- green does not mean it was asked to act
- latency is a sum of intervals
basics
~20 sWalk the chain in order and stop at the first stage whose state is wrong: did the ImageRepository scan see the tag, did the ImagePolicy select it, did the ImageUpdateAutomation commit it, and has the Kustomization applying those manifests reconciled since that commit.
solid answer
~50 sDiagnose stage by stage rather than guessing. `flux get image repository` tells you when the registry was last scanned, how many tags were found, and whether authentication is failing — if the scan is stale or erroring, nothing downstream can be right. `flux get image policy` shows the image the policy currently selects; if the scan saw the tag but the policy did not pick it, your semver range or `filterTags` pattern excluded it. `flux get image update` shows the last automation run; if the policy is right but no commit exists, look at `spec.update.path`, the setter markers in the manifests, and push permissions. Finally check the Git log and `flux get kustomization`, because the commit still has to be reconciled before a Pod changes. Two things explain most "but nothing is red" cases: the total latency is the sum of three independent intervals, and a missing marker is a silent no-op rather than an error. Use `flux reconcile` to force each stage instead of waiting.
code
bash · 7 linesflux get image repository -A
flux get image policy -A
flux get image update -A
flux get kustomization -A
flux reconcile image repository app --namespace flux-system
flux reconcile kustomization apps --with-source --namespace flux-systemgo deeper
Know the order of the chain — scan, select, commit, apply — and that you check each stage's status with the Flux CLI rather than jumping straight to editing YAML.
Map each symptom to its stage: a stale scan, a policy whose rule excludes the tag, a missing setter marker, a Kustomization that has not reconciled yet. Explain why total latency is a sum of intervals.
Lead the diagnosis under time pressure: rule out the uninstalled-controller case, localise the stall in one pass, and articulate why a silent no-op produces no error condition anywhere.
Make this diagnosable before the incident — alerting on scan and reconcile staleness rather than only on failures, sane interval choices, and a documented expectation of how long a promotion legitimately takes.
## Why an ordered walk beats guessing Image automation is a pipeline of four independent stages, each with its own controller, its own interval and its own status. The reason Flux models it as separate objects is precisely so you can localise a stall. Check them in order and stop at the first one that disagrees with your expectation; everything downstream of a broken stage is meaningless. ## Stage 0: are the controllers even installed? Both image controllers are optional. If someone created the custom resources on a cluster bootstrapped without `--components-extra=image-reflector-controller,image-automation-controller`, the objects exist and nothing reconciles them. The tell is objects with no status at all, and no corresponding Deployment in the `flux-system` namespace. This is the single most embarrassing cause and worth ruling out in ten seconds. ## Stage 1: did the scan see the tag? `flux get image repository` reports the last scan time and the number of tags found. Things that go wrong here: - **The scan is simply not due yet.** With `spec.interval: 5m` a tag pushed 20 minutes ago should be visible, but with a longer interval it may not be. - **Authentication is failing.** A registry credential that expired shows up as a ready condition of false with a registry error. This is red, but only on this object — everything else stays green and stale. - **The tag is excluded.** `spec.exclusionList` filters tags before anything else sees them. `flux reconcile image repository <name>` forces a scan immediately rather than waiting. ## Stage 2: did the policy select it? `flux get image policy` prints the selected image. If the scan found the tag but the policy still shows an older one, the fault is your ordering rule, not Flux: - a semver `range` that excludes the new version, for example `~1.4.0` when the new tag is `1.5.0`; - a prerelease tag that a plain range will never select; - a `filterTags.pattern` that the new tag does not match, so it never entered the competition; - alphabetical ordering placing `10` below `9`. This stage is the most common resting place of "automation is broken" reports, and in almost every case the policy is doing exactly what it was told. ## Stage 3: did anything get committed? `flux get image update` shows the last run and its message — typically either the commit it pushed or a note that there was nothing to update. If the policy is correct but there is no commit: - **No setter marker.** The manifest has no marker comment, or it names the wrong `namespace:name`. This produces no error at all: the automation walks the files, finds nothing to set, and reports success. - **Path scoping.** The marked file is outside `spec.update.path`. - **Push rejected.** A read-only deploy key or a protected branch produces a push failure — loud, and visible in the object's conditions and the controller logs. - **Suspended.** Someone set `spec.suspend: true` during a previous incident and never turned it back on. `flux get image update` shows the suspension. ## Stage 4: has the commit been applied? A commit in Git is not a running Pod. Check that the `GitRepository` source has fetched the new revision and that the `Kustomization` or `HelmRelease` owning those manifests has reconciled it — `flux get kustomization` shows the applied revision. If the applied revision predates the automation commit, you are simply waiting on an interval, or the Kustomization itself is failing on something unrelated. ## The latency arithmetic End to end, worst case is the sum of the scan interval, the automation interval and the applying Kustomization's interval. Three five-minute intervals is up to about fifteen minutes of entirely correct, entirely silent delay. Teams that expect a 30-second deploy from a push-based pipeline read this as a failure. If you need it faster, shorten the intervals, drive scans from a registry webhook through a notification-controller `Receiver`, and use `flux reconcile` in the moment. ## The two shapes of "nothing is red" It is worth stating the general principle, because it generalises past this leaf: Flux reports errors for actions it attempted and failed, not for actions it was never asked to take. A missing marker, a filter that matches nothing and a policy range that excludes the tag are all *correct* executions of an incorrect configuration. That is why the diagnostic is a walk through expected state at each stage, not a hunt for a red condition.
- The policy shows the new image but the repository has no commit. Where do you look?The automation side. Confirm the target file carries a setter marker naming the right `namespace:name`, that the file sits under `spec.update.path`, that the automation is not suspended, and that the push is not being rejected by a read-only key or a protected branch. A missing marker is silent; a rejected push appears in the object's conditions and the controller logs.
- Everything looks correct and the tag still reaches production only after 15 minutes. Is that a bug?No — it is the sum of three independent intervals: the registry scan, the automation run, and the Kustomization that applies the manifests. Shorten the intervals if the latency matters, trigger scans from a registry webhook through a notification-controller Receiver, or force a stage with `flux reconcile` when you need it now.
- How would you rule out the two image controllers not being installed at all?Look for their Deployments in the `flux-system` namespace and check whether your ImageRepository and ImagePolicy objects have any status written to them. Custom resources on a cluster without their controller are inert: the API server accepts them and nothing ever reconciles them, so there is no error anywhere to find.
saying these in an interview costs you the question
- Starts by editing manifests before checking the policy
- Treats an all-green Flux as proof nothing is misconfigured
- Forgets a commit still needs a Kustomization reconcile
- Expects push-pipeline latency from interval-based polling
- Never checks whether the image controllers are installed