skip to content

Image Update Automation

Flux scanning a registry for new tags and committing the chosen one back into the manifest repo, closing the CI-to-CD loop without a push-based pipeline. Interviewers like it because it exposes whether you have thought about who is allowed to write to the config repo.

on this pageshow

explore

questions

6

In Flux CD, which three custom resources make up image update automation, and what does each one contribute?

level: juniorimportance: must knowfreq 68%

answer

  1. three objects, one per stage
  2. scan, then select, then write
  3. the repository field carries no tag
  4. the last step is a Git commit
  5. two extra controllers, opt-in at bootstrap

basics

~20 s

Flux splits image automation across three objects: ImageRepository scans a container registry and lists its tags, ImagePolicy selects the newest tag matching a rule, and ImageUpdateAutomation writes that tag back into the Git manifests Flux deploys from.

solid answer

~40 s

Flux does not have a single "auto-update" switch; it composes three custom resources in the `image.toolkit.fluxcd.io` group. An `ImageRepository` names a repository without a tag, for example `ghcr.io/org/app`, and on `spec.interval` lists the tags it can see, using `spec.secretRef` if the registry is private. An `ImagePolicy` points at that ImageRepository and applies one selection rule — `semver`, `numerical` or `alphabetical` — publishing the winning image in its status. An `ImageUpdateAutomation` then clones the `GitRepository` you point it at, rewrites the marked fields in the manifests under `spec.update.path`, commits and pushes. Nothing here touches a running workload: the update lands in Git, and the ordinary Kustomization or HelmRelease reconcile applies it. The two controllers behind these objects, image-reflector-controller and image-automation-controller, are optional and must be installed explicitly.

go deeper

for a junior

Be able to name the three objects and say in one line what each does: scan the registry, pick a tag, commit it to Git. Also say plainly that the deploy itself still happens through normal Flux reconciliation.

for a middle

Explain why the work is split across three resources — shared scans, independent status you can inspect stage by stage — and name the two extra controllers that must be installed for any of it to run.

for a senior

Show you can reason about the end-to-end latency as a sum of independent intervals, and about the trust boundary the write step creates: something in the cluster now holds push access to the config repository.

for a principal

Own the decision of whether automated tag bumping belongs in this environment at all, and where the human checkpoint sits if it does — a narrow semver range, a pull request branch, or no automation above staging.

## The gap this fills A CI pipeline builds a container image and pushes it to a registry. Something still has to tell the cluster that a new tag exists. In a push-based pipeline, CI would hold cluster credentials and run a deploy command. Flux's image update automation closes the same loop the other way round: components running *inside* the cluster watch the registry themselves and rewrite the Git manifests, so the change arrives through the same reviewed, versioned path as every other change. Flux deliberately splits that job across three objects rather than one, because scanning, choosing and writing are different concerns with different failure modes and different owners. ## ImageRepository — the scanner `ImageRepository` names a *repository*, never a specific image version: `spec.image: ghcr.io/org/app`. Putting a tag in that field is a common first mistake and the object will be rejected. Other important fields are `spec.interval` (how often to list tags), `spec.secretRef` (a Kubernetes secret holding registry credentials), `spec.provider` for contextual cloud login, and `spec.exclusionList` for tags you never want considered. It is reconciled by **image-reflector-controller**, which caches the tag list and reports in status when it last scanned and how many tags it found. Scanning is polling by default; a notification-controller `Receiver` can also trigger a scan from a registry webhook. ## ImagePolicy — the chooser `ImagePolicy` references a scanner through `spec.imageRepositoryRef` and applies exactly one ordering rule under `spec.policy`: `semver` with a `range`, `numerical` with an `order`, or `alphabetical` with an `order`. An optional `spec.filterTags` narrows the candidate list first with a regular expression. The result is pure computation — the policy changes nothing anywhere. It simply publishes the selected image in its status, and you read it with `flux get image policy`. Separating this from the scanner means several policies can share one scan: production may track `>=1.0.0 <2.0.0` while staging tracks any tag, without listing the registry twice. ## ImageUpdateAutomation — the writer `ImageUpdateAutomation`, reconciled by **image-automation-controller**, is the only object that writes anything. It points at a `GitRepository` source through `spec.sourceRef`, checks out a branch, walks the files under `spec.update.path`, applies `spec.update.strategy: Setters`, and if any file changed it commits using the author and message template under `spec.git.commit` and pushes. ```yaml apiVersion: image.toolkit.fluxcd.io/v1beta1 kind: ImageUpdateAutomation metadata: name: flux-system namespace: flux-system spec: interval: 5m sourceRef: kind: GitRepository name: flux-system git: checkout: ref: branch: main commit: author: name: fluxcdbot email: [email protected] messageTemplate: 'chore: update images' push: branch: main update: path: ./clusters/production strategy: Setters ``` ## Nothing here deploys This is the point candidates miss most often. Image automation produces a **Git commit**, not a cluster mutation. The new tag reaches a Pod only when the Kustomization or HelmRelease that owns those manifests reconciles them in the normal way. That is also why end-to-end latency is a sum of independent intervals: the ImageRepository scan interval, plus the ImageUpdateAutomation interval, plus the deploying Kustomization's interval. It also explains why the design is safe by construction: if someone hand-edits a Deployment in the cluster, the reconciler puts it back, and the record of what should be running is always the file in Git. ## Installation is opt-in Neither controller ships in a default Flux install. You add them at bootstrap: ```bash flux bootstrap github \ --owner=my-org --repository=fleet-infra \ --components-extra=image-reflector-controller,image-automation-controller \ --read-write-key ``` The `--read-write-key` part matters: the default deploy key Flux creates is read-only, and automation has to push. ## Why three objects and not one Besides sharing scans, the split gives you three independently inspectable states. `flux get image repository` tells you whether the registry is reachable and authenticated. `flux get image policy` tells you whether your rule actually matched the tag you expect. `flux get image update` tells you whether a commit was attempted. When automation "does nothing", that ordering is the diagnostic path — and a single fused resource would have hidden which of the three stages stalled.

  • Why does Flux commit the new tag to Git instead of patching the running Deployment directly?
    Because Git is the declared desired state. A direct patch would be drift: the Kustomization that owns that Deployment would revert it on its next reconcile. Committing keeps the cluster and the repository in agreement, gives you an audit trail of exactly which tag was promoted and when, and makes rollback an ordinary Git revert rather than a manual cluster edit.
  • What has to be true of a Flux installation before any of these objects work?
    Both optional controllers must be installed — image-reflector-controller for ImageRepository and ImagePolicy, image-automation-controller for ImageUpdateAutomation. They are added with `--components-extra` at bootstrap. Without them the custom resources may be created but nothing reconciles them, so they simply sit with no status and no error anywhere obvious.
  • Roughly how long after a registry push should the new image be running?
    Add the intervals: the ImageRepository scan interval, the ImageUpdateAutomation interval, and the interval of the Kustomization or HelmRelease that applies the manifests. With 5m each that is up to about 15 minutes worst case. You shorten it with `flux reconcile`, or by driving scans from a registry webhook through a notification-controller Receiver.

saying these in an interview costs you the question

  • Says Flux watches registries out of the box, no extra controllers
  • Puts a tag in the ImageRepository image field
  • Claims Flux patches the Deployment in the cluster directly
  • Thinks one CRD does scanning, selection and writing
  • Assumes the new tag is live the instant it is pushed

context

open as a page

In a Flux ImageUpdateAutomation, how does the controller know which lines of a YAML manifest to rewrite when a new image tag is selected?

level: middleimportance: must knowfreq 56%

basics

~20 s

You mark the target lines yourself. Flux's Setters strategy only rewrites a YAML value carrying an inline comment that names an ImagePolicy, such as # {"$imagepolicy": "flux-system:app"}. Unmarked files under the update path are left untouched.

open as a page

How does a Flux ImagePolicy decide which tag in a scanned repository is the latest, and what do you do when the tags are not semver?

level: middleimportance: should knowfreq 50%

basics

~20 s

A Flux ImagePolicy applies exactly one rule under spec.policy — semver with a range, numerical with an order, or alphabetical with an order. For non-semver tags you first narrow and reshape them with spec.filterTags, whose regex capture group feeds extract.

open as a page

A new image tag was pushed to the registry 20 minutes ago, Flux reports nothing unhealthy, and the cluster still runs the old image. How do you work out where the image automation chain stopped?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Walk the chain in order and stop at the first stage whose state is wrong: did the ImageRepository scan see the tag, did the ImagePolicy select it, did the ImageUpdateAutomation commit it, and has the Kustomization applying those manifests reconciled since that commit.

open as a page

Flux's image-automation-controller pushes commits into your config repository. What write access does that require, and how do you keep those commits from re-triggering CI in a loop?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Image automation needs a Git credential with push rights to the config repository — Flux's default bootstrap deploy key is read-only, so you must bootstrap with a read-write key or supply a token. Break the CI loop by skipping runs from the bot author, branch or path.

open as a page

How does a Flux ImageRepository authenticate to a private container registry, and what changes when that registry belongs to a cloud provider such as Amazon ECR?

level: middleimportance: nice to knowfreq 32%

basics

~20 s

An ImageRepository authenticates either with spec.secretRef, pointing at a Kubernetes docker-config secret, or with spec.provider, which tells image-reflector-controller to obtain a short-lived token from the cloud platform it runs on instead of holding a static credential.

open as a page