skip to content

Flux's image-automation-controller pushes commits into your config repository. What write access does that require, and how do you keep those commits from re-triggering CI in a loop?

level: seniorimportance: should knowfreq 44%

answer

  1. the cluster now holds a push credential
  2. bootstrap makes a read-only key by default
  3. protected branches refuse a direct push
  4. a bot identity CI can recognise
  5. suspend before you debug

basics

~20 s

Image automation needs a Git credential with push rights to the config repository — Flux's default bootstrap deploy key is read-only, so you must bootstrap with a read-write key or supply a token. Break the CI loop by skipping runs from the bot author, branch or path.

solid answer

~50 s

The write step turns a read-only GitOps setup into one where something inside the cluster can change the declared desired state, so the credential deserves scrutiny. `flux bootstrap` creates a **read-only** deploy key by default; automation needs `--read-write-key`, or a secret holding a token, SSH key or app credentials with push access — scoped to that one repository. If the target branch is protected, the controller cannot push to it at all; the usual answer is `spec.git.push.branch` pointing at a separate branch, with a pull request opened by other tooling, since Flux itself does not open PRs. The loop risk is real: an automation commit lands, CI fires on push, builds an image, a new tag appears, automation commits again. You break it by giving the automation a distinct identity in `spec.git.commit.author` and having CI ignore it — a skip token in `messageTemplate`, or branch and path filters that exclude the manifests directory. `spec.suspend` is the kill switch while you sort it out.

go deeper

for a junior

Know that the automation writes to Git and therefore needs push access, and that Flux's default bootstrap key is read-only so this has to be configured deliberately.

for a middle

Explain where the credential comes from — the GitRepository the automation references — and describe the CI retrigger loop plus at least one concrete way to break it.

for a senior

Argue the trust boundary out loud: a cluster component that can write the desired state is a path to production. Cover branch protection, the push-branch pattern, scoping, and suspension during an incident.

for a principal

Decide the organisational policy: which environments may be written automatically, where the approval gate lives, what machine identity every team uses, and how those credentials are rotated and audited.

## The trust boundary you just moved In a read-only GitOps setup, the credential the cluster holds can only *fetch*. Compromise it and an attacker learns your manifests. Once image automation is enabled, the cluster holds a credential that can *write* the repository defining what the cluster runs — and that same repository is reconciled back into the cluster automatically. A write credential is therefore, transitively, a path to changing production. That is exactly what interviewers are probing when they ask about image automation: whether you noticed. The mitigations are ordinary least privilege, applied deliberately: - **Scope the credential to one repository.** A per-repository deploy key does this naturally; an organisation-wide personal access token does not. - **Prefer a machine identity you can audit and revoke** — a deploy key or an app installation — over a human's token, which vanishes when they leave and carries their whole access footprint while they stay. - **Keep the automation's path narrow.** `spec.update.path` bounds which files it can rewrite, and its markers bound which values. A production automation restricted to one directory with a tight semver policy has a small blast radius even if everything else goes wrong. ## Getting the credential in place `flux bootstrap` provisions a deploy key that is **read-only** unless you ask otherwise: ```bash flux bootstrap github \ --owner=my-org --repository=fleet-infra \ --components-extra=image-reflector-controller,image-automation-controller \ --read-write-key ``` If you already bootstrapped read-only, automation reconciles and then fails on push with a permission error from the Git host — a loud failure, which is the good case. The ImageUpdateAutomation uses the credential of the `GitRepository` it references through `spec.sourceRef`, so there is one secret to fix, not two. ## Branch protection: the collision nobody plans for Mature repositories protect the deploy branch: no direct pushes, review required. The automation controller is a direct push. These are incompatible by design, and the resolution is a choice about who approves an image bump: - **Push straight to the deploy branch.** Requires the branch to be unprotected, or the identity to be exempt. Fastest, no human in the path. - **Push to a separate branch** with `spec.git.push.branch`, then open a pull request from it. Flux does **not** open the pull request — that is a bot, a scheduled CI job, or Git host automation you provide. You get review and a merge audit trail, at the cost of a human step and some machinery to maintain. A common compromise is to push directly in staging and go through a pull request for production, which puts the approval gate exactly where the risk is. ## The loop The failure to name explicitly: 1. Automation commits an updated tag to the config repository. 2. CI is configured to run on every push to that repository. 3. The CI run builds and pushes an image. 4. The scanner sees a new tag, the policy selects it, automation commits again. With a mono-repo holding both application code and manifests this happens on day one. Three defences, usually combined: - **Give the bot an identity.** Set `spec.git.commit.author` to something like `fluxcdbot`, and have CI skip runs from that author. - **Put a skip token in the commit message.** `spec.git.commit.messageTemplate` is a template you control; most CI systems honour a marker such as `[skip ci]` in the subject line. - **Filter what triggers CI.** Path filters that exclude the manifests directory, or branch filters if automation pushes to its own branch, are the most robust because they do not depend on the message text surviving a squash or a rebase. The message template is also where you make the commits readable. A template that names what changed turns `git log` on the config repository into a deployment history you can actually read during an incident — which is half the argument for doing image updates through Git in the first place. ## Provenance If your repository requires signed commits, the automation must sign too; Flux supports giving it a signing key so its commits verify like any other. Independently of policy enforcement, pairing a dedicated author identity with signed commits makes "who changed production at 03:00" answerable from the log. ## The kill switch `spec.suspend: true`, or `flux suspend image update <name>`, stops the automation without deleting it. Reach for it the moment you suspect a loop or a bad policy: it halts new commits while leaving the deployed state exactly where it is, so you can reason about the repository history without a controller writing underneath you.

  • Your deploy branch requires pull requests. How do you run image automation against it?
    Point `spec.git.push.branch` at a separate branch so the controller never pushes to the protected one, then have a bot or scheduled job open a pull request from that branch. Flux does not open pull requests itself, so that piece is yours to provide. You trade update latency and a little machinery for review and an approval record.
  • Why is a read-only deploy key not enough once you enable image automation?
    Because the automation's whole job is to commit. A read-only key lets the controller fetch the manifests but the push fails with a permission error from the Git host, so the policy resolves correctly and nothing ever lands. Re-bootstrap with `--read-write-key`, or replace the credential with a token or app installation scoped to that one repository.
  • What is the risk if that write credential leaks?
    Whoever holds it can commit to the repository that defines the cluster's desired state, and Flux will faithfully apply whatever they write. Scope it to a single repository, prefer a revocable machine identity, keep the automation's update path narrow, and treat rotating that secret with the same urgency as a leaked cloud credential.

saying these in an interview costs you the question

  • Assumes the bootstrap deploy key can already push
  • Gives the controller an org-wide personal access token
  • Thinks Flux opens the pull request for you
  • Ignores that automation commits retrigger CI
  • Cannot name a way to stop automation without deleting it

context