skip to content

How does a Flux ImageRepository authenticate to a private container registry, and what changes when that registry belongs to a cloud provider such as Amazon ECR?

level: middleimportance: nice to knowfreq 32%

answer

  1. the controller scans, not the kubelet
  2. a docker-config secret in the same namespace
  3. cloud tokens expire in hours, not months
  4. provider means fetch a token, hold nothing
  5. stale results, one red condition

basics

~20 s

An ImageRepository authenticates either with spec.secretRef, pointing at a Kubernetes docker-config secret, or with spec.provider, which tells image-reflector-controller to obtain a short-lived token from the cloud platform it runs on instead of holding a static credential.

solid answer

~50 s

By default an `ImageRepository` scans anonymously, which only works for public repositories. For a private registry you set `spec.secretRef` to a Kubernetes secret of type `kubernetes.io/dockerconfigjson` — the same shape as an image pull secret, but read by image-reflector-controller, not by the kubelet, so a pull secret already attached to your Pods does not automatically cover the scanner. For cloud registries this static approach ages badly: Amazon ECR issues authorisation tokens that expire after roughly twelve hours, so a stored secret rots and teams used to refresh it with a CronJob. Setting `spec.provider` to the platform instead enables contextual login: the controller acquires and refreshes a token itself using the cloud identity it runs under, and no long-lived credential exists in the cluster. Two more fields matter occasionally — `spec.certSecretRef` for a registry behind a private CA, and `spec.insecure` for a plain-HTTP registry, which should stay off outside a lab.

code

yaml · 9 lines
yaml
apiVersion: image.toolkit.fluxcd.io/v1beta2
kind: ImageRepository
metadata:
  name: app
  namespace: flux-system
spec:
  image: 123456789012.dkr.ecr.eu-west-1.amazonaws.com/app
  interval: 5m
  provider: aws

go deeper

for a junior

Know that scanning a private registry needs credentials, that they are configured on the ImageRepository object, and that the usual form is a Kubernetes docker-config secret in the same namespace.

for a middle

Explain that the controller authenticates, not the kubelet, and why cloud registries push you towards spec.provider contextual login rather than a stored token that expires within hours.

for a senior

Recognise the stale-cache failure signature — one red condition on the scanner while policies and automations stay green — and argue for read-only, narrowly scoped scanner credentials.

for a principal

Set the standard across teams: contextual login wherever the platform supports it, IAM as the audit surface, per-team scoping of scanner access, and no long-lived registry credentials living in cluster secrets.

## Who is actually authenticating The first thing to get straight is that this credential belongs to **image-reflector-controller**, running in the cluster and talking to the registry's HTTP API to list tags. It is not the kubelet pulling an image. That distinction explains a recurring surprise: a workload that pulls fine using an `imagePullSecrets` entry can sit next to an ImageRepository that cannot scan at all, because nothing shares that secret with the controller. They are two different clients with two different credentials, and both must be configured. It also shapes the least-privilege answer. The scanner never needs push rights and never needs to pull layers; read access sufficient to list tags is the whole requirement, and the credential should be scoped that narrowly. ## Option one: a static secret ```yaml apiVersion: image.toolkit.fluxcd.io/v1beta2 kind: ImageRepository metadata: name: app namespace: flux-system spec: image: ghcr.io/org/app interval: 5m secretRef: name: registry-creds ``` The referenced secret is an ordinary `kubernetes.io/dockerconfigjson`, the kind `kubectl create secret docker-registry` produces. It must live in the same namespace as the ImageRepository. This is the right answer for a registry that issues long-lived robot accounts — a self-hosted registry, or a token-based account on a public registry. The cost is the usual one for static credentials: it sits in the cluster, it must be rotated, and rotating it is a change nobody remembers to schedule. If it is committed to Git it needs sealing or encryption like any other secret. ## Option two: contextual cloud login Cloud registries mostly do not issue long-lived credentials at all. Amazon ECR hands out an authorisation token valid for about twelve hours; Azure and Google have their own short-lived token flows. Storing one in a secret means it stops working the same day, which is why the older Flux pattern involved a CronJob that re-ran the cloud CLI every few hours and rewrote the secret — a moving part that fails quietly at 3 a.m. `spec.provider` removes it: ```yaml spec: image: 123456789012.dkr.ecr.eu-west-1.amazonaws.com/app interval: 5m provider: aws ``` The controller now authenticates using the identity it already has on the platform — the workload identity or instance role attached to it — and refreshes tokens on its own. There is no registry credential in the cluster to leak or rotate, and access is granted by cloud IAM policy, where it is auditable alongside everything else. The prerequisite is that the controller actually has such an identity: on a self-managed cluster off-platform, contextual login has nothing to bind to and a static secret is still the answer. ## Transport details Two smaller fields exist for awkward registries. `spec.certSecretRef` supplies a CA certificate for a registry serving a certificate your cluster does not trust, which is common for internal registries with a private PKI. `spec.insecure` allows plain HTTP; it disables transport security for that scan, which means credentials and responses cross the network in the clear, so it belongs in a lab and nowhere else. ## Failure signature An authentication problem surfaces on the ImageRepository object itself: a ready condition of false with the registry's error, visible from `flux get image repository`. What makes it worth recognising is the downstream effect. The ImagePolicy keeps ranking the tag list from the last successful scan and keeps reporting a perfectly plausible selected image; the automation keeps running and finds nothing to change. So the system looks healthy everywhere except one object, while quietly serving results from an increasingly stale cache. When someone says a new tag was never picked up, the scanner's own condition is the first thing to read. ## Scoping the blast radius Because the credential is per-ImageRepository through `secretRef`, you can give different teams' scanners different credentials in different namespaces rather than one cluster-wide registry account. With contextual login the equivalent control lives in IAM. Either way the goal is the same: a scanner credential that can list tags on the repositories that team owns and nothing else.

  • Your Pods already pull from this ECR repository. Why can the ImageRepository still fail to scan it?
    Different client, different credential. The kubelet uses the Pod's `imagePullSecrets` or the node's instance role; image-reflector-controller uses `spec.secretRef` or `spec.provider` on the ImageRepository itself. Nothing propagates one to the other, so a working pull tells you the registry and network are fine but says nothing about whether the scanner is authenticated.
  • What did teams do before contextual login existed, and why was it fragile?
    They ran a CronJob that called the cloud CLI to fetch a fresh registry token and rewrote the docker-config secret every few hours, because ECR tokens expire in roughly twelve hours. It works until the job fails, its own credential expires, or someone deletes it — after which scans go stale silently and the policy keeps serving the last known tag list.
  • What access should the scanner's credential have?
    Read-only, enough to list tags on the repositories it scans, and nothing more. It never pushes and never pulls layers. Scope it per repository or per team namespace rather than issuing one cluster-wide registry account, so a leaked scanner credential does not become a way to publish images.

saying these in an interview costs you the question

  • Assumes the Pod's imagePullSecrets cover the scanner
  • Stores a long-lived ECR token in a Secret
  • Grants the scanner push access to the registry
  • Turns on insecure to get past a TLS error
  • Thinks a failed scan makes the policy report nothing

context