skip to content

Your Kyverno verifyImages rule rewrote a tag to a digest and GitOps now reports permanent drift — why?

level: seniorimportance: should knowfreq 38%

answer

  1. the rule does not only decide
  2. check what mutateDigest defaults to
  3. desired in Git, live in the cluster
  4. the diff is one field, forever
  5. self-heal turns it into a loop

basics

~10 s

verifyImages mutates the admitted object by default, appending the resolved digest to the tagged reference. Git still says only the tag, so a GitOps controller diffing desired against live reports permanent drift.

solid answer

~50 s

A `verifyImages` rule does not only decide; on success it rewrites the image reference in the admitted object to include the resolved digest, because `mutateDigest` defaults to true. The object stored in the cluster therefore reads `repo:tag@sha256:...` while the manifest in Git still reads `repo:tag`. A GitOps controller reconciles by diffing the manifest it manages against the live object, so that one field differs on every sync and the application never reaches a synced state — and if it self-heals, it reapplies the tag, admission mutates it again, and you get a churn loop. There are three honest fixes: tell the controller to ignore that field for those resources, resolve digests at release time so Git already carries what the cluster will hold, or set `mutateDigest: false` and accept that what you verified and what the node later pulls are no longer pinned to the same bytes.

go deeper

for a junior

Know that a Kyverno image-verification rule can change the object it admits, appending the resolved digest to the image reference rather than just approving it.

for a middle

Explain the diff mechanically: mutateDigest defaults to true, the live object gains a digest the Git manifest lacks, and reapplying re-triggers the mutation.

for a senior

Weigh the three remedies out loud — scoped diff exclusion, digests resolved at release time, or disabling mutation — and say what each one costs.

for a principal

Own the boundary between two systems that both claim the object, and decide where in the release path digest resolution belongs so this never becomes a per-team workaround.

## Verification has a side effect Most policy rules answer yes or no. A `verifyImages` rule does something extra: after it verifies an image it *changes the object it admitted*. The field `mutateDigest` controls this and defaults to true, so unless you turned it off, an image written as `registry.example.com/apps/api:1.4.2` is admitted as `registry.example.com/apps/api:1.4.2@sha256:...`, with the digest the engine resolved while doing the verification. The reason is straightforward: the engine verified a specific set of bytes, and pinning the admitted reference to that digest is how the decision stays attached to the thing it was made about. But the side effect is that the object in the cluster is no longer the object you applied — and that is exactly the invariant a GitOps controller is built to defend. ## Why the drift is permanent A GitOps controller reconciles by comparing the desired state — the manifests in the repository — against the live state of the resources it manages. Its whole model assumes that applying the desired state makes live equal to desired. Admission-time mutation breaks that assumption for one field: - Git says `image: registry.example.com/apps/api:1.4.2`. - The cluster holds `image: registry.example.com/apps/api:1.4.2@sha256:...`. - The diff is non-empty, so the app is reported out of sync. - Reapplying does not fix it, because the reapplied object goes through admission again and gets the digest appended again. If the controller is configured to self-heal, this becomes a loop: it detects drift, applies the tag form, admission mutates it back, it detects drift again. The cost is not just noise on a dashboard — a permanently out-of-sync application trains the team to ignore the sync status, and the next time drift means something real, nobody looks. ## The three fixes, and what each costs **Ignore the field.** Configure the GitOps controller to exclude the image field of the managed resources from diffing. Cheap and effective, but it is a targeted blindfold: from then on, that controller will not tell you if the image in the cluster differs from the image in Git for any reason, mutation or otherwise. Scope it as narrowly as the tool allows. **Put the digest in Git.** Have the release process resolve the tag to a digest and commit the pinned reference, so the manifest already says what the cluster will hold and the mutation is a no-op. This is the cleanest end state: desired and live agree, the reference in Git is unambiguous, and the change history records exactly which bytes were deployed. It costs you an automated step in the release path and a repository that now churns on every image bump — and the humans reading those pull requests can no longer tell at a glance what changed, which is a real reviewability cost. **Turn mutation off.** Setting `mutateDigest: false` makes the drift disappear because nothing is rewritten. What you give up is the link between the reference that was verified and the reference the node resolves when it pulls, which is the whole reason the mutation exists. ## Second-order effects worth naming The rewrite changes more than a diff. Anything reading the live spec now sees a two-part reference: dashboards and scripts that parse the image string, alerting that groups by image, a human running `kubectl get` to see what version is deployed. A digest is not human-legible, so "which release is this?" gets harder to answer from the cluster alone — the tag is still present in the mutated form, which helps, but tooling that split on `:` will find more than it expected. It also affects how a rollback reads. The object that was admitted carries a digest, so re-applying an old manifest by tag produces a fresh resolution rather than the bytes that ran last time. If reproducing exactly what ran is a requirement, that argues again for digests in Git. ## Answering it well The question is really testing whether you know that a decision engine here is also a mutating one, and whether you can reason about the interaction between two systems that each believe they own the object. Name the default, explain the diff mechanically, then give the three options with their costs and say which you would pick — usually digests in Git for the services where reproducibility matters, and a narrowly scoped diff exclusion everywhere else while you get there.

  • Which of the three fixes would you pick and why?
    Resolve digests at release time and commit them, for anything where reproducing what ran matters — desired and live then agree and the history is unambiguous. A narrowly scoped diff exclusion is the pragmatic stopgap for the rest of the estate while the release path is changed. Disabling the mutation is last, because it removes the reason it exists.
  • What does committing digests to Git cost the humans reviewing those pull requests?
    Legibility. A bump from one digest to another says nothing about what changed, so reviewers lean entirely on the automation that produced it. Teams usually keep the tag alongside the digest in the reference so the pull request still reads as a version change, with the digest carrying the precision.
  • Does the drift also affect tooling that reads the running spec?
    Yes. Anything parsing the image string sees a tag plus a digest, so scripts that split on a colon or group dashboards by image need to cope with the longer form. It also makes "which release is running?" less readable directly from the cluster, which is a small but real operational cost.

saying these in an interview costs you the question

  • Believes verifyImages only allows or denies, never mutates
  • Thinks reapplying the manifest will clear the drift
  • Disables the mutation without saying what is lost
  • Blames the GitOps tool for a diff it correctly reported
  • Suggests ignoring all diffs on the resource rather than one field

context