skip to content

You are asked to make signature verification mandatory before any container image runs in production, across many teams and with substantial third-party images in use. How do you roll that out, and what will break?

level: principalimportance: should knowfreq 28%

answer

  1. enforce at admission, mutate tag → verified digest
  2. policy = which identity, for which images
  3. third-party: mirror + re-sign internally
  4. audit mode → burn down → fail closed per namespace
  5. break-glass: time-boxed, owned, audited

basics

~20 s

Enforce at admission, not in CI: a policy engine verifies signatures over digests against expected signer identities and rewrites tags to verified digests. Roll out in audit mode first, sign your own builds, mirror and re-sign third-party images internally, then fail closed per namespace with a documented break-glass path.

solid answer

~60 s

**Enforce where workloads are admitted**, because CI-side checks only bind teams that choose to run them. A policy engine (sigstore policy-controller, Kyverno `verifyImages`, or equivalent) intercepts workload creation, resolves the image to a **digest**, verifies a signature from an **expected identity**, and mutates the reference to the verified digest. Rollout sequence: 1. **Inventory** every image running in production, and how much is third-party. 2. **Sign your builds** — keyless in CI, signature made over the pushed digest. 3. **Third-party images**: mirror them into an internal registry and **re-sign them internally** after vetting. This turns 'upstream doesn't sign' from a blocker into a documented gate. 4. **Audit mode** for weeks; drive the violation count to zero; publish the dashboard per team. 5. **Fail closed namespace by namespace**, starting with a willing team; exclude system namespaces deliberately, not accidentally. 6. **Break-glass**: a time-boxed, audited exemption, with an owner and an expiry. What breaks: unsigned third-party images, mirrored images whose signatures were left behind, registry GC pruning signature objects, controllers pulling images the policy never sees, and cluster-wide outages if the policy fails closed on verification-service unavailability.

code

yaml · 22 lines
yaml
apiVersion: kyverno.io/v1
kind: ClusterPolicy
metadata:
  name: verify-internal-images
spec:
  validationFailureAction: Audit   # flip to Enforce per namespace after burn-down
  rules:
    - name: check-signature
      match:
        any:
          - resources:
              kinds: [Pod]
      verifyImages:
        - imageReferences:
            - "registry.internal/*"
          mutateDigest: true
          required: true
          attestors:
            - entries:
                - keyless:
                    subject: "https://github.com/acme/*/.github/workflows/release.yml@refs/tags/*"
                    issuer: "https://token.actions.githubusercontent.com"

go deeper

for a junior

Know that verification should happen automatically before workloads run, not manually, and that images should be deployed by digest.

for a middle

Describe an admission policy that verifies signatures over digests against an expected identity, and why audit mode comes before enforcement.

for a senior

Run the migration: inventory, pipeline signing, mirroring and re-signing third-party images, per-namespace enforcement, and the failure modes around mirrors, GC and verifier availability.

for a principal

Own the trust and risk model: which identities count, graduated fail-closed scope, break-glass governance, attestation requirements beyond a bare signature, and the metrics that show whether the programme is decaying.

## Enforce at admission, not in the pipeline A verification step in CI is a good hygiene check and a bad control: it protects only the teams that run it, and it does not stop anything that reaches the cluster by another route. The control must sit where workloads are **admitted** — an admission policy that verifies before the workload is allowed to exist. Concretely: intercept the object, resolve each image reference to a digest, verify signatures against a policy, and **mutate the reference to the verified digest** so what runs is what was verified. ## Decide the policy before the tooling The engineering that matters is not which engine you install, it is answering: **whose signature counts, for which images?** For internal builds, the answer is a specific workflow identity and issuer per repository — not 'any signature'. For third-party images, upstream signing coverage is uneven, and even where it exists you must decide whether upstream's identity is one you want to trust directly. The pattern that scales: **mirror third-party images into an internal registry and re-sign them internally** after whatever vetting you require (scan, licence check, human approval). Now there is exactly one trust rule — 'signed by us' — plus an auditable record of why each external image was admitted. It also removes an availability dependency on upstream registries. ## Sequence the rollout so it cannot fail loudly 1. **Inventory.** Enumerate every image actually running: registries, tags, whether signed. The share of third-party images determines the real cost. 2. **Produce signatures.** Get every internal pipeline signing by digest. This is usually weeks of small PRs, not a design problem. 3. **Audit mode.** Run the policy in warn/report mode cluster-wide. Publish a per-team dashboard of would-be denials. Nothing is blocked; everything is visible. 4. **Burn down.** Work the list: unsigned internal builds get pipelines fixed; third-party images get mirrored and re-signed; genuinely exceptional cases get named exemptions with owners. 5. **Fail closed incrementally.** Flip enforcement per namespace, starting with a team that helped build it. Keep system namespaces (cluster add-ons, CNI, CSI) explicitly scoped — decide whether they are exempt or covered, and write it down; an accidental deny there is a cluster outage. 6. **Break-glass.** A documented, time-boxed, audited bypass. Without it, the first production incident produces a permanent global exemption instead of a temporary one. ## Failure modes to design for - **Signatures lost in transit.** Signatures are separate objects in the repository; mirroring or promoting an image without carrying referrers leaves them behind and verification fails on a perfectly good image. - **Registry garbage collection.** Retention rules that prune untagged manifests can delete signature objects. This produces the worst kind of incident: images that verified yesterday and do not today. - **Availability of the control plane.** A fail-closed policy that also fails closed when its own webhook or the registry is unreachable can stop all deployments — and, worse, block pod recreation during an unrelated incident. Decide the failure policy per namespace deliberately, keep the verifier highly available, and cache verification results. - **Air-gapped and restricted networks.** Verification may need transparency-log and trust-root material that the cluster cannot reach; plan bundled/offline verification or a self-hosted stack. - **Paths the policy does not see.** Images pulled by node-level components, static pods, or anything created outside the admission path. Enumerate them. ## What good looks like Beyond a boolean, mature enforcement adds **attestations**: require not only a signature but a provenance attestation naming the source repo and build system, and possibly a scan attestation under a freshness bound. That moves the guarantee from 'someone we trust signed it' to 'this artifact was built from this source by this pipeline'. Measure it: percentage of workloads running verified digests, count of active exemptions with age, and time-to-remediate for expiring ones. Exemption age trending upward is the leading indicator that the programme is decaying. ## The honest tradeoff Mandatory verification adds a hard dependency in the deployment path and real friction for third-party software. It buys a genuine control against running artifacts nobody authorised. State that tradeoff plainly, and pick fail-closed scope accordingly — the credible answer is graduated enforcement with measured exemptions, not a flag flipped cluster-wide on day one.

  • Upstream publishes an image you depend on and it is not signed at all. What do you do?
    Do not weaken the policy to allow unsigned images. Mirror that image into your internal registry, pin it by digest, run it through whatever vetting you require, and re-sign it with your own identity. The policy still says 'signed by us', and the record of why that external artifact was admitted is auditable. It also removes a runtime dependency on the upstream registry.
  • How do you avoid the enforcement policy itself causing an outage?
    Scope failure behaviour deliberately: keep system namespaces explicitly handled, run the verifier highly available with cached results, and be clear about what happens when the registry or transparency log is unreachable. Fail-closed on unavailability can block pod recreation during an unrelated incident, so many teams fail closed on 'signature absent or invalid' but degrade gracefully on 'verifier cannot reach its dependencies', backed by alerting and a time-boxed break-glass.
  • Signature verification passes but the deployed manifest still names a tag. Is that a problem?
    Yes, unless the policy rewrote the reference. If the workload runs a tag, the tag can move between verification and pull, or resolve differently on a later reschedule, so the running artifact may never have been verified. Enabling digest mutation at admission closes that window and also makes the running artifact auditable.

It is like requiring ID at the door of a building that already has thousands of people inside: you first log who lacks a badge, issue badges, and only then start turning people away — floor by floor, with a documented way to let the fire brigade in.

saying these in an interview costs you the question

  • Putting the only verification step in CI, where teams can skip it and out-of-band deploys bypass it.
  • Turning enforcement on cluster-wide in one step without an audit phase.
  • Relaxing policy to 'any valid signature' to accommodate unsigned third-party images.
  • Forgetting that mirroring and registry garbage collection can remove the signature objects.
  • Having no break-glass procedure, guaranteeing a permanent blanket exemption after the first incident.
  • Ignoring system namespaces and node-level image pulls that never pass through admission.

context