skip to content

You discover that an image already pushed to a shared registry contains a live cloud credential inside one of its layers. What is your response, and why is deleting the file in a follow-up layer, or squashing the image, not a fix on its own?

level: seniorimportance: should knowfreq 34%

answer

  1. rotate first — it is disclosed the moment it is pushed
  2. whiteout hides, blob remains
  3. squash/rebuild = new digest, old one still pullable
  4. tags mutable, digests permanent
  5. delete manifests + registry GC + purge caches

basics

~20 s

Rotate the credential first — treat it as compromised. Layers are immutable and content-addressed, so a later deletion only hides the file, and squashing produces a new digest while the old one stays pullable from the registry, mirrors, caches and node image stores until it is deleted and garbage-collected.

solid answer

~50 s

**Rotate first, clean up second.** The moment a credential lands in a shared registry, assume disclosure: revoke it, issue a replacement, and check the provider's audit log for use you did not authorize. Then understand why cleanup alone fails. Layers are immutable content-addressed blobs. A `RUN rm` in a later layer adds a whiteout entry; the earlier blob still contains the bytes. Squashing or rebuilding creates a *new* image with a new digest — it does not retract the old one. Anyone who recorded the old digest, any pull-through cache or mirror, any node with the image in its local store and any registry backup still has it. So cleanup is: rebuild with a BuildKit secret mount, repoint tags, delete the affected manifests, run registry garbage collection, and purge mirrors and node caches. Finally add prevention — CI checks on build history and layers, and short-lived tokens so nothing long-lived is available to bake in.

code

bash · 9 lines
bash
docker save app:1.0 -o img.tar && mkdir -p x && tar -xf img.tar -C x
grep -rl 'AKIA' x/ | head        # bytes remain in the earlier layer blob

docker history --no-trunc app:1.0 | grep -i -e key -e token

# after rotating, remove the manifest by digest and collect blobs
crane digest registry.example.com/app:1.0
crane delete registry.example.com/app@sha256:<digest>
# then run the registry's garbage collection, and purge mirrors/node caches

go deeper

for a junior

Know the headline: rotate the credential, and understand that deleting a file in a later layer does not remove it from the image.

for a middle

Explain whiteouts, content-addressed digests versus mutable tags, and why a rebuild does not retract the old image.

for a senior

Run the incident: rotate, assess with provider audit logs and registry pull logs, rebuild, delete plus garbage-collect, purge caches, and add pipeline detection.

for a principal

Design the response so it is survivable by default — short-lived scoped credentials, blast-radius limits, mandatory artifact scanning, and a documented rotation SLA.

## Order of operations **1. Rotate.** A credential in a shared registry is disclosed. Revoke it and issue a replacement before touching the image, because everything else takes time and the credential is live the whole while. **2. Assess.** Determine blast radius: what the credential could access, and whether it was used. Pull the provider's audit log for the window between the image push and the revocation, and check who could pull the image — an internal registry usually means "anyone in the org", a public registry means everyone, plus the automated scrapers that watch public registries for exactly this. **3. Rebuild.** Fix the Dockerfile so the credential never enters a layer — a BuildKit secret mount, or a builder stage that never reaches the final image — and publish a clean image. **4. Remove what you can.** Delete the offending manifests, then run **garbage collection** on the registry to free the underlying blobs. Many registries require deletion to be explicitly enabled, and blobs shared with other images are not collected. Repoint tags, and purge pull-through caches, CI caches, mirrors and node-local image stores, all of which can serve the old content. **5. Prevent.** Add a CI gate that inspects build history and layers for credential patterns, forbid `--build-arg` for anything sensitive, and move builders onto short-lived tokens obtained per build so there is no durable credential available to leak. ## Why the popular "fixes" do not work **Deleting in a later layer.** Overlay filesystems represent a deletion as a whiteout marker in the newer layer. The lower layer's tarball is untouched and is still pulled by every client; `docker save` plus `tar` recovers the file. The final `docker run` view looks clean, which is exactly what makes this mistake persistent. **Squashing.** Squashing produces a new image whose layers are merged. It does not modify or delete the previously pushed image — that digest still resolves in the registry, and any pipeline, deployment manifest or SBOM that pinned it keeps pulling it. It is also easy to squash and still ship the credential if it was set as `ENV`, since the config travels with the new image. **Moving the tag.** Tags are mutable pointers; digests are not. Repointing `:latest` at a clean build leaves the old digest fully accessible. **Deleting the tag without GC.** Untagging typically leaves the manifest and blobs on disk until garbage collection runs, and some registries retain them by policy for a period. ## The honest framing What you can actually control is the *credential*, not the *bytes*. Copies of a distributed artifact cannot be recalled: mirrors, caches, developer laptops, backups and any third party who pulled it are outside your reach. This is why the correct instinct is rotation-first, and why prevention has to be structural — build-time secret mounts, CI scanning of history and layers, and credentials scoped and short-lived enough that disclosure is survivable. A candidate who leads with "delete the layer" has missed the model; one who leads with "rotate, then assess, then clean" has it. ## Detection, so it does not recur Scan images rather than only source: run a secret scanner over the built image's history and layer contents in the pipeline, fail the build on a hit, and repeat the scan periodically across the registry to catch images built before the gate existed.

  • The image was only pushed to a private internal registry. Does that change the response?
    It changes the urgency and the blast-radius estimate, not the decision to rotate. An internal registry is typically readable by the whole engineering org and by every CI job, which is a far wider audience than the service the credential belongs to. Rotate, then use the registry's pull logs to bound who actually fetched it.
  • How would you detect this class of mistake automatically in a pipeline?
    Scan the built artifact, not just the source: run a secret scanner over the image's build history and the contents of each layer, and fail the build on a hit. Complement it with a policy that rejects sensitive values passed as build arguments, and a periodic sweep of existing registry images so pre-existing leaks surface rather than waiting for an incident.

It is a misprinted newspaper: you can stop the presses and print a correction, but every copy already delivered stays on someone's doorstep — so you change the story, not the copies.

saying these in an interview costs you the question

  • Leading with 'delete the layer' instead of rotating the credential
  • Believing squashing or rebuilding retracts the previously pushed digest
  • Assuming an untagged manifest is gone before garbage collection has run
  • Forgetting mirrors, pull-through caches, node-local image stores and backups
  • Treating a private registry as a safe audience and skipping rotation

context