skip to content

When a host fetches an image, who presents the read credential, and what breaks first once it expires?

level: middleimportance: should knowfreq 48%

answer

  1. the puller, not the workload
  2. bytes precede the process
  3. short-lived and scoped to read
  4. warm hosts hide the expiry
  5. it breaks at the next cold pull

basics

~20 s

The host-side component performing the pull presents it — not the workload, which does not exist until the image is on disk. When it expires, running copies keep serving and every cold host, replacement or new revision fails to obtain the image.

solid answer

~40 s

Pulling happens before any workload process exists, so the credential cannot belong to the workload. It belongs to whatever fetches on the host — the platform's host-side agent or runtime — and is usually a configured secret or machine identity exchanged at pull time for a short-lived credential scoped to reading that repository. The failure shape is what makes this worth knowing: at the moment of expiry, nothing breaks. Copies already running hold their layers on disk and need the registry for nothing. The break arrives at the next cold pull — a scale-out, a replaced host, a rolled-out reference — and under a resolve-every-time policy it also catches restarts on warm hosts. That delay is why an expired pull credential is usually discovered during a scale-out at peak.

go deeper

for a junior

Remember that something on the host does the pulling, before the application exists, and that it needs read access to the repository to do it.

for a middle

Explain the ordering that forces this — image before process — and why running copies are untouched while cold pulls fail.

for a senior

Show that you plan for latency between cause and symptom: monitor the refresher, exercise a cold pull on a schedule, and know which start-up policy the fleet runs.

for a principal

Own the rotation design across the estate: credential lifetime, scope, who refreshes it, and what proves the cold path still works between incidents.

## Who actually pulls The image has to be on disk before a process can be started from it. That ordering settles the question: the entity presenting the read credential is the **host-side component that performs the pull**, running as part of the platform, not the application. Where its credential comes from varies — a secret configured for the host, a machine identity the host proves and exchanges, or a credential the platform hands the puller alongside the workload's placement — but in every design it is the puller's credential, resolved at pull time. The best practice underneath the variation is the same everywhere: exchange whatever long-lived thing you hold for a **short-lived credential scoped to read that one repository**, so a leak is bounded in both time and reach. ## Why the workload's own identity cannot do it - **It does not exist yet.** The identity a workload is given is given to a running process, and there is no process until the layers are expanded. - **The grant is the wrong shape.** A workload's identity authorises what the application calls at runtime; read access to an image repository is an infrastructure concern belonging to the host. - **It would be circular.** If pulling required the workload's credential, obtaining that credential would require the workload to be running, which requires the image. Candidates who miss this usually propose fixing a pull failure by granting the application more permissions, which changes nothing at all. ## The shape of the failure | Event | Valid credential | Expired credential | |---|---|---| | a copy already running | keeps serving | keeps serving — its layers are on disk | | restart on a warm host, use-local-copy policy | starts | starts; nothing is requested | | restart on a warm host, resolve-every-time policy | starts | blocked at resolution | | scale-out onto a cold host | pulls and starts | cannot obtain the image | | a rolled-out new reference | pulls and starts | blocked on every host | Two directions in that table are worth stating out loud, because they are the ones people get backwards. An expired **pull** credential never stops a running copy — the bytes are already local and the registry is not in the serving path. And a fully warm fleet can hide the expiry completely, right up until it needs a host it does not already have. ## Why it is found late 1. The credential expires, or whatever refreshes it quietly stops. 2. Nothing changes: every serving copy is on a host that already holds the blobs. 3. Time passes — often days — with a green dashboard and no failing request. 4. The first cold pull happens: a host is replaced, a scaling decision fires, a new reference rolls out. 5. Capacity that was supposed to arrive does not, and it happens at exactly the moment demand made it necessary. That is a latent fault, not a slow one. The window between cause and symptom is set by how long the fleet can go without needing a new host. ## Making the path robust - **Keep credentials short-lived and refreshed by something you monitor.** The refresher's silence is the actual fault; alert on it rather than on the pull. - **Scope the credential to read, on the repositories that host actually needs.** It is an infrastructure credential, not an application one. - **Exercise the cold path on a schedule.** Have something pull the reference on a host with an empty store regularly, and alert when it fails. Nothing else surfaces this before a scale-out does. - **Know which start-up policy is configured.** It decides whether ordinary restarts are exposed to the registry at all, and therefore whether the blast radius includes warm hosts. - **Do not confuse it with the workload's own configuration.** A pull credential delivered as part of the application's settings tends to be rotated on the application's schedule, which is not the schedule the puller needs. The habit worth building is asking, of any registry-related failure: *did this need the network, or was it already on disk?* Pull-time credentials matter only on the first side of that line, and that single question sorts most of these incidents in seconds.

  • Why can the workload's own identity never authorise its image pull?
    Because the pull comes first. The layers must be on disk before any process exists, and the workload's identity is issued to a running process. Using it would be circular, and it is the wrong grant anyway — reading an image repository is the host's concern, not the application's.
  • Why does the failure appear at a scale-out rather than at the moment of expiry?
    Hosts that already hold the blobs ask the registry for nothing under a use-local-copy policy, so the serving fleet is unaffected. The first thing that must authenticate is the first cold pull — a replacement, a scale-out or a new reference — which can be days after the credential lapsed.
  • How would you catch this before a scale-out does?
    Exercise the cold path deliberately: on a schedule, pull the reference on a host with an empty content store and alert if it fails. Pair that with monitoring whatever refreshes the credential, because a refresher that stopped silently is the underlying fault rather than the pull itself.

saying these in an interview costs you the question

  • Thinks the workload's runtime identity authorises the pull
  • Assumes an expired credential stops copies already running
  • Believes a pull that worked yesterday proves today's credential
  • Thinks every restart re-contacts the registry regardless of policy
  • Expects one credential to work for repositories it was never scoped to