skip to content

A fleet of 40 CI build agents behind a single NAT gateway intermittently fails image pulls during peak hours with HTTP 429 from the registry. Walk through how you diagnose it and what you change so it stops recurring.

level: seniorimportance: should knowfreq 38%

answer

  1. 429 ≠ network flake — read the log line
  2. one NAT IP = one anonymous bucket
  3. ephemeral agents = cold cache every job
  4. login → mirror → pre-bake → pay
  5. alert on 429 rate + cache hit ratio

basics

~20 s

Confirm it is registry throttling, not the network: check the 429 and remaining-quota headers, and whether pulls are anonymous. Then authenticate CI with a machine account, put a pull-through cache in front of the fleet, pin and pre-bake hot base images, and alert on 429 rate and cache hit ratio.

solid answer

~60 s

**Diagnose first.** Confirm the failure is a registry 429 rather than DNS/proxy flakiness — read the actual daemon or build log line. Check whether the agents are pulling **anonymously**: if so, all 40 share one per-IP bucket keyed to the NAT address, which explains why it correlates with peak concurrency rather than with any single job. Query the remaining-quota headers from an agent to see the bucket drain. **Then fix in layers.** 1. `docker login` with a dedicated machine account and a read-only access token, so quota is per-account and per-team, not per-egress-IP. 2. Stand up a **pull-through cache** with authenticated upstream credentials; the fleet's misses collapse to roughly one upstream pull per image. 3. Reduce demand: pin bases by digest, bake hot images into the runner image, and stop `docker pull`-ing on every step. 4. Make it visible: alert on 429 rate and cache hit ratio, so the next regression is caught before builds fail. Avoid retry storms — they extend the throttle window rather than escaping it.

code

yaml · 9 lines
yaml
steps:
  - name: Registry login (machine account, read-only token)
    run: echo "$REGISTRY_TOKEN" | docker login -u "$REGISTRY_USER" --password-stdin
    env:
      REGISTRY_USER: ${{ secrets.CI_REGISTRY_USER }}
      REGISTRY_TOKEN: ${{ secrets.CI_REGISTRY_TOKEN }}

  - name: Build
    run: docker build --pull=false -t app:${{ github.sha }} .

go deeper

for a junior

Recognise the 429/toomanyrequests message and know that authenticating the CI pull is the first step.

for a middle

Explain the shared-NAT anonymous bucket, check remaining quota, and configure login plus a mirror.

for a senior

Run the full diagnosis: confirm attribution, quantify demand, layer login/cache/pre-bake, and add metrics and alerting so it cannot silently return.

for a principal

Treat upstream registry availability as a fleet-level dependency: policy on internal registries of record, budget for paid tiers, and per-region cache topology as the fleet scales.

## Read the actual error before theorising HTTP 429 from a registry is a specific claim: the server accepted your request and refused it for quota. That rules out most of the usual suspects — DNS, MTU, proxy interception, TLS. Get the raw line from the build log or `journalctl -u docker`; `toomanyrequests` names the cause unambiguously. Distinguish it from 401/403 (auth) and 5xx (upstream trouble), because the remediation is completely different. ## Why 'intermittent, at peak' is the tell Anonymous registry quota on Docker Hub is bucketed **per source IPv4 address**. Forty agents behind one NAT gateway present one address, so the fleet shares a single counter that refills on a time window. Under light load the fleet stays under the ceiling; at peak — a merge queue, a nightly matrix, an incident-driven redeploy — concurrency multiplies pulls per unit time and the bucket empties. The failures then look random, hitting whichever job happens to pull next, which is why teams misdiagnose it as 'flaky network'. A second amplifier: ephemeral agents. If every job starts from a clean container/VM with an empty image cache, *every* job pays full pull cost. Persistent agents with warm caches quietly hide the problem until you switch to ephemeral runners — a classic 'it broke when we migrated the runners' story. ## Confirm attribution Check whether pulls are authenticated at all: is there a `docker login` step, does `~/.docker/config.json` on the agent contain an auth entry, does the Kubernetes-based runner have an image pull secret? Then measure: fetch a token for the throwaway `ratelimitpreview/test` repository and read the `ratelimit-limit` / `ratelimit-remaining` headers from an agent during peak. Watching remaining go to zero closes the case, and repeating the check after `docker login` proves the change worked. ## The fix, in cost order **1. Authenticate.** A machine account with a scoped read-only access token, injected as a CI secret, moves the fleet off per-IP accounting. This alone often ends the incident. Never use a human's personal credentials — rotation and offboarding will bite you. **2. Cache.** A pull-through mirror in the same network as the agents turns N pulls into approximately one upstream pull per image per cache window, and cuts pull latency substantially. Give the cache upstream credentials, or it becomes the new single IP hitting the anonymous ceiling. **3. Reduce demand.** Bake the ten images that account for most pulls into the agent base image or a warm cache volume. Pin base images **by digest** so builds are reproducible and cheap to reason about. Audit for pathological patterns: a `docker pull` in every matrix leg, a compose file re-pulling sidecars, `--pull=always` where it is not needed. **4. Buy headroom.** A paid registry plan or a hosted registry with generous limits is usually cheaper than the engineering time lost to flaky builds. Frame it that way to management. ## Make it stay fixed Export metrics: count of 429 responses seen by agents, mirror hit ratio, upstream pull volume. Alert on the 429 rate crossing zero rather than waiting for red builds. Add a CI preflight that fails fast with a clear message ('registry quota exhausted') instead of a confusing mid-build failure. Document the machine account and its token rotation, because an expired token silently reverts the whole fleet to anonymous pulls — and the symptom will look exactly like this incident again. ## What not to do Do not add blind retries: the window is time-based, so retries prolong exhaustion and hide the signal. Do not spread agents across extra NAT addresses to dodge per-IP counting — it is quota evasion, it violates terms, and it fails again the moment the fleet grows.

  • You added `docker login` and it still fails occasionally. What now?
    Check that every pull path is authenticated — build stages, compose-launched service containers, and any Kubernetes-based runner pulling with node credentials rather than the job's login. Then check the token itself: an expired or revoked token makes the client fall back to anonymous pulls silently. If all paths are authenticated, you are simply over the account tier and need a mirror or a higher plan.
  • How would you prove the mirror is actually helping rather than just existing?
    Instrument it: the mirror's hit ratio and its upstream pull count are the two numbers that matter. Pull a fresh image on two agents and confirm only the first produces an upstream fetch in the mirror's logs. Compare weekly upstream pull volume before and after; if it did not drop, references are probably still resolving to the upstream host.

saying these in an interview costs you the question

  • Treating 429 as a transient network error and 'fixing' it with retries or longer timeouts.
  • Not checking whether the pulls were anonymous before proposing an architecture change.
  • Rotating or adding NAT egress IPs to evade per-IP quota.
  • Assuming a mirror removes all upstream pulls, so no authentication is needed on the mirror.
  • Fixing one agent's config by hand and calling the fleet fixed.

context