skip to content

Sixty cold hosts pull the same large image at once and the source starts throttling — what do you change?

level: seniorimportance: should knowfreq 55%

answer

  1. sixty empty stores, one source
  2. limits can count requests
  3. lockstep retries re-trigger the limit
  4. move bytes nearer, or pre-warm
  5. jitter, cap, shrink, cache

basics

~20 s

A burst of identical cold pulls is being rate-limited at the source, so scale-out stalls behind the slowest transfer. Cut demand rather than retry harder: a nearer cache, pre-warmed hosts, a smaller image, capped concurrency, and backoff with jitter.

solid answer

~50 s

Every one of those hosts has an empty content store, so each transfers the whole image, and they all hit one source in the same second. A registry limits what any one client may ask for — commonly per credential, per address or per repository — and answers over-rate clients with a throttling response. The trap is the retry loop: sixty clients that failed together retry together and re-trigger the limit, so naive retrying makes the burst last longer. Four levers actually help. Put the bytes nearer, so hosts read from a cache that fetches from the source once. Pre-warm hosts so scale-out finds the layers already present. Shrink what must move, ideally by sharing a base the fleet already holds. And cap concurrent pulls per host while backing off with jitter so the retries spread out.

code

pseudocode · 16 lines
pseudocode
attempt = 0
while attempt < maxAttempts:
    waitForPullSlot(maxConcurrentPullsPerHost)      # cap simultaneous transfers
    result = pull(imageReference)
    releasePullSlot()

    if result.ok:
        return success
    if not result.throttled:
        return fail(result)                         # missing repo or bad credential: waiting cannot help

    attempt = attempt + 1
    delay = min(baseDelay * 2 ^ attempt, maxDelay)
    sleep(delay * random(0.5, 1.5))                 # jitter: sixty hosts must not retry together

return fail("still throttled after maxAttempts")

go deeper

for a junior

Recognise that a throttling response is a deliberate refusal by the source, not an outage, and that retrying immediately makes it worse.

for a middle

Explain why the burst is correlated, what the limit is keyed to, and why jittered backoff behaves differently from a longer fixed delay.

for a senior

Show the diagnosis and the ordering of fixes: remove hosts from the source's path first with a cache or pre-warming, then shrink the image, then cap concurrency to survive the rest.

for a principal

Treat distribution as a capacity plan: how much of the fleet may go cold at once, what warm floor you fund, and whether a cache is a component you are willing to own.

## Why sixty pulls are worse than sixty times one pull One cold pull is bounded by the image size and the link. Sixty simultaneous cold pulls of the *same* image are bounded by something else: a single source serving sixty copies of identical bytes, under a limit that exists precisely to stop one client doing that. The transfer that took ninety seconds alone does not take ninety seconds sixty times over — it queues, gets refused, and retries. The demand is also perfectly correlated. Nothing staggers it: a batch of work lands, the scaling decision fires once, and every new copy is placed within seconds of the others onto hosts that are all equally empty. ## What the source is actually limiting - **Requests, not only bytes.** One pull is many requests — a resolution, a manifest, then one per missing blob. A fleet can trip a request-rate limit while its total transferred volume looks modest. - **Scoped to something.** Limits are commonly keyed to a credential, a source address or a repository. Which one is in play decides whether adding hosts behind the same egress helps or hurts. - **Anonymous versus authenticated.** Many registries hold unauthenticated clients to a tighter ceiling than authenticated ones, so pulling with a read credential can raise the limit without changing anything else. - **Bandwidth as well.** Even under the limit, sixty simultaneous transfers share the source's outbound capacity and your own egress, so each one slows down. ## The four levers | Lever | What it reduces | What it costs | |---|---|---| | a nearer cache the hosts read from | requests reaching the source, and the distance bytes travel | another component to run and keep healthy; it still fetches once, cold | | pre-warming hosts before the burst | the transfer at scale-out time | disk on every host, and staleness when the reference moves | | a smaller image, ideally a shared base | bytes and time per host | build effort; a unique small image can lose to a shared larger one | | capped concurrency plus jittered backoff | the peak request rate | a longer tail before the last copy is serving | The first two are the structural fixes, because they change how many hosts must reach the source at all. The last is damage control that makes the burst drain instead of thrashing. Ordering matters when you have one afternoon: pre-warming or caching removes the problem, concurrency caps merely survive it. ## Retrying is the trap 1. Sixty clients are refused within the same second. 2. Each waits a fixed interval and retries — so all sixty retry in the same second again. 3. The limit fires again, having burned a full round trip per host, and the queue is no shorter. 4. Randomising each wait around its nominal delay breaks the lockstep: some hosts land in each window, succeed, and stop competing. The delay length matters less than the spread. A long fixed delay still keeps the fleet synchronised; a short jittered one drains steadily. Growing the delay between attempts on top of the jitter keeps a persistent refusal from hammering the source indefinitely. One more discipline belongs here: only back off on a **throttling** answer. A refusal for a missing repository or an invalid credential will not become true by waiting, and retrying it wastes the budget that a genuinely throttled pull needs. ## What this means for adding capacity under load On a cold host, the pull sits between the placement decision and the process starting. Capacity added during a spike therefore arrives one full transfer late — and *later than that* when the whole fleet is competing for the same source. Three consequences follow: - **Measure scale-out latency on an empty store.** The figure from a warm host is not the one you will get when it matters. - **Keep a floor of warm capacity** rather than scaling from near zero, so the first burst of work is not waiting on the network. - **Treat the image size as a latency budget item.** Halving what has to move halves the cold portion of every scale-out for as long as the image lives. The underlying point is that at burst scale, distribution — not the runtime, not the scheduler — is the bottleneck, and every fix is some form of moving the bytes closer or having fewer of them to move.

  • Why does jitter matter more than simply waiting longer?
    A fixed delay preserves the synchronisation that caused the problem — clients that failed together retry together, whatever the delay. Randomising each wait spreads attempts across the window, so a few succeed on every round and the queue drains. A longer fixed delay just makes the same collision happen later.
  • Fewer layers or smaller layers — which helps this burst more?
    Smaller helps the transfer, since total bytes dominate the time. Fewer helps where the limit counts requests. The larger win is usually neither: a base the hosts already hold means only the small unique layers move, turning a cold pull into a partly warm one.
  • The source is healthy and the limit is generous — sixty pulls are still slow. What now?
    Then it is bandwidth, not policy: sixty simultaneous transfers share the source's outbound capacity and your own egress link. A nearer cache still helps, because it turns sixty long-haul transfers into one plus sixty local ones, and capping concurrency stops the transfers starving each other.

saying these in an interview costs you the question

  • Retries immediately and harder when a pull is throttled
  • Thinks adding more hosts speeds up a throttled scale-out
  • Reads a throttling response as the registry being down
  • Assumes only total bytes matter, never the request count
  • Assumes a nearer cache removes the first transfer too
  • Plans autoscaling as if a cold host started instantly