Why does the delay between restart attempts grow, and what makes that growing delay reset?
answer
- failing fast should not cost full speed
- each consecutive failure waits longer
- the growth is capped, not endless
- starting is not what resets it
- reset needs a stability window survived
basics
~20 sThe delay grows so that a workload which cannot stay up stops consuming the host and its dependencies at full speed. It resets only once an instance has stayed up long enough to count as stable, not merely because it started.
solid answer
~50 sWithout a delay, a workload that fails immediately would be restarted as fast as the host can create instances — burning processor time, re-pulling and re-opening everything it touches, and drowning the very dependency that may be causing the failure. So the wait between attempts grows, typically by doubling, up to a cap: `10s, 20s, 40s, 80s, 160s`, then the cap. The subtle part is the reset. The counter goes back to its starting value only after an instance has **run for long enough to count as stable** — not when an instance starts, and not when it merely reaches a running state. That distinction is what a telemetry ingester alive for forty-five seconds per attempt fails: it starts fine every time, so uptime looks non-zero, yet it never survives the stability window, and the delay keeps growing until it sits at the cap.
code
pseudocode · 15 linesdelay = initialDelay # 10s
loop:
startedAt = now()
run(instance) # returns when the instance ends
ranFor = now() - startedAt
if ranFor >= stabilityWindow: # e.g. 10 minutes
delay = initialDelay # stayed up long enough: counter resets
else:
delay = min(delay * growthFactor, maxDelay)
sleep(delay)
# ingester lives ~45s per attempt, stabilityWindow is 10m
# -> ranFor < stabilityWindow every time -> the growth branch always firesgo deeper
Recall that repeated failures are retried with a wait that gets longer each time, up to a limit, so a workload that cannot stay up does not consume the host at full speed.
Explain the doubling-to-a-cap schedule and, more importantly, the reset rule: the counter clears only after an instance has stayed up for a stability window, not when one starts.
Use the delay as a diagnostic. A delay sitting at its cap summarises a long history of consecutive failures, and it also tells you how wide a window you have to observe the workload while it is briefly present.
Weigh the tuning: a short window resets too eagerly and hides a chronic loop, a long one delays recovery from a transient cause, and the cap sets the worst-case time-to-recover that dependent teams will experience.
## What the growing delay is for A restart policy that says "try again" with no other constraint is a busy loop. A workload that fails in fifty milliseconds would be restarted twenty times a second, forever, and the damage is not limited to that workload: - it consumes processor time on the host that neighbouring workloads needed; - it re-establishes every connection it opens on each attempt, which hammers whatever it connects to — often the exact dependency whose slowness caused the failure; - it produces a flood of near-identical output that makes the useful evidence harder to find and costs money to collect; - it can keep a shared resource — a lock, a queue position, a claim — churning between taken and released. So platforms pair the policy with a **growing delay**: each consecutive failure waits longer than the last before another attempt is made. The growth is usually multiplicative and always capped, because an uncapped doubling would eventually mean the workload is retried once a day. ## The shape of the schedule With a ten-second start, a doubling factor and a five-minute cap, the waits fall out like this: | attempt | wait before this attempt | |---|---| | 1 | none | | 2 | 10s | | 3 | 20s | | 4 | 40s | | 5 | 80s | | 6 | 160s | | 7 and after | 5m (the cap; the uncapped value would be 320s) | Two things follow from that table. First, the loop becomes cheap quickly — after six failures the workload is costing the host almost nothing. Second, it becomes **slow to observe**: from attempt seven onward, an ingester that lives forty-five seconds per attempt is present for 45 seconds out of every 345, under fifteen percent of the time. Any client depending on it sees mostly absence, and any human waiting to catch it in the act waits five minutes per look. ## What resets the counter — and what does not This is the part interviews probe, because the intuitive answer is wrong. - **Does not reset it:** the instance starting. Creating a new instance is what the delay was waiting to do; it proves nothing. - **Does not reset it:** the instance reaching a running state. The ingester reaches running every single time and still dies a minute later. - **Resets it:** the instance running for long enough to count as stable — a **stability window** measured from the start of the attempt. The rule is deliberate. The counter is trying to estimate whether the workload is *healthy*, and the only cheap evidence of that is sustained uptime. A window is the mechanism that separates "it started" from "it is working". Platforms genuinely differ here in the details — in how long the window is, in whether the delay resets fully or decays, and in whether a workload with a defined unit of work is eventually abandoned rather than retried forever. What does not differ is the shape: growth on consecutive failure, a cap, and a reset conditioned on sustained running rather than on starting. ## Why this makes uptime a bad stability signal A workload in this state produces a genuinely confusing dashboard. Its instances start successfully; they report running; a naive uptime check sampled at the wrong moment finds them up. Meanwhile the delay between attempts is telling the true story — a delay sitting at the cap means *the workload has failed consecutively many times without a single stable run*. That number is often the most useful single fact available early in an incident, and it is one that no individual instance's own state can give you. ## What it does not do The growing delay rations the attempts. It fixes nothing, diagnoses nothing, and tells you nothing about why the attempts fail. A workload at the capped delay is failing exactly as badly as one on its second attempt; it is simply failing more cheaply and more slowly, which is the whole and sufficient point.
- Why is the growing delay capped rather than left to double indefinitely?Because an uncapped doubling would eventually retry the workload once an hour, then once a day, long after whatever broke it was fixed. The cap keeps recovery from a transient cause bounded while still making the loop cheap.
- An instance has been up for forty-five seconds. Why is that not evidence of stability?Because starting is not the thing being measured. Stability is defined by surviving a window, and forty-five seconds is exactly the uptime this workload has achieved on every previous attempt before dying.
- Why is the current delay between attempts often more useful early in an incident than the instance's own state?The instance's state is a snapshot and is 'running' for part of every cycle. A delay sitting at its cap is a summary of history: many consecutive failures with no stable run in between.
It behaves like a circuit breaker you reset by hand: each re-trip makes you wait longer before trying again, and only a long stretch of the circuit holding under load convinces you the fault is really gone — flipping the switch back on proves nothing by itself.
saying these in an interview costs you the question
- Believes the delay resets every time a new instance starts
- Reads a long gap between attempts as the platform having given up
- Thinks a growing delay fixes the cause rather than rationing the attempts
- Treats forty-five seconds of uptime as proof the workload is now stable
- Assumes the delay grows without bound instead of reaching a cap