Your machine-token cache is cold on every site controller after a deploy and all wake on the same half-hour boundary — how do you avoid a burst at the authorization server?
answer
- one miss should cause one acquisition
- replace before it dies, not after
- margin must beat skew plus latency
- jitter or the fleet re-synchronises
- empty cache means fail closed
basics
~20 sCoalesce concurrent misses so one acquisition per cache key serves every waiting caller, refresh ahead of expiry in the background while the current token is still served, jitter the refresh point and the wake boundary so instances do not re-synchronise, and warm the cache before an instance is declared ready.
solid answer
~60 sThree mechanisms, each fixing a different burst. **Single-flight** bounds one process to one in-flight acquisition per cache key: forty workers finding the entry empty produce one request and share its result, rather than forty. **Refresh-ahead** replaces the token at a fraction of `expires_in` — a margin comfortably larger than clock skew plus your worst acquisition time — in the background, while the still-valid token continues to be served, so expiry never lands on a poll. **Jitter** is what keeps the fleet from re-synchronising: instances deployed together refresh together unless the refresh point is spread across a window, and the half-hour wake boundary needs the same treatment. Pre-warm on startup, before readiness, so a deploy does not move an acquisition round trip onto the first poll's tail latency. When the authorization server throttles, keep serving the valid token you already hold, back off with a cap and jitter, and alarm as the remaining life shrinks — an empty cache and a failing acquisition means failing closed, not calling the meter service without a token.
code
pseudocode · 22 linesfunction tokenFor(key):
entry = cache.get(key)
if entry != null and now() < entry.serveUntil:
return entry.token # fresh enough: nothing to do
if entry != null and now() < entry.notAfter:
singleFlight.startIfAbsent(key, acquire) # background, do NOT wait
return entry.token # still valid: keep serving it
return singleFlight.join(key, acquire) # cold or expired: one acquisition, all wait
function acquire(key):
for attempt in 1..maxAttempts:
result = obtainToken(key) # the request itself is out of scope here
if result.ok:
return cache.put(key, result, serveUntil = refreshPoint(result, jitter()))
if result.throttled:
sleep(min(cap, backoff(attempt)) + jitter())
else:
raise AcquisitionFailed(result) # joiners fail with this, they do not retry individually
raise AcquisitionFailed("exhausted") # caller has no token: fail the poll closedgo deeper
Know the shape of the problem: many workers, one empty cache entry, and a token endpoint that does not want one request per worker.
Explain how coalescing and refreshing early differ, and why the refresh margin must be bigger than clock skew plus the time an acquisition actually takes.
Demonstrate the operating picture: what you serve while the authorization server throttles, when a degraded refresh becomes an outage, and which metrics tell you which of the three bursts you are looking at.
The judgment is how much availability you are buying from a dependency you do not run: token lifetime, refresh budget, fleet-wide stagger and whether a shared cache is worth its own failure mode.
## Three bursts, three different fixes A fleet of site controllers polling a utility's meter service produces token-endpoint load in three distinct shapes, and they are commonly confused with one another. 1. **The in-process stampede.** One instance, many workers, one empty entry. Every worker misses, every worker acquires. 2. **The expiry cliff.** The token expires while requests are in flight, so the acquisition latency lands inside a user-visible — or schedule-visible — path, and every worker hits it at once. 3. **The fleet convergence.** Every instance started at the same deploy, so every instance's token expires in the same second, and the whole fleet acquires together. Single-flight fixes the first, refresh-ahead fixes the second, jitter and staggered starts fix the third. None of them substitutes for another: single-flight bounds each **process** to one acquisition per key, so a hundred-instance fleet still produces a hundred simultaneous acquisitions unless something spreads them. ## Single-flight The rule is one in-flight acquisition per cache key. The first caller to miss starts the acquisition; every later caller that misses the same key **joins** it rather than starting its own, and they all receive the same result — success or failure. Two details are easy to get wrong: - **Failure must be shared too.** If the acquisition fails, the joiners must fail with it and not immediately each start their own. Otherwise the failure path is exactly the stampede you removed. - **The refresh-ahead path must not join-and-wait.** When a valid token is still in the entry, the background refresh is started if absent and the caller returns the existing token immediately. Waiting there reintroduces the cliff. ## Refresh-ahead Refresh-ahead replaces the token before it dies. The refresh point is a fraction of `expires_in` — say three-quarters — subject to a floor: the remaining margin must exceed clock skew plus the slowest acquisition you are prepared to tolerate, or the refresh is late by construction on a slow day. Short lifetimes change the arithmetic: a five-minute token refreshed at 75% leaves a 75-second margin, which a throttled authorization server can eat easily. What refresh-ahead buys is a **budget**. Between the refresh point and real expiry, acquisition can fail repeatedly and nothing breaks, because a still-valid token is worth serving right up to its actual expiry. That budget is the difference between a refresh problem and an outage, and it should be a metric: *remaining life of the token currently being served*, not merely a success counter. ## Jitter and the cold start | pattern | what it prevents | |---|---| | randomised refresh point inside a window (for example 70–85% of `expires_in`) | instances deployed together refreshing together, forever | | jitter on the scheduled wake boundary | every controller's poll — and therefore every miss — landing on the same second | | staggered or rate-limited rollout | a deploy turning into a fleet-wide cold cache in one instant | | pre-warm at startup, before readiness | the acquisition round trip landing inside the first poll's tail latency | Pre-warming has a second benefit: an instance that cannot acquire a token has nothing useful to do, and failing readiness is a more honest signal than accepting work it will fail. ## When the authorization server throttles or is unreachable The correct behaviour depends entirely on what is in the cache. - **A valid token is held.** Keep serving it. Back off — capped exponential with jitter, honouring a `Retry-After` when one is present — and raise the severity of the alarm as the remaining life shrinks towards your acquisition time. This is a degraded state, not an outage. - **The cache is empty or the token has genuinely expired.** Fail closed. Do not call the meter service without a token and do not send the expired one: the callee will refuse it, and you have converted one failure into two plus load on a service that is not at fault. Shed the poll, record the gap, and let the next cycle retry. There is no third option in which the caller decides its own token is still good. The meter service verifies; the caller only holds. Any design that survives a refusal by relaxing something on the caller's side is deciding a question that belongs to the callee. ## What to measure Acquisitions per minute per key (should be roughly one per lifetime per process), coalesced-wait count, refresh-ahead failures, remaining life of the served token, and throttling responses received. The first and the fourth together tell you, in one glance, whether the cache is working and how much budget is left before it stops mattering that it was.
- Refresh-ahead has failed four times but the held token is valid for nine more minutes. Do you alarm, and do you fail polls?Do not fail polls — the token is valid and the meter service will accept it. Do alarm, at a severity that rises as the remaining life approaches your worst acquisition time, because the nine minutes are a budget being spent. The page-worthy moment is not the first failed refresh; it is the point where a further failure becomes an outage.
- Your cache is per process and you run forty instances. Does single-flight help at all?It bounds each process to one acquisition per key, which turns per-request load into per-instance load — a real reduction. What it cannot do is bound the fleet: forty instances still produce forty simultaneous acquisitions on a cold start. That floor is lowered by jitter, staggered rollout and pre-warming, or by accepting a shared cache and the dependency it adds.
- Why jitter the refresh point when every instance already refreshes at the same fraction of expires_in?Because a fixed fraction preserves whatever alignment the fleet started with. Instances deployed together acquire together, expire together and refresh together, in the same second, for as long as they run. Spreading the refresh point across a window breaks the alignment once, permanently, and costs nothing.
- The meter service refuses a call and the caller believes its token is still valid. What is the caller allowed to conclude?Nothing about the token's validity — that is the callee's determination, not the caller's. The caller may re-acquire once, in case the refusal was about a token it should have replaced, and must then treat a second refusal as a signal about what it is requesting rather than a reason to keep retrying.
A maintenance crew starting a shift at the same door. If every technician queues at the key desk for their own set, the desk is the bottleneck at exactly nine o'clock. One person collects a set the crew shares — and collects the next set a few minutes before the current one is handed back at shift change — so nobody ever stands at a locked door waiting for the desk. If the desk is busy, the crew keeps working with the keys it already has, and only worries as the handback time approaches.
saying these in an interview costs you the question
- Thinks single-flight alone bounds acquisitions across the whole fleet
- Refreshes only after a call has already been refused
- Fails every poll the moment one refresh-ahead attempt fails
- Retries a throttled token endpoint in a tight loop until it answers
- Sends the expired token anyway, hoping the callee is lenient
- Uses a fixed refresh fraction fleet-wide and wonders why the burst returns