A shared session tier comes back empty mid-sailing and all eleven boarding lanes re-authenticate within the same minute — how do you plan the login path's capacity for that in advance?
answer
- everyone lost it at the same instant
- the return, not the outage
- credential checks are deliberately expensive
- bounded wait beats timeout-and-retry
- stagger any invalidation you control
basics
~20 sRecovery from a session-tier fault is a synchronised step, not a retry curve: a whole population hits the most expensive path in the system at once. Size the credential path against that number, admit traffic deliberately, and rehearse it.
solid answer
~50 sThe outage is the easy half; the **return** is the capacity event. Everybody lost their reference at the same instant, so everybody presents a credential inside the same minute, and they land on the path that verifies credentials — deliberately costly work per attempt, compute-bound, and followed by a write into the tier that has only just recovered. Plan against the stampede number, not the steady-state login rate: the signed-in population divided by how long people will keep retrying, which is routinely two or three orders of magnitude apart. The levers are **headroom**, **admission control with a bounded wait** so the path completes at its measured rate instead of accepting everything and timing out, **jittered retry hints** so shed traffic does not return as a second wave in phase, and a **read-only degraded mode** that shrinks who must re-authenticate at all. Then rehearse it, and measure how long until lanes are moving vehicles again.
code
pseudocode · 17 lines# admission control in front of the credential-verification handler
MAX_IN_FLIGHT = 40 # measured: where verification stops completing
QUEUE_LIMIT = 400 # about one sailing's staffed population
WAIT_BUDGET = 20s
on login_request:
if in_flight < MAX_IN_FLIGHT:
return verify_and_open_session()
if queue.length >= QUEUE_LIMIT:
return busy(retry_after = jittered(5s, 20s)) # shed, and spread the return
ticket = queue.enqueue(login_request)
if not ticket.admitted_within(WAIT_BUDGET):
return busy(retry_after = jittered(5s, 20s)) # bounded wait, never an open one
return verify_and_open_session()go deeper
Recall that signing in costs the server far more work than serving an already-authenticated request, so a moment when everyone signs in at once is not the same kind of load as ordinary traffic.
Explain why the load arrives all at once rather than spreading out: the loss was shared, so every caller discovers it simultaneously and every retry timer starts in the same second.
Show how you keep the credential path completing under that load — a measured in-flight limit, a bounded wait, shedding with a jittered retry hint — and say what you measured to pick each number.
Own the trade-offs: idle headroom against admission control, a read-only fallback's stale-authority window, who gets admitted first when capacity is finite, and the drill that turns all of it from assertion into a measured time-to-serve.
## Recovery is a step, not a curve Most failures recover along a curve: some callers retry, some give up, load returns over minutes as caches warm and clients back off. A session tier that emptied does not behave that way, because the thing it broke was shared. Every caller lost its reference in the same instant, so every caller discovers it in the same instant, and the load that follows is a **step function** aimed at one path. That path is the worst possible target. Verifying a credential is intentionally expensive work per attempt — that expense is the entire point of how credentials are stored — so the path is compute-bound in a way the session-read path never is. It also fans out: a successful login writes a new record into the tier that has just recovered, and it may reach further dependencies before it completes. The component that failed is therefore the component the recovery load lands on. ## The arithmetic you are planning against Steady-state logins at a boarding desk are the joiners at the start of a shift: a handful a minute, and whatever number your capacity plan was built around. A stampede is a different quantity entirely: **the signed-in population, divided by how long people will tolerate retrying before the operation is considered failed.** Eleven lanes is not a large population, but the ratio is what matters, and for anything larger the two numbers are two or three orders of magnitude apart. Write both down. The gap between them is the capacity decision, and a plan that only knows the first number has not been made. ## Why the spike is worse than its headcount - **Retries multiply it.** A clerk who sees nothing happen presses the button again, and a client that auto-retries adds its own. The request rate exceeds the headcount by whatever those two factors are. - **Autoscaling is the wrong instrument.** The spike lands in seconds; capacity arrives in minutes. New instances also land on the same recovering tier, and the scale-out only helps if the population is still retrying when it completes. - **Everything is in phase.** Fixed retry intervals re-synchronise the herd: the traffic returns as a second wave with the same shape, at a time nobody chose. - **The recovery writes to the thing that just broke.** Every completed login creates a record, so the tier receives its heaviest write burst immediately after it came back. ## Levers, and what each one costs | Lever | What it buys | What it costs | |---|---|---| | Headroom on the credential path | Absorbs the spike with no behaviour change | Capacity that is idle almost always | | Admission control with a bounded wait | The path keeps completing at its measured rate | Some callers are shed; the limit must be measured, not guessed | | Jittered retry hints | The shed traffic returns spread out, not in phase | An individual's worst case gets slower | | Read-only degraded mode | Shrinks the population that must re-authenticate at all | A window in which possibly-stale authority is honoured | | Staggering an invalidation you chose | Turns the step into a ramp | Only available when you own the clock | | Prioritising who is admitted first | The lanes that block the sailing get back first | A business call that must be agreed before the incident | Admission control is the one teams reach for last and should reach for first. An unbounded queue does not absorb a stampede; it accepts work, holds it past the caller's own timeout, and then spends scarce capacity completing requests nobody is waiting for — which is how a sixty-second event becomes a twenty-minute one. ## What the drill must measure Rehearsing this is cheap, and it is the only way the numbers above stop being guesses. 1. **Time until lanes are serving again** — measured from a clerk's side, not from the tier's. "The tier accepts writes" is not recovery; "eleven lanes are moving vehicles" is. 2. **The rate at which the credential path stops completing.** That number is the admission limit; without it, admission control is a guess wearing a configuration value. 3. **The retry amplification you actually see**, because it is always higher than the headcount suggests. 4. **Whether the degraded mode was usable.** A read-only fallback nobody has practised is not a fallback. ## The trade you are accepting Every lever that shrinks the stampede widens some other window. A read-only fallback honours authority that may already have been withdrawn, for as long as the fallback lasts. Longer-lived records mean fewer re-authentications and a longer interval in which an unwanted session persists. Prioritising lanes means somebody is deliberately served last. None of these is free, and a plan that presents any of them as free has not been costed — the lead's job here is to name the number attached to each and choose, before the sailing on which it is discovered.
- Why does autoscaling not rescue this on its own?Because the timescales do not match. The spike lands in seconds and new capacity arrives in minutes, by which time the population has either been served or has given up and started a second wave. Verifying a credential is also compute-bound per attempt, so the extra instances help only in proportion to cores, and they all write their new records into the tier that just recovered.
- What single number should a drill produce?The time from killing the tier to lanes serving vehicles again, measured from the clerk's side. Everything else is an input to it: the rate at which the credential path stops completing, the retry amplification you observed, and whether the read-only fallback was usable by someone who had not practised it. "The tier accepts writes" is not the number anyone cares about.
- If you must deliberately invalidate every stored session for a population, what changes?You own the clock, which is the one case where the step function is a choice. Spread the invalidation over a window long enough that the re-authentication load arrives as a ramp the login path is already sized for, and start it when you are watching. The load does not disappear — the whole population still signs in inside that window — but the peak becomes something you picked.
A building evacuation. Getting everyone out is drilled, signposted and quick. What nobody plans is re-entry: four hundred people back through one badge reader in the same two minutes, having all left at once. The exit was never the capacity problem — the return was, and it is the half that is rehearsed least.
saying these in an interview costs you the question
- Autoscaling absorbs it; the spike only lasts a few minutes.
- Logins are cheap, so the login path scales like any read path.
- Retries sort themselves out without backoff or jitter.
- Steady-state login rate is the number to plan capacity against.
- Forcing a whole population to sign out is free if it happens off-peak.
- Recovery is complete once the session tier accepts writes again.