The session records for eleven staffed lanes at a ferry boarding desk live in a tier another team operates — what does a failover of it cost you, and how do you bound that in advance?
answer
- not one request, a population
- who else shares that tier?
- acknowledged write, lost at promotion
- memory pressure evicts live records
- choose the degraded mode beforehand
basics
~20 sA session-tier fault invalidates a whole signed-in population at once rather than degrading one request. Bound the radius in advance: keep irreplaceable records off the regenerable tier, and choose the degraded mode before the incident.
solid answer
~50 sTreat the tier holding server-side session records as a **tier-one availability dependency**, because its failure shape is unlike anything else on that list: a lost record does not slow one caller down, it signs out everyone whose reference pointed into the lost range at the same instant, so all eleven lanes stop together. Three events produce that with no outage — a failover promoting a replica that never received writes the primary had already acknowledged, a capacity-pressure eviction dropping live records to make room for entries the service could have recomputed, and a routine restart of a tier nobody labelled stateful. Bound it beforehand: put the irreplaceable records in a tier of their own, answer a lookup failure differently from a genuine `401 Unauthorized`, and decide whether a fault degrades to read-only service or to a deliberate sign-out — knowing a cached validation delays revocation by exactly its own window.
code
pseudocode · 16 lineson each authenticated request:
ref = cookie["sid"] # an opaque reference, never the state
result = session_tier.lookup(ref) # FOUND | NOT_FOUND | UNREACHABLE
if result is FOUND:
local_cache.put(ref, result.record, expires_in = 30s)
return serve(result.record)
if result is NOT_FOUND:
return challenge() # genuinely unauthenticated -> 401 Unauthorized
# UNREACHABLE: a fault, not an answer -- the two are not the same event
cached = local_cache.get(ref)
if cached is PRESENT:
return serve_read_only(cached) # writes to the session record are refused
return challenge()go deeper
Recall that the cookie carries only a reference and the real record lives on the server. If that record is gone, the request arrives as an unauthenticated one even though the browser still holds its cookie.
Explain why one tier fault signs out many people at once: every reference resolves through the same lookup, so an emptied range invalidates a whole population simultaneously instead of failing caller by caller.
Name the three events that empty the tier with no outage — writes lost at promotion, live records evicted under a memory limit, a restart of a tier nobody called stateful — and state the degraded mode you chose, including the revocation delay a cached validation buys with.
Argue the blast radius as a design property: which state is allowed to share a tier with which, how large a population a single fault may touch, and the cases where signing everyone out is the correct chosen behaviour rather than the accident.
## The dependency this is really about Every authenticated request carries a **reference**, not the state. The server lifts an opaque value out of the `Cookie` header, looks up the record filed under it, and only then knows who is calling. That lookup sits on the path of every authenticated request in the system, which makes the tier holding those records a **tier-one availability dependency** — a component the entire authenticated surface fails with. What makes it unlike the other names on that list is the **shape** of the failure, not its likelihood. A slow database degrades one endpoint. A failed downstream call degrades one feature, for the callers who happen to ask for it, spread over time. A session tier that loses records does something no other dependency does: it invalidates a **population simultaneously**. Eleven staffed lanes do not fail one at a time as each clerk reaches the broken path — they stop inside the same second, and the sailing window closes whether or not the software is working. ## Three events that empty it without an outage The planning mistake is preparing only for "the tier is down". Each of these leaves it up and answering: 1. **A failover that promotes a replica missing recent writes.** Where a write is acknowledged before it has reached a replica, promotion keeps the gap. Sessions opened or refreshed inside that window are absent afterwards, and the tier reports them as **not found** rather than as an error — indistinguishable, to your code, from a session that was properly destroyed. 2. **A capacity-pressure eviction.** A tier under one memory limit holding both a disposable read cache and live session records will, when it fills, drop whatever its policy selects. That policy sees keys, ages and access counts; it does not know which values the service can rebuild from the database and which are irreplaceable. Logins get evicted to make room for something that was recomputable all along. 3. **A restart of a tier nobody labelled stateful.** The team operating it calls it a cache. Caches get restarted, resized and moved during ordinary maintenance. Everyone signed in is signed out, and the change record reads "restarted cache". ## Choosing a tier is choosing a blast radius | Where the records live | Radius of one fault | What fires it | Hot-path cost | |---|---|---|---| | Process memory on each instance | One instance's share of the population | Every ordinary deploy, crash or scale-in | None; nothing leaves the process | | A shared in-memory tier | The whole population | Failover gap, eviction under a memory limit, restart | A network round trip per request | | A relational database | The whole population | Rarer; durability is that tier's normal promise | A read per request, a write whenever the record changes | No row is right in the abstract. Process memory has the **smallest radius and the highest frequency** — it guarantees a sign-out on every deploy, which is at least a rate you can measure and schedule. A shared tier trades a frequent partial loss for a rare total one. The table's point is that the choice is not "which is fastest" but **how many people one fault may sign out, and how often you are willing to have that happen**. ## What you decide before the incident - **Separate the irreplaceable from the regenerable.** Records and cache entries sharing one memory limit means the cache's growth can evict your logins. Two tiers is two blast radii, and the second one is allowed to fail. - **Distinguish "no such record" from "cannot reach the tier".** They deserve different answers. The first is genuinely unauthenticated and gets a challenge (`401 Unauthorized`); the second is a fault. Collapsing them turns a thirty-second hiccup into a mass sign-out. - **Pick the degraded mode.** Serve already-authenticated callers read-only from a short, bounded cached validation, or sign everyone out deliberately. Both are defensible; only a chosen one is defensible at three in the morning. - **Size the population one fault can touch.** Partitioning records so a fault reaches a fraction of the desks costs nothing at steady state and is the only lever that shrinks the radius itself rather than the consequences. - **Write down that signing everyone out is sometimes correct.** For a lane amending a vehicle manifest, acting on authority that may be stale is worse than an interruption — but that has to be a decision with a paper degraded path behind it, not something that emerges because a tier happened to restart. ## What the bounding costs you The cached validation is the main lever and it is not free. A session ended during the fault — by a clerk signing out, or by an operator ending it deliberately — **keeps being honoured until that cached copy expires**. Thirty seconds of grace is thirty seconds of delay on every ending, all the time, not only during an incident. Choose the number against how fast you must be able to stop a session, say it out loud in the design, and keep the fallback read-only so the extra window cannot be used to change anything.
- Does pinning each lane's terminal to one instance remove the problem?No. Pinning is a routing decision about which instance serves a caller; it changes who loses the state, not whether it is lost. Holding records in process memory makes the radius one instance's share rather than the whole population, but it also guarantees a sign-out on every ordinary deploy, crash and scale-in. The balancing algorithm and its own trade-offs are a load-balancing subject.
- How do you find out whether your tier can lose a session write it already acknowledged?Read its durability posture rather than its throughput numbers: whether a write is acknowledged before it is replicated or persisted, how far a replica is allowed to lag, and what promotion does with the gap. Then test it — force a failover in a quiet window, and count the records that were acknowledged beforehand and are missing afterwards.
- When is signing everybody out the right chosen behaviour?When acting on possibly-stale authority costs more than the interruption does — a lane amending a vehicle manifest or a hazardous-cargo declaration, for instance. The point is that it is chosen, written down and paired with a degraded path staff can actually work through, rather than emerging from a tier that was restarted by someone who thought it held nothing.
A cloakroom at a concert hall. The ticket in your pocket is a reference; the coat is the state, and only the tag board says which is which. Lose one coat and one person is unhappy. Wipe the board and every ticket in the building becomes meaningless in the same instant — nobody is served more slowly, everybody is stuck at once, and the queue that forms is the real event.
saying these in an interview costs you the question
- It is only a cache, so an outage just makes requests slower.
- Replication makes losing a session record impossible.
- A memory limit evicts cache entries only, never live session records.
- Everyone simply signs in again, so the cost is nothing.
- Pinning each terminal to one instance makes the session state durable.
- Restarting the session tier is safe because it holds no real data.