Your relying party runs on several instances: sign-in starts on one and the callback lands on another — where does the pending-login record live?
answer
- the callback may land elsewhere
- one instance's memory is not shared
- three parking places, three costs
- now a dependency on the login path
- take-and-delete is the atomic part
basics
~20 sPark the record where any instance can reach it: a shared short-TTL store keyed by the value sent in state, or a signed cookie the browser carries back. An instance's own memory is correct only behind routing affinity, and a deploy drops every in-flight login.
solid answer
~50 sThere are three places, and they differ in what they assume. The instance's own in-memory session is free and correct only if routing pins a browser to one instance — and instance replacement during a deploy still drops in-flight sign-ins. A signed, short-lived cookie scoped to the callback path makes the browser carry the record, so any instance can serve the callback with nothing shared behind it; the costs are that single use is not yours to enforce and the cookie must actually survive a cross-site return. A shared store keyed by the `state` value gives you take-and-delete, which is atomic single use plus a TTL the store enforces — at the price of a real availability dependency on the login path. When that store is down, fail closed: new sign-ins fail, existing sessions are untouched.
code
pseudocode · 15 lineson GET /auth/start:
rec = PendingLogin(state=random(32), nonce=random(32),
verifier=random(32),
returnTo=localPathOrDefault(query.next))
store.put("login:" + rec.state, rec, ttl=10.minutes)
redirect(authorizationUrlFor(provider, rec))
on GET /auth/callback:
rec = store.takeAndDelete("login:" + query.state) // atomic; null if absent or expired
if rec == null:
return renderRestartSignIn() // never a server error, never continue
tokens = exchangeCode(query.code, verifier=rec.verifier, redirectUri=config.callbackUrl)
subject = verifyIdToken(tokens.id_token, expectedNonce=rec.nonce)
sessionLayer.issue(subject) // this leaf ends here
redirect(rec.returnTo)go deeper
Know that a web service usually runs as more than one copy, and that anything one copy holds in memory is invisible to the others when the next request happens to land somewhere else.
Name the three parking places — the instance's own session, a signed cookie, a shared short-TTL store — and say what each one assumes about routing and about dependencies.
Quantify it: with N instances and no affinity, roughly 1 − 1/N of sign-ins fail. Then say what happens when the shared store is unavailable, and defend failing closed rather than degrading.
The real decision is which dependency the login path is permitted to have and who owns it when it fails at 3 a.m. If sticky routing is load-bearing, that is a deployment constraint someone must record.
## The two halves of a login need not be the same machine The depot dashboard runs several instances behind a load balancer. One serves the sign-in page and builds the authorization request; some minutes later the browser returns from the fleet operator's identity provider to whichever instance the balancer picks. If the record was written into memory only the first instance has, the second sees a callback carrying a `state` value it has never heard of, and a perfectly good login fails. The failure is statistical, which is what makes it nasty. With N instances and no affinity, roughly 1 − 1/N of sign-ins fail: half at two instances, over eighty per cent at six, which reads like an outage. At one instance — a laptop — it never happens at all, so it ships. ## Three places to park it | Where | What it costs | When it is right | |---|---|---| | In-process memory, or the instance's own session | Correct only with routing affinity; every deploy drops in-flight logins | A single instance, or a session already externalised | | A signed, short-lived cookie held by the browser | No server-side dependency, but single use is not yours and the cookie must survive a cross-site return | Deployments that would rather add nothing to the login path | | A shared short-TTL store keyed by the `state` value | A real availability dependency on the login path | Anything multi-instance that already runs a shared store | **In-process.** Free, no new dependency, and correct only while every request from one browser reaches one instance. Routing affinity buys that, at the price of uneven load and of losing in-flight logins whenever an instance is replaced — on a service that deploys several times a day, a steady trickle of unexplained failed sign-ins nobody can reproduce. If you externalise the whole session instead, you have chosen the third option under a different name. **Signed cookie.** The browser carries the record, so any instance can serve the callback with nothing shared behind it. Three costs come with it. It rides on every request to the path it is scoped to, so scope it to the callback path rather than to the whole site. Sign it at minimum, and encrypt it if you would rather the browser not read the `nonce` and `code_verifier` inside. And it has to arrive on a return navigation from another site: a cookie marked `SameSite=Strict` is not sent on that navigation at all, `SameSite=Lax` is sent on a top-level GET callback, and a provider configured to return via `response_mode=form_post` produces a cross-site POST that needs `SameSite=None; Secure`. Mark it `HttpOnly` and `Secure`, give it the same short expiry as the record, and clear it at the callback. **Shared store.** One small row per in-flight login, keyed by the `state` value, with an expiry the store enforces for you. Single use becomes a conditional delete that tells you whether the row existed, which is as close to atomic consumption as this problem gets. ## Single use is where the three differ most - The shared store gives it to you directly: take-and-delete either returns a record or does not. - In-process memory gives it to you within one instance and not across them. - The cookie does not give it to you at all — the browser holds the bytes and can present them again. Either keep a small server-side consumed-marker keyed by the `state` value, which reintroduces a shared dependency in a smaller form, or accept the token endpoint as the backstop, because the authorization code is itself single-use and a replayed exchange returns `invalid_grant`. That last position is defensible. What is not defensible is believing the cookie enforces single use when it does not. ## When the parking place is unavailable A shared store on the login path is a genuine availability dependency, and the honest answer is to own it rather than design around it: 1. **Fail closed.** No record means no login. The right response is a page asking the supervisor to start again, never a code path that proceeds without the record because the store was slow. 2. **Know the blast radius.** Staff already signed in are unaffected; their sessions do not touch this store. Only new sign-ins fail, which is a smaller incident than it first looks and worth saying out loud during one. 3. **Keep it separate from the large caches.** The dataset is tiny and short-lived, so an eviction storm in an unrelated cache should not be able to empty it. ## Choosing, and writing the choice down Multiple instances with a shared store already in the architecture: use the store. Multiple instances and a strong preference for adding nothing to the login path: use the cookie, and state in the design that the code's own single use is what stops a replay. Sticky routing as the only reason an in-process record works is a decision that must be written down, because the day someone disables affinity to balance load, sign-ins break and nothing in the login code has changed. That is the shape of this leaf's worst incident: a change in routing, a failure in authentication, and no diff between them.
- The record is in a signed cookie. Can the callback still enforce single use?Not from the cookie alone — the browser holds the bytes and can present them again. Either keep a small server-side consumed-marker keyed by the `state` value, which reintroduces a shared dependency in a smaller form, or accept that the authorization code is single-use and a replayed exchange returns `invalid_grant`. Both are defensible; say which one you chose and why.
- How long should the record live, and what sets the floor?Minutes. The floor is a realistic authentication at the provider — a password, a second factor, sometimes a forgotten-password detour — so ten to fifteen minutes is defensible. The ceiling is that every live record is a half-finished login you are holding open; hours buy nothing measurable and leave a large pile that only expiry ever clears.
- What does replacing an instance mid-deploy do to sign-ins already in flight?With in-process records it drops every login that started on the replaced instance, and those users see a callback that finds nothing. With a shared store or a cookie, nothing happens — the record outlives the process. This is the quiet cost of the in-process option, and it scales directly with how often you deploy.
A cloakroom ticket only works if the coats hang behind a counter every attendant can reach. Give each attendant a private rack behind their own desk and a ticket handed to a different desk buys nothing — which is exactly what the returned state value is worth at an instance that never saw the sign-in start.
saying these in an interview costs you the question
- Assumes the browser returns to the instance that started the sign-in
- Treats sticky routing as a design rather than a constraint accepted
- Thinks a cookie-held record enforces single use by itself
- Falls back to local memory when the shared store is down
- Scopes the login cookie to the whole site instead of the callback path
- Marks the login cookie SameSite=Strict and expects a cross-site return