skip to content

A per-user session version is compared on every request; where do you hold it, and what revocation lag do you accept?

level: principalimportance: should knowfreq 35%

answer

  1. two operands, one of them shared
  2. free if the user is already loaded
  3. cache lifetime is the lag
  4. write through, expiry is the backstop
  5. decide fail open or closed

basics

~20 s

Read it wherever the request already loads the user, and the kill is immediate. Cache it and the kill lags by the cache lifetime, so pick that lag against the threat behind the button, write through on the revoke path, and publish the number.

solid answer

~50 s

The comparison is only as fresh as the value it compares against, so where that value lives is the whole decision. If the request already loads the user record for other reasons, read it there and the revocation takes effect on the very next request at no extra cost. If it does not, an authoritative read per request is a second hop on the hottest path, and the usual answer is a short-lived cache — which converts the cost into a **lag**: a stale copy keeps a revoked session alive until it expires. Bound that deliberately rather than discover it mid-incident. Write the new value through on the bump path, keep the lifetime as the backstop for copies you cannot reach, and choose the number against the worst case the button is for. What you must not do is compare against a copy frozen in the session record: the bump never reaches it.

code

pseudocode · 18 lines
pseudocode
VERSION_CACHE_TTL = 60s        # this number IS the revocation lag you are promising

function current_version_of(subject):
    hit = shared_cache.get("sessver:" + subject)
    if hit != null:
        return hit                                  # may be up to TTL stale
    v = user_store.read_session_version(subject)    # authoritative
    shared_cache.set("sessver:" + subject, v, VERSION_CACHE_TTL)
    return v

function bump_version(subject):
    v = atomic_increment(user_store.session_version, subject)
    shared_cache.set("sessver:" + subject, v, VERSION_CACHE_TTL)   # write through
    return v

# endpoints that must not run on a stale value read past the cache
function current_version_authoritative(subject):
    return user_store.read_session_version(subject)

go deeper

for a junior

Recall that the check needs the user's current version on every request, and that holding a copy of that number in a cache means a revoked session can keep working until the copy expires.

for a middle

Explain the placements and their costs, and why storing the current value inside the session record defeats the mechanism entirely. Be able to say what write-through on the bump path buys.

for a senior

Show that you would measure the lag rather than describe it, invalidate on the bump instead of waiting for expiry, and treat per-instance caches as a separate and longer lag than the shared one.

for a principal

Own it as a commitment: pick the lag against the threat the button exists for, split endpoints so the operations that matter read authoritatively, decide fail-open versus fail-closed in advance, and put both numbers in the runbook.

## The number has to be on hand for every request Invalidating a whole population of server-side sessions with a monotonic per-subject version turns revocation into a comparison. The comparison has two operands. One is already in the session record the request just loaded. The other — the subject's **current** version — has to come from somewhere, on every authenticated request, and that is where the design cost moved to. This is a lead's decision rather than an implementer's, because the choices differ not in correctness but in what each one commits the organisation to when someone presses the button during an incident. ## Four placements and what each costs | Placement | Extra cost per request | Lag after a bump | Main risk | |---|---|---|---| | Read the user record, which the request already loads | None | None | Only viable if that load genuinely already happens | | A dedicated authoritative read per request | One extra round trip on the hot path | None | Doubles the dependency's request rate | | Short-lived cache, plain expiry | Amortised to near zero | Up to the cache lifetime | The lag is invisible until an incident measures it | | Cache plus write-through on the bump | Amortised to near zero | Propagation time, lifetime as backstop | A dropped write-through silently reverts to full lag | And one non-option that looks like an option: **copying the current version into the session record and comparing against that copy**. It removes the read entirely, and it removes the mechanism with it — the bump changes a value the copies never consult, so nothing is invalidated and you are back to enumerating rows. If a design review produces this, it has misunderstood which operand is authoritative. A per-instance in-process cache deserves its own note. Write-through reaches a shared cache; it does not reach a copy held inside every application instance unless each is notified. With several instances behind a load balancer, the honest statement of the lag is the in-process lifetime, not the propagation time of the shared write. ## Choosing the lag, rather than discovering it The question the lag answers is: *between the moment the button is pressed and the moment the last revoked session stops working, how long may that be?* The answer is not a technical constant, it is a commitment, and it should be chosen against the worst case the button exists for: - **A tablet left in a wagon overnight.** The realistic exposure window is hours; a lag of a minute is irrelevant. - **A supervisor's account known to be in someone else's hands right now.** Every second of lag is a second of an attacker acting as them, and a minute may be indefensible. That split is why many systems carry both: a cached comparison as the ordinary path, and a narrow uncached path for the operations that matter most — the ones that move money, change credentials or change who has access. Those endpoints read authoritatively and pay for it, and every other endpoint runs off the cache. The result is a lag that is long where it is harmless and zero where it is not, which is a far better answer than one number applied everywhere. Whichever you choose, three things follow: 1. **Write it down as a number.** Support will be asked *how long until the other tablets stop working*, and the answer must not be a shrug. 2. **Make the bump path invalidate the cache**, rather than waiting for expiry. Expiry is the backstop, not the mechanism. 3. **Test it the way it will be used.** Bump, then hit a protected endpoint from another session in a loop, and measure when it starts failing. A lag nobody has measured is a lag nobody can defend. ## When the value cannot be read at all The interesting failure is the dependency holding the authoritative version being unreachable while a cached copy has expired. Now the check has no second operand, and there are only two answers: - **Fail closed.** Refuse to authenticate anyone whose version cannot be established. Correct, and it converts one dependency's trouble into everybody being signed out. - **Fail open on the stale value.** Keep honouring the last value seen, and accept that a session revoked during the outage keeps working until the value can be read again. Neither is wrong in the abstract; what is wrong is not having decided. Write the choice down, bound it (fail open on a stale copy for a stated maximum, then close), and make sure the incident runbook says which behaviour the responder should expect, because an engineer discovering it under pressure will assume the opposite.

  • The user store is unreachable and the cached version has expired. Do you let the request through?
    Decide it in advance and bound it. Failing closed signs out the whole population for the duration; failing open on the last known value keeps the yard working and keeps a session revoked during the outage alive. A common middle path is to honour a stale value for a stated maximum and then refuse. The unacceptable answer is discovering the behaviour during the incident.
  • Support asks how long after pressing the button the other tablets stop working. What do you answer?
    A measured number, not an architecture description. With write-through to a shared cache it is effectively the next request; with per-instance copies it is up to the in-process lifetime. Measure it by bumping and polling a protected endpoint from another session until it fails, and keep that figure in the runbook.

saying these in an interview costs you the question

  • Caches the version and calls the revocation immediate anyway
  • Copies the current version into the session record and compares that
  • Quotes a shared-cache lag while instances hold their own copies
  • Has no answer for an unreachable store beyond hoping
  • Applies one lag to every endpoint including credential changes