skip to content

Where does your service provider's assertion replay cache live when several replicas sit behind a load balancer, and how long do entries last?

level: seniorimportance: must knowfreq 45%

answer

  1. single use needs somewhere to write
  2. one map per process is a hole
  3. check and insert as one operation
  4. time-to-live from the validity window
  5. eviction reopens what you closed

basics

~20 s

In a store every replica shares, claimed with one atomic set-if-absent keyed on issuer plus message id, with a time-to-live running to the end of the message's validity window plus the clock skew you allow. A per-process cache leaves a replay hole at the load balancer.

solid answer

~50 s

Single use is a stateful property, so it needs somewhere to keep state that outlives one process. Each replica keeping its own map means a captured message replayed a few times lands on a replica that has not seen the id and is accepted, so the cache has to be shared — a key-value store the whole fleet reads and writes. The claim has to be **one** atomic set-if-absent, not a read followed by a write, or two simultaneous posts of the same message both see nothing and both succeed. The key is the issuing counterparty plus the message id, because ids are only unique within an issuer's namespace. The time-to-live is derived from the end of the message's validity window plus the skew I tolerate — the period during which the message would still pass my other checks — not from my local session lifetime. And if the store is unavailable I reject sign-ins rather than accept messages I cannot prove are fresh.

code

pseudocode · 11 lines
pseudocode
claim_single_use(store, issuer_entity_id, message_id, window_end, skew, now):
    key = "acs:replay:" + issuer_entity_id + ":" + message_id
    ttl = (window_end - now) + skew          # NOT the local session lifetime

    if ttl <= 0:
        return false                          # already outside its own window

    # one round trip: create-if-absent AND set expiry together
    won = store.set_if_absent(key, marker = now, ttl = ttl)

    return won                                # false => already consumed, reject

go deeper

for a junior

Know that a sign-in message is meant to be usable once, and that enforcing once requires the server to remember which ones it has already accepted.

for a middle

Explain why a per-process cache does not hold behind a load balancer, and why the check and the insert have to be one operation rather than two.

for a senior

Derive the time-to-live from the message's validity window plus skew, size the store from peak sign-in rate, treat eviction as a security event, and state the fail-closed decision explicitly.

for a principal

Own the consequence: single-use enforcement makes a shared store a tier-one dependency of authentication. Decide what availability that buys, what it costs, and what you accept during an outage.

## Why this is infrastructure rather than a field A service provider is required to reject a message it has already consumed. The only way to know it has already consumed one is to have written that fact down somewhere. That makes the replay cache a piece of shared, availability-critical infrastructure with a capacity, a latency budget and a failure mode — the same class of thing as a session store — and not, as it first appears, a flag on a row. In the permit system, every contractor firm's sign-in passes through it, several times a minute at shift change, across however many replicas are running. ## A per-process cache is a hole, not an optimisation The tempting implementation is an in-memory map with an expiry, one per process. It is fast, it needs no dependency, and it passes every test you are likely to write, because a test runs one process. Behind a load balancer it does not enforce single use across the fleet. A replayed message is routed by the balancer, not by whoever holds the cache entry, so a replay lands on whichever replica has not seen that id and is accepted there. With a handful of replicas, an attacker who can post the same body repeatedly gets accepted with high probability, and the defect is invisible in logs unless you correlate accepted message ids across instances. It also fails on deploy: a rolling restart empties every map at once. The correct shape: - a **shared key-value store** the whole fleet reads and writes; - entries created by a single **atomic set-if-absent** operation; - a per-entry expiry the store enforces for you; - a key of **issuer plus message id**, because two counterparties may legitimately mint the same id string. ## The claim must be atomic A read-then-write implementation has a lost-update race. Two posts of the same message arriving milliseconds apart at two replicas both read *not present*, both proceed, and both issue a session. This is not theoretical under a deliberate replay, because a replay is exactly a burst of identical requests. The operation you want is the store's native *create only if absent* with a time-to-live set in the same call, returning whether you won. The claim also has to happen **before** the session is issued. Claiming afterwards means the window between validation and issuance is a window in which a concurrent copy also succeeds. ## Sizing the time-to-live from the right clock An entry only needs to exist for as long as the message it names could still pass your other checks. That is the end of the message's own validity window, extended by the clock skew you tolerate. Concretely: ``` ttl = (window_end - now) + skew_allowance ``` Three consequences worth being able to state: 1. **Widening skew widens retention.** Tolerating more clock difference keeps every entry alive longer and enlarges the cache proportionally. It is not a free knob. 2. **The local session lifetime is irrelevant.** A session may last eight hours; the message that started it stopped being acceptable minutes after it was issued. Pinning the entry to the session lifetime wastes memory by orders of magnitude. 3. **A fixed hour is a guess.** Derive it, do not choose it. Capacity follows directly: entries in flight are roughly sign-ins per second multiplied by the window in seconds. That number is small, which is the good news — this store is cheap if it is sized from the window and expensive if it is sized from a habit. ## Eviction, and what it actually costs A store configured to evict under memory pressure rather than strictly by expiry will, when it fills, quietly drop entries that are still inside their window. Every dropped entry is a message that can now be replayed. This is the failure that does not page anyone: sign-ins keep working, throughput is fine, and the property you thought you had is gone. So: configure expiry-driven eviction, size for the real peak rather than the average, and alarm on evictions and on the store's memory headroom. Treat an eviction event as a security event with a blast radius you can name — every message whose window had not closed when it was dropped. ## Failing closed If the shared store is unreachable, you have a choice, and it should be made in advance rather than by a timeout default. Accepting messages you cannot claim means abandoning single use for the duration of the outage. Rejecting them means federated sign-in is down for every counterparty while the store is down. For a system that authorises work on live infrastructure, failing closed is the defensible answer, and it makes the replay cache's availability a first-class dependency of sign-in — which is exactly what it is. Say so in the design, and put it on the same monitoring tier as the session store.

  • Why key the cache on the issuing counterparty as well as the message id?
    Because the id is only unique inside the issuer's own namespace. Two contractor firms' identity providers may mint the same id string, and a shared key space would let one firm's ordinary sign-in block another's, producing a rejection nobody can explain. Prefixing with the issuer entityID keeps the namespaces separate.
  • The shared store goes down at shift change. What does your endpoint do?
    Reject federated sign-ins and say so, rather than accept messages whose freshness cannot be established. That makes the store a tier-one dependency of sign-in, which it already was; the alternative is silently suspending single use across the whole estate for the length of the outage. Decide it in advance and monitor accordingly.
  • Your store is configured to evict the least recently used entries when memory is tight. What is wrong with that here?
    Eviction under pressure drops entries whose validity window has not closed, so those messages become replayable while sign-ins carry on looking healthy. Configure expiry-driven eviction, size from peak sign-in rate multiplied by the window, and alarm on evictions and headroom — an eviction here is a security event, not a capacity note.

A cloakroom ticket is only good once. If each of five attendants keeps a private list of tickets already redeemed, a duplicate ticket works on the fourth attempt. The list has to hang on the wall where all five can see it, and it only needs to hold tickets from tonight.

saying these in an interview costs you the question

  • An in-memory map per instance is enough; replays are rare anyway.
  • Read the cache, then write it — the race window is too small to matter.
  • Pin the replay cache time-to-live to the local session lifetime.
  • Key the cache on the message id alone across all counterparties.
  • If the shared store is down, accept the sign-in and carry on.
  • Claim the id after issuing the session, to keep the sign-in path short.