Rate-limit counters, an idempotency-key deduplication table, and a response cache are all server-side data that persists between requests. Explain which of these are compatible with a stateless HTTP API and what makes the difference.
answer
- Two tests: self-describing, and reachable by all
- Cache = optimisation, must be correct when cold
- Idempotency-Key: client supplies the correlation
- Local counters: limit scales with instance count
- In-progress marker + request fingerprint
basics
~20 sAll three are fine if they live in shared storage every instance can reach and the request still describes itself. They break statelessness only when kept locally per instance, so correctness depends on which node the request happens to hit.
solid answer
~50 sThe constraint is about **per-client context the request depends on**, not about persistence. Each of these passes or fails on two questions: is it reachable identically from every instance, and does the request still carry its own meaning? - **Response cache**: compatible. It is derived data and an optimisation; a cold instance must still answer correctly, just slower. HTTP explicitly assumes caching. - **Idempotency-key dedupe table**: compatible, and it is exactly the kind of shared server-side state statelessness permits. The client supplies the key in the request (the `Idempotency-Key` header), so the request remains self-describing; the server merely records what it already did. Kept in shared storage, any instance can recognise a replay. - **Rate-limit counters**: compatible when shared; a per-instance counter makes the effective limit depend on routing and instance count, which is the interchangeability failure. The test throughout: would a different instance handle this request correctly?
code
http · 21 linesPOST /payments HTTP/1.1
Host: api.example.com
Idempotency-Key: 6f1a2c9e-4b77-4a1d-9a10-2f0c5f3b8d21
Content-Type: application/json
{"amount":2500,"currency":"EUR"}
HTTP/1.1 201 Created
Location: /payments/p_88213
POST /payments HTTP/1.1
Host: api.example.com
Idempotency-Key: 6f1a2c9e-4b77-4a1d-9a10-2f0c5f3b8d21
Content-Type: application/json
{"amount":2500,"currency":"EUR"}
HTTP/1.1 200 OK
Content-Type: application/json
{"id":"p_88213","amount":2500,"currency":"EUR"}go deeper
Say that shared storage is fine and give the test: could another instance handle this request correctly?
Separate derived caches from bookkeeping such as dedupe records, and explain why local rate-limit counters make the effective limit depend on instance count.
Cover the operational details - atomic reservation for concurrent keys, request fingerprints, key scoping and expiry, and the local-bucket-with-shared-budget compromise.
Frame it as choosing where each class of state lives against accuracy, latency and blast radius, and set explicit tolerances for approximate limiting versus exact deduplication.
## Restating the constraint precisely Statelessness does not say the server remembers nothing. It says the server must not require **per-client interaction context carried over from that client's earlier requests** in order to understand the current one. Everything else - durable resources, derived caches, operational bookkeeping - is allowed. Two tests decide any candidate: 1. **Self-description**: does the request still mean the same thing on its own, without the server recalling the client's previous call? 2. **Interchangeability**: is the data reachable identically by every instance, so a different node handles the request correctly? ## Response caches A cache holds copies of data that can be recomputed. Correctness must never depend on a hit: a freshly started instance with an empty cache must return the same answer, only slower. That makes caches a pure optimisation, and HTTP is designed around them - `Cache-Control`, `ETag` and validators exist precisely so intermediaries can cache. A local in-process cache is fine under this rule, with two caveats. First, invalidation is now per-instance, so stale windows differ between nodes and a user can see values flip as routing changes; that is a consistency problem, not a statelessness violation. Second, if the cache holds *per-client* data that exists nowhere else, it stops being derived data and becomes hidden session state - that is a violation. ## Idempotency-key deduplication A client sends a unique key with a non-safe request, typically in the `Idempotency-Key` header, so a retry after a timeout does not create a second charge or order. The server records the key with the outcome and, on seeing it again, returns the stored result instead of re-executing. This is server-side state that persists across requests - and it is entirely compatible, for a clean reason: **the client supplies the correlation**. The request still fully describes itself; the server does not need to remember anything about the client to interpret it. The dedupe record is bookkeeping about a completed operation, closer to a resource than to a conversation. It must, however, live in storage every instance reads, because retries frequently land elsewhere - that is the whole point of retrying after a failure. A per-instance dedupe map would let a retry through on another node and defeat the mechanism. Practical details: store the key plus a fingerprint of the request so the same key with different content is rejected rather than silently returning the wrong result; scope the key to the caller so tenants cannot collide or probe; record an in-progress marker so two concurrent requests with one key do not both execute; and expire records after a defined window. ## Rate-limit counters Counters are the interesting case, because they are per-client by nature. Statelessness is not violated by counting - the request is still self-describing, since the client's identity comes from its own credential. What matters is placement. A **local** counter makes the enforced limit depend on how many instances exist and how requests are routed: with ten instances and round-robin, a per-instance limit of 100 permits roughly 1000. Scale out and the limit silently loosens; scale in and it tightens. Worse, a client's observed behaviour changes with routing, which is exactly the interchangeability failure. A **shared** counter with atomic increments enforces one global limit regardless of routing and instance count. The usual production compromise is a local token bucket synchronised periodically against a shared budget: it removes a network hop from the hot path while keeping the global limit roughly accurate. That is a deliberate accuracy-for-latency trade, and worth naming as such. ## The general rule Allowed server-side state: resources, derived caches, dedupe and audit records, counters and locks - provided they are shared and the request remains self-describing. Disallowed: anything a request's *meaning* depends on that only one instance knows. When in doubt, ask what a cold instance would do with this exact request.
- Two requests with the same idempotency key arrive concurrently on different instances. How do you prevent both from executing?Insert the key first with an in-progress state using a conditional or unique-constrained write, so exactly one insert wins. The loser sees the existing record and either waits briefly for the outcome or returns 409 telling the client a request with that key is in flight. Without that reservation step, both requests pass the check-then-act window and execute.
- Is an in-process response cache ever a statelessness violation?Only when it holds per-client data that exists nowhere else, so a different instance would answer differently in substance rather than just more slowly. A cache of shared, recomputable data is an optimisation and is fine; the risk it does carry is per-instance invalidation, which lets users observe stale values inconsistently as routing changes.
saying these in an interview costs you the question
- Claiming any server-side storage between requests violates statelessness
- Keeping rate-limit counters in local memory and calling the limit global
- Storing an idempotency key without a request fingerprint, so the same key with different content returns the wrong result
- Checking then writing the idempotency record without an atomic reservation, allowing concurrent duplicates
- Letting correctness depend on a cache hit, so a cold instance answers differently