A REST API is described as stateless so that any instance can serve any request. Explain concretely what that means for load balancing and horizontal scaling, and what kinds of server-side data break the property.
answer
- Any instance, any request
- Per-client memory vs shared store
- Sticky routing = sharded instances
- Capacity = instances x throughput
- Test: would another pod answer this?
basics
~20 sEach request carries everything needed to process it, so no instance holds per-client memory. The load balancer can send any request to any instance and you scale by adding instances. Per-client data in one instance's local memory breaks it.
solid answer
~50 sStateless means the server keeps no per-client conversational state between requests: every request carries its own credentials, parameters and position (token, resource id, page cursor). Because no instance is special to a given client, the load balancer can use round-robin, least-connections or random routing, and capacity becomes roughly a multiple of instance count. What breaks it is per-client state in one process's memory: an in-memory HTTP session map, a locally cached auth decision keyed by session id, an in-progress multi-step wizard, an open server-side cursor, a per-user counter in a local map. Any of those forces sticky routing, which turns interchangeable instances into shards. Shared resource state - the database, a shared cache, an idempotency table - does not break it, because every instance reaches the same store. The distinction is per-client memory local to one node versus shared state any node can read.
go deeper
Say that every request carries its own credentials and parameters, so the load balancer can send it to any server, and give one example of what would break it (an in-memory session).
Distinguish per-client interaction state from shared resource state, and name concrete violators such as in-memory sessions, server-held cursors and local counters.
Connect it to operations: disposable instances, free routing, safe autoscaling and rolling deploys; note that the shared store becomes the new bottleneck and single point of failure.
Discuss it as a placement decision for state - client, shared store, or node - and reason about the cost of each in latency, blast radius and cross-region topology.
## What the constraint actually says Statelessness is a rule about **where per-client interaction state lives**, not a claim that the server stores nothing. Each request must contain everything the server needs to understand it, and the server must not depend on remembering anything about *that client* from an earlier request. Durable data still lives on the server - that is resource state, and it is shared. ## Why it produces free horizontal scaling If no instance holds anything private about a client, then all instances are functionally identical. That has three consequences. **Routing is unconstrained.** A layer-4 or layer-7 load balancer can pick any backend by round-robin, least-connections, least-latency or random choice. It does not need to inspect a cookie, hash a session id, or maintain a client-to-backend table. Two consecutive requests from the same browser can land on different instances and both succeed. **Capacity is additive.** Throughput is approximately instances x per-instance throughput, until a shared dependency (database, cache, downstream service) becomes the bottleneck. Adding a pod adds capacity immediately, with no warm-up of client affinity. **Instances are disposable.** Because losing an instance loses nothing a client needs, you can kill, replace, drain or preempt them freely. This is what makes autoscaling, spot instances, rolling deploys and chaos-style restarts safe. ## What breaks it The practical test: *if this request landed on a different instance, would it still work?* Common violators: - **In-memory HTTP sessions.** The classic server-side session map. The session id in the cookie only means something to the node that created it. - **Local auth or permission caches keyed by session id** rather than by a self-contained token. - **Multi-step flows held in memory** - checkout wizards, upload assembly buffers, pending objects that exist only in RAM. - **Server-held pagination cursors** - an open database cursor or result set the client resumes by index. - **Local counters** used for rate limiting, quota or dedupe: correctness silently degrades as instance count changes, and results jump when routing changes. - **Local filesystem writes** treated as durable (uploaded files on the pod's disk). Each of these forces **sticky sessions** (source-IP hash or a load-balancer affinity cookie). Sticky routing is not fatal, but it costs you: uneven load because clients are not uniformly expensive, an inability to shed load from a hot instance, session loss on instance death, and deploys that must either drain slowly or drop users. ## What does *not* break it Shared state reachable identically from every instance is fine: the database, a shared Redis cache, an object store, an idempotency-key table. Read-through caches are also fine as long as they are an optimisation - a cold instance must be able to serve correctly, just slower. Server-side **resource** state (the order, the user record) is exactly what a REST API exists to manage. ## How to say it in an interview Frame it as: statelessness moves per-client context either to the client (self-contained token, cursor in the URL) or into shared storage, so instances stop being special. That is what makes load balancing trivial and instances disposable. Then name the concrete cost: bigger requests, repeated authentication work per call, and no in-process shortcut for expensive per-client setup.
- Does using a database make an API stateful?No. The database holds resource state, which every instance reads identically, so instances remain interchangeable. Statelessness constrains per-client conversational state held between requests, not durable domain data. The test is whether a different instance could serve the next request without loss.
- If you must keep per-client state, what is the cheapest way to keep instances interchangeable?Push it into the request itself (a self-contained token, an opaque cursor in the URL) or into shared storage every instance can reach, such as Redis or a table. Both keep routing free; the client-side option also removes a shared dependency, at the cost of larger requests and harder invalidation.
A stateless service is a bank branch where any teller can serve you because you carry your own passbook; a stateful one is a teller who keeps your file in their own drawer, so you must queue for that person.
saying these in an interview costs you the question
- Claiming a stateless API cannot use a database or any storage
- Believing statelessness means each request opens a new TCP or TLS connection
- Thinking sticky sessions make a stateful design equivalent to a stateless one
- Assuming a shared Redis session store restores statelessness rather than relocating the state
- Calling an API stateless while caching per-user permissions in a local in-process map keyed by session id