What is session affinity (sticky sessions) in a load balancer, how is it typically implemented, and what breaks when the backend a client is pinned to becomes unavailable?
answer
- cookie pins client to one backend
- requires L7 to read/set cookies
- IP-hash alternative is fragile behind NAT
- pinned backend dies -> state dies with it
- shared store/JWT makes affinity optional, not required
basics
~20 sSession affinity means the load balancer sends all of one client's requests to the same backend server, usually to keep server-stored session data (like a login session) working. It's typically done with a cookie identifying the server. If that server goes down, the client loses that session state unless it's shared elsewhere.
solid answer
~1 minSession affinity pins a given client's requests to the same backend instance across multiple requests, most commonly implemented via a cookie the load balancer sets (an L7 mechanism, since it requires reading and setting HTTP headers) identifying which backend served the first request, or less commonly via a hash of the client's source IP at L4. It exists because many applications keep per-session state - login session, shopping cart, in-memory cache - in the memory of the specific server process that handled the first request, so subsequent requests must land on that same server to see consistent state. The trade-off against pure statelessness is real: sticky sessions undermine even load distribution (a server that happens to accumulate many active sticky sessions stays hot even if new connections would prefer another backend) and, critically, when the pinned backend dies, its in-memory session state dies with it - the client either gets logged out / loses their cart, or the balancer must silently re-pin them to a different backend that has no idea who they are. The standard fix is to move session state out of the backend process entirely into a shared store (Redis, a database, or a client-held JWT), making backends stateless so affinity becomes an optimization rather than a correctness requirement.
go deeper
Should know sticky sessions mean a client keeps talking to the same server, usually via a cookie.
Should explain why it's needed for in-memory session state and name cookie-based vs IP-hash implementations.
Should articulate the load-skew trade-off and the concrete data-loss failure mode when the pinned backend dies.
Should argue for stateless backends with externalized session state as the architectural fix, and reason about how affinity-dependence conflicts with autoscaling and rolling deploys at fleet scale.
## What sticky sessions are, and why they exist Session affinity, commonly called 'sticky sessions,' is a load-balancing policy that overrides the normal per-request distribution algorithm to instead route every request from a given client to the same backend instance for the duration of their session. It exists because many applications, especially older or simpler ones, hold session state in the memory of the process that first handled a client: - a logged-in user's session object, - the contents of a shopping cart, - an in-progress multi-step form, - or a server-side cache keyed by user. If the client's second request lands on a different backend that has never seen them before, that state simply isn't there, and the user appears logged out or their cart is empty. ## How affinity is implemented The most common implementation is **cookie-based** and requires an L7 (application-layer) load balancer, because it depends on reading and writing HTTP headers. On a client's first request, the balancer picks a backend using its normal algorithm, then sets a cookie in the response - either an application-generated session cookie the balancer inspects, or a cookie the balancer itself injects (commonly named something like `AWSALBAPP` or similar depending on vendor) whose value encodes which backend handled the request. On every subsequent request, the balancer reads that cookie and routes directly to the identified backend, bypassing the normal algorithm entirely. A cruder alternative, usable at L4, is **IP-hash affinity**: the balancer hashes the client's source IP and consistently maps that hash to the same backend. This requires no application cooperation but is fragile in practice: - many clients share a source IP behind NAT or a corporate proxy, causing unwanted clustering onto one backend; - mobile clients that switch networks (WiFi to cellular) change IP mid-session and lose affinity anyway. ## The first trade-off: uneven load The first real trade-off is against even load distribution. A well-tuned algorithm like least-connections balances instantaneous load across the fleet, but sticky sessions override that per client: if, by chance, a disproportionate number of long-lived, chatty sessions get pinned to one backend early on, that backend stays hotter than its peers for as long as those sessions live, and the balancer cannot rebalance them away without breaking affinity. This is a slow-building form of load skew that doesn't show up in short benchmarks but becomes visible over the life of a long-running deployment with many concurrent user sessions. ## The second trade-off: the pinned backend goes away The second, more serious problem is what happens when the pinned backend actually goes away - a crash, an out-of-memory kill, a rolling deployment replacing it, or the health check pulling it for being unhealthy. Because the session's state lived only in that backend's process memory, that state is gone the instant the instance is. The load balancer's typical fallback is to re-pin the client to a different backend on their next request (since the original is no longer in the pool), but that new backend has no record of who this client is - the practical result depends on the application: - a poorly designed one shows the user as logged out or their cart empty; - a well-designed one degrades gracefully by re-authenticating from a token or re-fetching cart contents from a database. This is precisely the failure mode that makes stateful, affinity-dependent architectures operationally fragile: every backend restart, deploy, or crash becomes a mini data-loss event for whichever sessions happened to be pinned there, and it directly conflicts with elastic autoscaling, since scaling a stateful fleet down means deliberately evicting sessions. ## The fix: make the backends stateless The standard, and now near-universal, fix is to eliminate the need for affinity at the correctness layer by making backends **stateless**: session state is moved into a shared, external store that every backend instance can reach equally. - a Redis or Memcached cluster for session objects, - a database table for cart contents, - or, avoiding server-side storage entirely, a signed client-held token (a JWT) that carries the session's claims and is verified independently by whichever backend receives it. Once backends are stateless in this sense, session affinity becomes purely a performance optimization rather than a correctness requirement - it can still help by improving cache locality (a backend that has recently handled a user's requests may have warm in-memory caches for their data) without the risk that losing the pinned backend causes any actual data loss, since the shared store or token survives independently. This statelessness is also precisely what makes horizontal autoscaling and rolling deployments safe: any backend can be added, removed, or replaced at any time without special-casing which sessions were 'living' on it. Practically, teams should treat 'do we need sticky sessions for correctness, or only for locality' as a design question to answer explicitly, because defaulting to stateful sticky sessions without a shared backing store is a common source of subtle, hard-to-reproduce production bugs around deploys and autoscaling events.
- Why does IP-hash affinity break down for mobile clients or users behind a corporate NAT?IP-hash pins a client to a backend based on their source IP, but many clients behind a shared corporate or carrier NAT present the same source IP, so they all get clustered onto one backend regardless of individual load. Mobile clients switching between WiFi and cellular change their source IP mid-session, which silently breaks their affinity and can drop their pinned session state.
- How does storing session state in Redis instead of backend process memory change the failure mode when a backend dies?With Redis-backed sessions, the backend process itself holds no unique session state, so if it dies the load balancer can route the client's next request to any other backend, which fetches the same session data from Redis and continues seamlessly. The failure mode shifts from 'lost session' to, at worst, a brief request retried against a different backend - a much smaller-blast-radius problem, though it does make Redis itself a new critical dependency that needs its own availability plan.
- If backends are fully stateless, is there still any reason to keep session affinity enabled?Yes, purely as a performance optimization for cache locality - a backend that recently served a user may have warm in-process caches (e.g. hydrated user objects, precomputed data) that a cold backend would need to rebuild or refetch, so keeping the same client on the same backend can reduce latency and shared-store load, even though correctness no longer depends on it.
Sticky sessions are like always being served by the same barista who remembers your usual order by memory. It's convenient until that barista goes home sick - the next one has no idea what you usually order, because nothing was ever written down anywhere else.
saying these in an interview costs you the question
- Thinks sticky sessions are the only way to preserve login state
- Doesn't know it typically requires an L7 balancer with cookies
- Can't explain what happens to session data when the pinned backend dies
- Believes IP-hash affinity is reliable for all clients
- Doesn't mention moving state to a shared store as the standard fix