A web application runs on two identical instances behind a load balancer that spreads requests round-robin. Users report being logged out at random. What is causing that, and what can the load-balancing layer do about it?
answer
- two instances, two separate memories
- the next request lands somewhere else
- pin the client to one backend
- affinity is a workaround, not storage
basics
~20 sEach instance keeps session state in its own memory, so round-robin sends the next request to an instance that never saw the login and treats the user as anonymous. At the proxy layer the fix is session affinity: pin a client to one backend.
solid answer
~50 sThe session lives in one instance's local memory. Round-robin does not know or care about that, so roughly every other request lands on the instance that has no entry for that session id, and the app falls back to "not logged in". It looks random because whether a given click works depends on which backend the balancer picked. Two layers can fix it. The load-balancing layer offers **session affinity**: either the balancer issues its own cookie identifying the chosen backend and honours it on later requests, or it hashes a stable client attribute such as the source IP to pick the same backend each time. That works, but it is a workaround — it pins traffic and creates its own problems with scaling and deploys. The durable fix is application-side: move session state off the instance into a shared store or a signed token so any backend can serve any request.
go deeper
Be ready to say plainly that the session lives in one instance's memory and the balancer sent the next request elsewhere. Name sticky sessions as the proxy-side option and a shared session store as the real fix.
Explain the two affinity mechanisms — a balancer-issued cookie versus hashing the client IP — and the granularity difference between them, including why NAT makes IP hashing coarse.
Show how you would diagnose it in production: expose the serving backend in a header or log line, correlate the login and the failure, and decide whether to ship affinity now while scheduling state externalisation.
Frame it as a platform rule rather than one service's bug: decide whether affinity is allowed at all in your edge tier, what it costs the fleet in scheduling freedom, and how you stop it becoming a permanent hidden dependency.
## The symptom and the mechanism Nothing is broken. The load balancer is doing exactly what round-robin means: request 1 to instance A, request 2 to instance B, request 3 to instance A, and so on. The application, however, stored the user's session — the server-side record created at login and keyed by the session id in the user's cookie — in the memory of whichever instance handled the login. Instance B has no such record. When B receives the cookie it looks up an id it has never seen, finds nothing, and renders the logged-out state. The intermittency is the giveaway. A configuration error is consistent; a state-locality problem is a coin flip, and its failure rate tracks the number of instances: with two instances about half of requests miss, with five about 80% miss. A page that issues several requests (HTML plus a few XHR calls) may render half-authenticated, which is why bug reports say "random". ## Confirming it rather than guessing Make the backend visible. Most proxies can add the chosen backend's name or address to a response header or to the access log line; correlate a failing request with the instance that served it and with the instance that served the login. If every failure is on a different instance than the login, the diagnosis is complete. A second confirmation: scale the deployment down to one instance and see whether the problem disappears entirely. ## What the load-balancing layer can offer **Session affinity** (sticky sessions) makes the balancer route a client's requests to the same backend for as long as that backend is available. It comes in two broad flavours. *Balancer-issued cookie.* The proxy sets its own cookie on the first response, encoding or referencing the backend it chose, and reads that cookie on later requests. This is precise — it identifies one browser rather than one network path — and it works through NAT and mobile carriers. It requires HTTP-level (L7) proxying, because the balancer must read and write headers, and it adds a cookie of the proxy's own to your responses. *Hash of a client attribute.* The proxy hashes something stable about the request — most commonly the source IP address — and maps the hash to a backend. It needs no cookie and works even for non-HTTP traffic, but the granularity is wrong: thousands of users behind one corporate NAT or mobile gateway hash to a single backend, and a user whose IP changes (Wi-Fi to cellular) is silently re-pinned and logged out anyway. Either way, affinity is a routing hint, not durable storage. If the pinned instance is removed, restarted, or fails a health check, the client is re-pinned to a fresh instance and loses the session exactly as before. Affinity narrows the window; it does not close it. ## Why the durable answer is not at the proxy The proxy is compensating for the application holding state where only one machine can see it. Move the state and the balancer's choice stops mattering: keep sessions in a store both instances read (a shared cache or database), or carry the session in a signed, tamper-evident token the client presents on each request so no server-side lookup is needed at all. Once any instance can serve any request, you can balance on load rather than on identity, replace instances freely, and scale out without re-teaching the balancer anything. A reasonable interim posture, and one interviewers like to hear, is: turn affinity on now to stop the bleeding, log how many requests actually depend on it, and treat the state externalisation as the real remediation with a date on it. The failure mode to avoid is affinity becoming permanent and invisible, so that a year later nobody remembers the service cannot survive being re-balanced. ## The adjacent traps Two things commonly confuse people here. First, a persistent (keep-alive) connection already keeps a series of requests on one backend, so a client that reuses one connection may appear to work while a client that opens fresh connections fails — that is a symptom of the same cause, not a fix. Second, session replication between instances (each instance broadcasting session changes to the others) is technically possible and occasionally used, but it trades a routing problem for a distributed-consistency problem and gets expensive as the fleet grows.
- If we turn on IP-hash affinity, why might a large customer still report the problem?Because IP-hash pins a network address, not a user. A customer behind a corporate NAT or a mobile carrier gateway can change egress IP mid-session and be re-pinned to a different backend, losing the session — while at the same time all of that customer's users hash to one backend, concentrating their load.
- The session survives normal traffic but is lost every deployment. Why does affinity not help there?Affinity only holds while the pinned backend exists. A deployment replaces instances, so every session pinned to a replaced instance is re-pinned to a new one that has no copy of the state. Affinity narrows the window between failures; it does not make the state durable.
- How would you prove which instance served a failing request?Have the proxy stamp the chosen backend into a response header or the access log, then correlate the failing request with the login request. Different backends on the two lines confirms state locality. Scaling to a single instance and seeing the failures vanish is a fast second confirmation.
saying these in an interview costs you the question
- Says the load balancer is misconfigured or dropping cookies
- Assumes the session cookie itself carries the session data
- Thinks sticky sessions make the application stateless
- Believes HTTPS or keep-alive guarantees the same backend
- Proposes replicating sessions across all instances without cost analysis