Why must per-user request state leave the application instance once a second instance exists, and what does the move cost?
answer
- memory belongs to one process
- the next request lands elsewhere
- pin the user, or move the state
- a round trip per authenticated request
basics
~20 sState held in one instance's memory is readable only by that instance, so the next request may land on a stranger. A shared tier makes all instances equivalent, at the price of a round trip per authenticated request.
solid answer
~50 sPer-user state written into an application instance's own memory is invisible to its peers, and with more than one instance behind a load balancer the next request from that user can land anywhere. There are two ways out. Pin the user to the instance that holds their state: reads stay local and free, but a deploy, a crash or a scale-in takes that state with the process. Or move the state onto a shared volatile tier that every instance reads and writes: the instances become interchangeable, which is exactly what rolling deploys and autoscaling need. What that costs is a network call to the tier in front of *every* authenticated request, plus encoding the state out and decoding it back — on a store that keeps values as opaque bytes, the whole entry travels both ways each time.
go deeper
Recall that each instance has its own memory: state written in one cannot be read by the others, so the next request may arrive at a process that has never seen this user.
Explain both exits — routing the user back to the instance that holds the state, or moving the state onto a shared tier — and say what each does to deploys, scale-in and the request path.
Show that you have priced the move: a call to the tier before an authenticated request can proceed, its latency added to every one of them, and the design choices that keep it at a single round trip.
Weigh what putting a whole signed-in population's state behind one shared component does to the failure domain, and decide what degradation you will accept there instead of an outage.
## Why the problem exists at all An application instance is a process, and a process's memory belongs to it alone. State a request handler puts in a field, a map or a framework's in-memory holder is readable by every thread of that process and by nothing else. Run a single instance and this is invisible: the same process handles every request from every user, so what one request wrote is there for the next. Run two, behind a load balancer that spreads requests, and a user's next request can arrive at a process that has never seen them. Nothing is broken — the second process simply has no record of the first one's work. That is the whole of the problem, and it appears the moment a deployment stops being one process, which for most systems is the moment they want to restart without downtime. ## The two exits There are only two shapes of answer, and they differ in what they make interchangeable. **Pin the user to the process.** Routing is configured so every request from one user goes back to the instance that first served them. The state stays in local memory: reads cost a memory access, nothing is serialised, and no extra component exists. **Move the state out.** Every instance reads and writes that state in a shared volatile tier all of them can reach. The instances become interchangeable, and the state becomes a thing with its own address, its own lifetime and its own capacity. | | Pinned to the instance | On a shared tier | |---|---|---| | Who can serve the next request | only one instance | any instance | | Cost of a read | a local memory access | a network round trip | | A deploy, crash or scale-in | takes that instance's state with it | leaves the state where it is | | Load balancing | constrained by who holds what | free to send a request anywhere | | New failure mode | losing one process affects its users | the tier sits on every authenticated request | Pinning's bill arrives on the most ordinary days — a rolling deploy, an instance replaced by the scheduler, scale-in after a quiet hour. Each of those takes a slice of the signed-in population's state with it. Affinity also fights load balancing: a heavy user stays on one process however busy that process gets. ## What the move actually costs Externalising does not make the application stateless. It moves the state onto a component that is honest about being stateful, and the bill is charged per request: - **A network call before any work.** Deciding whether a request may proceed now needs a call to the tier, and it sits in front of the work rather than beside it. The tier's latency, including its tail, is added to every authenticated request. - **Serialisation in both directions.** The state must be encoded to be stored and decoded to be used. On a store that treats the value as opaque bytes and hands them back untouched, that is the whole entry every time; a store that understands the value's structure may return only the part you asked for. A large share of this class does the former, so do not design assuming the latter. - **A second cost when the entry's deadline moves.** A fixed lifetime (counted from the write) leaves the entry alone while the user is active. An access-extended lifetime (pushed forward by each use) turns every request into a write as well as a read. - **A shared failure domain.** Every signed-in user now depends on one component, which is a capacity and availability conversation to have in advance with whoever operates it. ## Keeping it to one round trip 1. **Hold one user's state under one key**, so a request fetches it in a single operation instead of assembling it from several. 2. **Read it once per request** and pass it down, rather than re-reading in every layer that happens to want it. 3. **Keep the entry small**, because it crosses the network on every request — and on an opaque-value store it crosses whole. ## What an interviewer is listening for The junior answer is "the servers do not share memory", which is correct as far as it goes. The stronger answer names both exits and says why most systems chose the second: not because it is cheaper per request — it is not — but because it makes instances disposable, which is what deploys, autoscaling and failure recovery all depend on. The strongest answer then prices the choice out loud: a round trip in front of every authenticated request, a serialised payload each way, and a component whose bad minute belongs to everyone who is signed in.
- If the state moved to a shared tier, in what sense is the application now stateless?Only in the sense that the instances are interchangeable, which is what deploys and autoscaling need. The state did not disappear — it moved to a component with its own capacity, latency and failure behaviour, and each instance still holds it for the duration of a request. "Stateless" describes where the state is not kept between requests, not that the system has none.
- The state is small — why not carry it in the request itself and skip the tier entirely?You can, and then there is nothing to move: any instance can read it because the client brings it. The trade is that you have swapped a round trip for a copy the server does not hold and cannot remove. Ending that user's access stops being a deletion and becomes something every request path has to check.
- A team configures user affinity in the load balancer and calls the problem solved — when does that bill arrive?On ordinary operations: rolling deploys, a crashed or replaced instance, scale-in. Each takes the state of the users pinned to that process. Affinity also unbalances load, because a long-lived heavy user stays on one instance no matter how busy it is, and it makes the request path's behaviour depend on routing configuration that is easy to lose.
saying these in an interview costs you the question
- Assumes instances share one memory space because they share a deployment
- Claims a load balancer keeps a user on one instance by default
- Treats the shared tier as free and forgets the round trip it adds
- Says affinity is equivalent, ignoring that a deploy takes the state with the process
- Calls this state a copy of something else that can simply be recomputed