skip to content

A team enabled cookie-based session affinity at their load balancer to fix a stateful service. A year later, instances added by autoscaling during peak receive almost no traffic, and every rolling deploy produces a burst of errors. Explain both symptoms.

level: seniorimportance: should knowfreq 58%

answer

  1. the algorithm runs once per session
  2. only new sessions can reach new instances
  3. load tracks instance age, not capacity
  4. a replaced instance re-pins a whole cohort
  5. draining protects requests, not clients

basics

~20 s

Affinity pins existing sessions, so a new instance can only receive brand-new sessions — during a spike most traffic is already pinned and scale-out adds unusable capacity. A deploy destroys every session pinned to each replaced instance at once, producing error bursts.

solid answer

~60 s

Both symptoms come from the same property: once affinity is on, the balancing algorithm only chooses a backend for a client's *first* request. Everything after that follows the cookie. So when autoscaling adds instances mid-peak, the new ones are eligible only for sessions that start after they join. Existing users — who are most of your load during a spike — stay on the hot instances, and the fleet grows without the load moving. Rolling deploys hit the other end of the same rule: replacing an instance invalidates every affinity cookie pointing at it simultaneously, so a whole cohort is re-pinned in one step and each of them lands on a backend with no copy of their in-memory state, which surfaces as a burst of errors or logouts sized to your batch. Affinity has effectively bought a routing constraint that fights every elasticity mechanism you own. The mitigations at the proxy are palliative — bounded affinity lifetime, affinity only on the routes that need it, longer drain windows — while the fix is that the state should not be instance-local.

go deeper

for a junior

Know that sticky sessions tie a user to one backend, and that if that backend goes away or a new one is added, the pinning is what decides who gets traffic — not the balancing algorithm.

for a middle

Explain that the algorithm only runs for a session's first request, and derive both symptoms from that: new instances see only new sessions, and a replaced instance re-pins its whole cohort at once.

for a senior

Demonstrate the diagnosis and the staged response: correlate per-backend load with instance age, ship bounded affinity and route-scoped stickiness now, and schedule externalising the state as the actual fix.

for a principal

Own it as a platform constraint: whether affinity is permitted at the edge at all, what it costs you in deploy strategy and autoscaling responsiveness, and how you prevent a temporary workaround from becoming an architectural assumption.

## The one property behind both symptoms Session affinity converts load balancing from a per-request decision into a per-session decision. The algorithm — round-robin, least-connection, whatever it is — runs once, for the request that arrives with no affinity token. From then on the token dictates the backend and the algorithm is bypassed. Everything below follows from that. ## Symptom one: scale-out that does not relieve anything During a peak, most requests belong to sessions that already exist. New instances joining the pool are eligible for exactly one category of traffic: sessions that begin after they join. If your session lifetime is long relative to the spike — a shopping journey, a logged-in dashboard, an operator console — the new capacity absorbs a small trickle while the original instances stay saturated. Autoscaling reads the fleet-average metric, sees it still high, and adds more instances that inherit the same problem. You pay for capacity you cannot route to. The distribution is also cumulative. Instances that have been in the pool longest have had the most opportunities to be assigned new sessions, so load correlates with instance age rather than with instance capacity. Reading per-backend request counts against instance start time usually makes this obvious immediately. Scale-in is the mirror image and is worse than useless: removing an instance terminates every session pinned to it. An autoscaler that scales in on a metric dip is now an availability event generator. ## Symptom two: deploys that emit error bursts A rolling deploy replaces instances in batches. The moment a batch goes away, every affinity token pointing at those instances is dead. The clients behind them are re-pinned — all at once — to surviving or new instances that have no record of their session, so whatever the application does when it cannot find server-side session state (401, redirect to login, 500 from a null lookup) happens to that entire cohort simultaneously. The burst size is your batch size as a fraction of the fleet, and the burst recurs once per batch, which is why the error graph shows a staircase during every deploy. Connection draining does not save you here, and it is worth being precise about why. Draining lets *in-flight requests* on a departing instance complete before it is stopped. It says nothing about the *next* request from that client, which arrives after the instance is gone. Draining protects requests; affinity is about clients. A second-order effect: the re-pinning is a burst of first-requests, which all get balanced by the underlying algorithm at the same instant. With least-connection this is a stampede toward whichever backend momentarily has the fewest in-flight requests, so the cohort can land unevenly on top of everything else. ## Why it was not visible for a year Affinity works perfectly under steady state with a stable fleet. Its costs are paid only during elasticity events: scale-out, scale-in, deploys, instance replacement, spot reclamation, zone failover. A service that deploys weekly and never scales will look fine. The bill arrives when the team adopts autoscaling or increases deployment frequency, which is exactly the point at which nobody remembers that affinity is load-bearing. ## What to do at the proxy layer These are mitigations, and you should present them as such: - **Bound the affinity lifetime.** An affinity token with a short expiry means the population re-balances continuously, so newly added instances start receiving traffic within minutes rather than at session end. The trade is that a re-pin costs the user the same state loss, just earlier and in smaller amounts. - **Scope affinity to the routes that need it.** Usually a minority of endpoints depend on instance-local state. Applying affinity only there lets the rest of the traffic balance normally and shrinks the affected cohort. - **Prefer soft affinity.** "Use the pinned backend when it is available, otherwise balance normally" fails over gracefully; strict affinity that errors when the target is gone converts an instance replacement into user-visible failure. - **Consider hashing on the session identifier instead of a stored mapping**, so that membership changes move only a fraction of clients rather than everyone attached to the departing instance. - **Sequence deploys and scaling around it.** Smaller batches with longer waits spread the same total pain over more time, which for a user-facing error rate genuinely matters. ## What actually fixes it None of the above removes the constraint. The instance-local state is the problem; the balancer is only enforcing the consequence. Once session state is externalised, affinity becomes a performance optimisation you can disable at any time — and that is the property to aim for, because it restores per-request balancing, makes instances interchangeable, and makes deploys and autoscaling uneventful. The senior-level answer names the mitigation you would ship today *and* the remediation you would schedule, and says explicitly which one is which.

  • Would shortening the affinity cookie's lifetime fix the scale-out problem?
    It improves it: with a short lifetime, clients are re-pinned frequently, so instances added mid-peak start receiving traffic within minutes instead of waiting for sessions to end. But each re-pin costs the same state loss the deploy causes, just spread thinner. It buys elasticity by paying the affinity penalty continuously.
  • Why does connection draining not prevent the deploy error burst?
    Draining lets requests already in flight on the departing instance finish before it stops. The failing requests are the ones that arrive after it is gone: the client's next click, routed to a backend that never held its session. Draining protects in-flight requests; affinity is a promise about future ones.
  • How would you confirm that affinity, and not something else, is causing the uneven load?
    Plot per-backend request rate against each instance's start time. Under affinity, load correlates with age — the longest-lived instances hold the most sessions — while under working per-request balancing it does not. Comparing new-session rate per backend against total request rate per backend confirms it directly.
  • Which endpoints would you exclude from affinity first?
    Everything that does not read instance-local session state: static assets, health and metrics endpoints, token-authenticated API calls, and any idempotent read. That leaves the affected cohort as small as the genuinely stateful surface, and it usually reveals that only a handful of routes ever needed affinity.

saying these in an interview costs you the question

  • Blames the autoscaler's metrics rather than the routing constraint
  • Thinks connection draining prevents session loss
  • Assumes new instances get their fair share immediately
  • Treats sticky sessions as a permanent architecture, not a workaround
  • Proposes bigger instances instead of addressing the pinning

context