A queue consumer's pooled connections keep reaching a replaced replica of a scoring service — why did the stable name not protect it?
answer
- the pick happens once
- per connection, not per request
- pools outlive the replicas they chose
- removal stops new connections only
- bound connection age to re-pick
basics
~20 sBecause a member is chosen when a connection is established, not when a request is sent. A pool holds its connections for hours, so the choice made at open time is never revisited, and the abstraction only helps callers that open new connections.
solid answer
~40 sThe virtual address in front of a replica set is consulted at connect time: dispatch rules pick one current member and pin that flow to it for its lifetime. A long-lived consumer with a connection pool therefore makes its choices once, at start-up, and then keeps reusing them — so membership changes behind the name never reach it. Removing the replaced replica from the member set stops *new* connections going there, which the consumer never opens. The fixes are on both sides: cap how long a pooled connection may live and validate it on checkout, so the pick is remade periodically; honour a close sent by the server; and on the serving side, leave the member set and then drain, closing connections gracefully before the process exits.
code
pseudocode · 11 linesonCheckout(pool, serviceName):
conn = pool.take()
if conn == null or not conn.isUsable() or now() - conn.openedAt > maxConnectionAge:
if conn != null:
conn.close() // give up the member chosen earlier
address = resolve(serviceName) // the service's own stable address
conn = open(address) // opening re-picks from the CURRENT set
conn.openedAt = now()
return conngo deeper
Take away the core fact: a connection picks its replica when it opens, so a connection that is never reopened never notices new replicas.
Explain why membership updates do not touch established flows, and what a pool does that makes that permanent.
Diagnose it end to end: failures arriving after a rollout, sockets breaking on the next write, and the fix split across bounded connection age on the caller and leave-then-drain on the server.
Decide where this responsibility lives across the estate: bounded connection lifetime as a platform default in shared clients, versus paying per request for a hop that re-picks every time.
## The granularity of the pick The stable name and its virtual address do one job well: they let a caller open a connection without knowing which replicas exist. The pick of a member happens **once per connection**, at the moment it is established, and the platform's connection tracking then keeps that flow going to the same member until it closes. Everything sent on that connection afterwards goes to a replica that was chosen with a membership snapshot that may be hours old. A queue consumer is the worst case for this, because everything about it is long-lived: one process running for days, a pool of connections opened at start-up, a steady stream of work pushed over connections that never need to be reopened. It benefits from the abstraction exactly once and then opts out of it by accident. ## Why removal from the set does not reach it When the replaced replica leaves the member set, what changes is the list dispatch rules choose from for **new** connections. The consumer opens none. So: - the membership update propagates correctly to every host, and changes nothing for this caller; - its established flows keep their existing destination rewriting; - while the old replica is still shutting down it keeps answering, so nothing looks wrong; - when the process finally goes, the sockets break — as resets on the next write, or as a silent half-open connection that only surfaces as a timeout. The consumer's symptom is therefore a burst of failures a little *after* the replacement, not during it, which is why this gets misdiagnosed as an unrelated wobble. ## The worse variant: a pinned member address The same shape gets much worse if the caller resolved the name in the mode where it answers with member addresses, or built its own list once and stored it. Then it is not dialling a stable service address at all; it is dialling addresses that belonged to replicas which no longer exist. Two failure modes follow: 1. Every attempt fails, because nothing is at that address any more. 2. Occasionally an attempt *succeeds* and reaches something else entirely, because the address was released back to the pool and later handed to an unrelated workload. The second is far nastier than an outage: the caller is talking confidently to the wrong service. This is the concrete reason the guidance is to hold the name, resolve it again rather than storing the answer forever, and let the platform hold the members. ## What actually fixes it On the calling side: - **Cap connection lifetime.** Give pooled connections a maximum age and close them past it, so the member choice is remade regularly even when traffic never pauses. - **Evict idle connections** rather than holding a large pool of rarely used ones, which are the most likely to be stale. - **Validate on checkout** and treat a broken connection as a reason to open a new one and retry, keeping in mind that a retry is only safe for work that can be repeated. - **Honour a close from the other end** promptly instead of trying to keep the connection alive through a shutdown. On the serving side: - **Leave the member set first, then drain.** Stop reporting ready so new connections go elsewhere, keep serving what you already accepted for a grace period, then close those connections deliberately so the caller's pool is forced to re-open — a clean close is the signal the pool needs. - **Allow for propagation.** Membership changes reach every host eventually, not instantly, so the drain window has to cover that gap. Designs differ on whether anything re-picks per request. Some platforms and some hops in the path make a choice for every request rather than every connection, which removes the pinning entirely at the cost of doing more work per request; where that behaviour comes from, and what it costs, belongs to the traffic-routing material rather than here. The fact to hold on to for this leaf is the default: **per connection**. ## Two claims to keep straight | claim | true? | |---|---| | Removing a member cuts its established connections | No — it stops new ones being dispatched there | | A stable name re-picks a member for an open connection | No — the pick happened when it was opened | | A pooled caller can pin a replica for the process's lifetime | Yes — that is the defect being described | | Capping connection age makes churn visible to the caller | Yes — a re-opened connection picks from the current set | The interviewer is usually testing one thing with this scenario: whether you know that a service abstraction is a *connection-time* mechanism, and therefore that the caller's connection lifecycle is part of the design rather than an implementation detail of some library.
- Why does the consumer's failure burst arrive after the replacement rather than during it?During the replacement the old replica is still serving what it accepted, so the pool keeps working normally. Only when that process finally exits do the pooled sockets break, as resets on the next write or as timeouts on a half-open connection. The gap between the two makes the failures look unrelated to the rollout.
- The caller stored resolved member addresses instead of the service's address — what new failure appears?Besides dialling addresses where nothing exists, it can eventually reach an unrelated workload that was later given the same address. That is worse than an outage, because the calls succeed at the transport level and fail in confusing ways. Holding the name and letting the platform hold the members is the fix.
- Does adding a retry to the consumer solve this?Only partly, and only if the work is safe to repeat. A retry on a fresh connection does re-pick a current member, so it masks the symptom, but a pool that never ages its connections will keep re-pinning and keep producing bursts at every replacement. Bounded connection lifetime is the actual fix.
saying these in an interview costs you the question
- Assumes the stable name re-picks a member for an already-open connection
- Says removing a member immediately cuts its established connections
- Believes a bigger connection pool spreads traffic across new replicas
- Thinks resolving once at start-up is always safe
- Blames the platform's balancing rather than the caller's connection lifetime