A replica stops reporting ready but keeps running — what happens to its membership of the set behind the service's virtual address?
answer
- ready decides who is in
- out of the set, still running
- new connections only, not existing ones
- rejoining is silent and automatic
- empty set still resolves, still refuses
basics
~20 sIts address is removed from the member set, so no new connections are dispatched to it, while the process itself keeps running untouched. When it reports ready again the address is added back silently, with no caller involvement and no event a caller sees.
solid answer
~40 sMembership of the set behind a service's name is driven by readiness: while a replica reports ready its address is in the set, and while it does not, the address is out. Leaving the set is not a restart and not a replacement — the replica keeps running, keeps its address, and keeps serving whatever connections it already accepted; it simply stops being dispatched new ones. Rejoining is just as quiet: once it reports ready again the platform puts the address back and it starts taking a share of traffic, with nothing on the caller side to notice. The practical consequences are that a flapping replica silently rotates back into service, and that capacity behind a name can drop while every replica still counts as running.
code
pseudocode · 14 linesmembers = [] // addresses currently behind the service's virtual address
on readyStateChanged(replica):
if replica.readyCheck == PASS and replica.address not in members:
members.add(replica.address) // silently starts taking new connections
if replica.readyCheck != PASS and replica.address in members:
members.remove(replica.address) // stops NEW connections only
// the replica is not stopped, not restarted, and keeps serving
// whatever connections it already accepted
pickMemberForNewConnection():
if members is empty:
return CONNECTION_REFUSED // the name still resolves; nothing can serve
return members[nextIndex()]go deeper
Hold on to the basic pairing: a replica that says it is not ready stops getting new work, but nobody kills it for saying so.
Explain membership as a live list edited by readiness reports, and be precise that removal stops new connections rather than stopping the process.
Show the operational consequences: alarm on ready members rather than running replicas, expect a propagation window, and recognise a resolving name with refused connections as an empty set.
The trade-off you own is how deep readiness reports should go across an estate: shallow checks keep broken replicas in the set, deep ones take whole services out together when a shared dependency wobbles.
## What readiness controls here A readiness report is a per-replica statement of one thing: *this replica can serve requests right now*. The platform uses that report, in this part of the model, for exactly one purpose — **membership of the set behind the service's name and virtual address**. Address in the set while the replica reports ready; address out of the set while it does not. (The check that produces the report, how often it runs and how many failures it takes, belongs to the lifecycle material; what matters here is the consequence for addressing.) So the member set is not a static list of the replicas you declared. It is a continuously edited list of the ones that say they can work. ## What leaving the set does and does not do This is the pairing candidates most often state backwards, so be explicit about both halves: - **New connections stop being dispatched to it.** Dispatch rules on every host are updated to choose from the remaining members. - **The process keeps running.** It is not stopped, not restarted, not replaced, and nothing has failed from the scheduler's point of view. - **It keeps its address and its identity.** Logs and metrics keep flowing; you can still reach it directly if you know where it is. - **Connections it already accepted generally continue.** Removing a member edits which addresses new connections may be sent to; it does not reach into established flows and cut them. - **Capacity behind the name silently drops** by one replica's worth, while the count of running replicas is unchanged. Contrast that with the other consequence a failing check can have in this model: a failing liveness check **restarts** the container, whereas a failing readiness report **removes it from routing and leaves it running**. Naming which one you mean, and matching it to the right consequence, is most of the answer. ## The rejoin nobody sees Recovery is symmetric and silent. The moment the replica reports ready again, its address goes back into the set and the dispatch rules on every host start including it. No caller does anything. No caller is told. Two things follow that bite in production: - A **flapping** replica — one that passes, fails, passes — rotates back into service between failures and takes a share of traffic each time it is in. From the caller's side this looks like a small, wandering fraction of failing requests with no single culprit. - A replica whose readiness report is **shallower than its real health** (it answers a trivial check but its downstream dependency is broken) is a full member of the set the whole time. The operational lesson is to alarm on the number of *ready* members behind a name, not on the number of running replicas — those two counts diverge exactly when this material matters. ## When the set empties 1. The last member stops reporting ready and its address leaves the set. 2. The service declaration is untouched, so **the name still resolves and the virtual address still exists**. 3. New connections to that address fail — refused or timing out — because there is no member to rewrite them to. 4. The caller therefore sees a connection failure against a name that resolves perfectly well. That asymmetry is a genuinely useful diagnostic. A name that fails to resolve points at the declaration or at the caller's permission to see it; a name that resolves cleanly while every connection fails points at an empty or nearly empty member set — that is, at readiness, not at addressing. ## Propagation is not instantaneous A membership change has to reach the dispatch rules on every host that might send traffic. Between the moment a replica stops reporting ready and the moment the last host has been updated, some connections still arrive at it. This is why a replica that is shutting down should keep serving what it has accepted for a grace period rather than closing its listener the instant it decides to leave, and why platforms sequence removal from the set ahead of termination. Treat the removal as *eventually* effective, not as a switch. | what happened to one replica | in the member set | process running | what acts on it | |---|---|---|---| | reports ready | yes | yes | nothing | | stops reporting ready | no | yes | traffic stops arriving; nothing restarts it | | removed, then reports ready again | yes, silently | yes | traffic resumes with no caller action | Read that table in both directions before an interview: the middle row is the one people invert.
- Every replica behind a name is out of the ready set — what does a caller actually see?Name resolution still succeeds, because the declaration and its virtual address are untouched. What fails is the connection: there is no member to rewrite it to, so the caller gets a refusal or a timeout. A resolving name with failing connections is a strong signal that the problem is membership, not addressing.
- A replica keeps flapping in and out of the ready set — why is that worse than one that stays out?A replica that stays out is simply lost capacity, visible as a ready count below the replica count. A flapping one is added back silently each time it passes and takes a share of traffic while it is in, so a fraction of requests fails intermittently with no single failing replica to point at.
saying these in an interview costs you the question
- Says the platform restarts a replica that stops reporting ready
- Thinks removal from the set kills established connections immediately
- Believes an operator must return a recovered replica to the set
- Assumes the name stops resolving once the member set empties
- Treats running replica count as the number serving traffic