skip to content

One worker in a round_robin gRPC channel stops accepting connections — what happens to its subchannel and to the channel's own state?

level: seniorimportance: must knowfreq 55%

answer

  1. four states, one per connection
  2. the picker filters on one of them
  3. failure is retried, not discarded
  4. the channel is optimistic about itself
  5. reachable is not the same as serving

basics

~10 s

That worker's subchannel moves to TRANSIENT_FAILURE and the picker drops it from rotation until it reconnects and reports READY. The channel itself stays READY while at least one other subchannel is READY.

solid answer

~40 s

Each subchannel carries its own connectivity state: `IDLE`, `CONNECTING`, `READY` or `TRANSIENT_FAILURE`. The failed worker's subchannel goes to `TRANSIENT_FAILURE` and stays there, retrying with backoff, until an attempt succeeds and it reports `READY` again — it is never dropped from the address list for that reason alone. A `round_robin` picker rotates only over subchannels in `READY`, so the failed one is simply skipped. The channel aggregates: while at least one subchannel is `READY` the channel is `READY` and calls keep working. Calls that were already in flight on the dead subchannel fail with `UNAVAILABLE (14)`. Note that `READY` describes the transport only — a process that accepts connections but cannot serve still looks ready, so ejecting it requires consuming a health signal.

code

pseudocode · 10 lines
pseudocode
on subchannel_state_change(sc, state):
    if state == READY:
        rotation.add(sc)
    else:
        rotation.remove(sc)        // IDLE, CONNECTING or TRANSIENT_FAILURE

on pick(call):
    if rotation.is_empty():
        return UNAVAILABLE         // or queue the call until something becomes READY
    return rotation.next()         // one endpoint chosen per call, not per connection

go deeper

for a junior

Know that a channel holding several connections keeps working when one endpoint dies, and that the call which was on it fails rather than waiting.

for a middle

Name the four connectivity states and say which one the picker requires. Explain that a failed subchannel retries rather than being removed.

for a senior

Show that connectivity state is a transport fact, and that ejecting an unwell-but-reachable endpoint needs a health signal the balancer subscribes to. Alarm per endpoint, not per channel.

for a principal

Decide where the health signal is produced and who is accountable for it. A fleet-wide rotation rule is only as good as the weakest service's definition of serving.

## Two layers of state, and only one of them is about health A `round_robin` channel is a small state machine on top of a set of subchannels. A **subchannel** is one managed connection to one resolved address, and it reports one of four connectivity states: - `IDLE` — no connection, and none being attempted; it will connect when a call needs it; - `CONNECTING` — an attempt is in progress; - `READY` — a usable transport exists and calls can be sent; - `TRANSIENT_FAILURE` — the last attempt failed; it will be retried with backoff. The picker consumes exactly this: it rotates over the subchannels in `READY` and skips everything else. That is the whole of the failure handling, and it is why a single dead worker in a pool of eight is invisible to callers. ## What the failing worker's subchannel does When the worker stops accepting connections, its subchannel's next attempt fails and it enters `TRANSIENT_FAILURE`. Three things follow: 1. The picker rebuilds its rotation without that subchannel, so no new call is routed to it. 2. Any call already in flight on it ends with `UNAVAILABLE (14)` — the status that means the endpoint could not be reached or could not serve. Whether that call may safely be sent again is a property of the method, not of the balancer. 3. The subchannel keeps retrying on a backoff schedule and **stays in `TRANSIENT_FAILURE` for the whole of it**. It leaves that state only by succeeding, at which point it reports `READY` and rejoins the rotation. It is not discarded from the address list; only a re-resolution that no longer returns the address removes it. A failed subchannel also prompts the channel to ask its resolver for a fresh answer, which is the mechanism by which a replaced worker's new address is discovered. ## How the channel aggregates The channel reports a single state derived from its subchannels. The rule that matters in an interview is the optimistic one: **if any subchannel is `READY`, the channel is `READY`**. Only when no subchannel is usable does the channel report a failure state of its own. So the natural instinct — one endpoint down, therefore the channel is degraded and should say so — is wrong, and deliberately so: a channel with seven of eight workers up is not degraded from the caller's point of view. The practical consequence is that channel state is a bad alarm signal. By the time the channel reports `TRANSIENT_FAILURE`, every endpoint is gone; the interesting failure — one worker in four dead, a quarter of capacity missing — never changes it. ## The case `READY` cannot see Now the harder half. Consider the fingerprint-matching pool again, and a worker whose process is up, whose listener is accepting, and whose matching engine has deadlocked on a corrupt template cache. Its subchannel connects, so it is `READY`. The picker keeps handing it a quarter of every caller's traffic, and those calls sit until their deadlines expire. Connectivity state is a **transport** fact. It answers "can bytes reach this address", not "can this process do the work". Closing that gap is a separate mechanism: the server reports a health signal per service, the balancer subscribes to it, and an endpoint that reports itself not serving is held out of the rotation exactly as if its subchannel were not ready — while the connection stays up, ready to carry calls again the moment the signal turns round. The balancer *consumes* that signal; defining the health-reporting surface itself is a different subject. ## What this looks like when you operate it - A single dead endpoint should be invisible in call success rates and visible only in per-endpoint metrics. - Alarms should be built on per-subchannel state or per-endpoint success, not on channel state. - A vanished peer that never sends a reset — a host that lost power, a network partition — leaves a subchannel sitting in `READY` with no traffic flowing, because nothing told the transport anything was wrong. Detecting that is the job of keepalive, not of the balancer. - A restarted worker rejoins rotation on its own, and the first sign is usually its traffic share climbing back without anyone acting. The short version a candidate should be able to give: the transport state machine handles endpoints that are *unreachable*, the health signal handles endpoints that are *unwell*, and confusing the two is how a pool keeps routing a quarter of its work into a process that cannot do it.

  • The worker is reachable but reports itself as not serving. What stops the picker sending it calls?
    Only a balancer that consumes a per-endpoint health signal. Connectivity state says the transport works, nothing more, so a process that accepts connections while unable to serve stays in the rotation until the health signal is wired in and an endpoint reporting itself unwell is held out of it.
  • What status does a call already in flight on the failed subchannel receive?
    UNAVAILABLE (14), the status for an endpoint that could not be reached or could not serve. Whether the caller may send that call again is a property of the method's own semantics rather than something the balancer can decide.
  • Why is channel state a poor thing to alarm on?
    Because the channel reports READY while any single subchannel is READY. The failure operators care about — a quarter of the pool gone — never changes it, and by the time the channel reports failure every endpoint is already unreachable. Alarm on per-endpoint state instead.

saying these in an interview costs you the question

  • Thinks a subchannel in TRANSIENT_FAILURE is dropped from the address list for good
  • Assumes the channel reports failure as soon as one endpoint fails
  • Believes a process that accepts connections can necessarily serve calls
  • Thinks the picker keeps sending calls to a subchannel that is not READY
  • Assumes a peer that vanished silently is detected without keepalive