skip to content

When deploying a WebSocket gateway fleet holding millions of connections, how do you drain nodes without triggering a reconnect storm?

level: seniorimportance: should knowfreq 45%

answer

  1. stop new arrivals first
  2. close at a chosen rate
  3. clients dropped together retry together
  4. randomness in the backoff
  5. the server can say not now

basics

~20 s

Take a node out of rotation first, close its connections gradually with a randomized reconnect hint, roll a few nodes at a time, and pair client backoff with jitter with server-side admission control on new connections.

solid answer

~40 s

Fail the node's readiness check so the load balancer stops sending it new connections, and let clients that leave on their own land elsewhere. Then close the rest at a controlled rate: 500,000 connections over 10 minutes is about 833 per second, each told to reconnect after a randomized delay rather than being cut off cold. Roll nodes in small batches that spare capacity can absorb. Clients reconnect with **exponential backoff plus full jitter**, and each gateway enforces **admission control**, such as a token bucket on new connections that rejects excess upgrade requests with `503` and `Retry-After`, so a surge is turned away cheaply and told when to return instead of piling up half-finished sessions. Without these, 500,000 clients reconnecting within about a second hit the fleet at once.

code

pseudocode · 15 lines
pseudocode
attempt = 0
loop:
    result = connect()
    if result is connected:
        opened_at = now()
        wait until the connection closes
        if now() - opened_at > STABLE_PERIOD:
            attempt = 0
    else if result is HTTP 503 with Retry-After:
        sleep(Retry-After + random(0, JITTER))
        attempt = attempt + 1
        continue
    cap = min(MAX_DELAY, BASE_DELAY * 2 ^ attempt)
    sleep(random(0, cap))
    attempt = attempt + 1

go deeper

for a junior

Recall that restarting a gateway disconnects all of its users, and that those users all try to reconnect straight away.

for a middle

Explain the drain steps in order and why exponential backoff needs random jitter to break up clients that were disconnected together.

for a senior

Show you can pick a drain rate and batch size from the numbers, and that admission control protects the fleet from clients that ignore backoff.

for a principal

Frame deploy speed against disruption as a policy choice, and argue for a thin gateway so the expensive drain procedure runs rarely.

## Why deploying a gateway is different Deploying a stateless service is routine: stop routing requests to an old instance, wait a few seconds for in-flight work, stop it. A **gateway node** in a real-time messaging tier may hold 500,000 connections, some of them open for days. Stopping it disconnects every one of those clients, and each client reacts the same way, by reconnecting to the remaining nodes. ## Anatomy of a reconnect storm A **reconnect storm**, a form of **thundering herd**, happens when many clients that lost their connections at the same moment reconnect at the same moment. Illustrative arithmetic: - One node with 500,000 connections restarted abruptly: if its clients all reconnect within about a second, the fleet sees roughly **500,000 connection attempts per second**. - The same clients spread evenly over 60 seconds: 500,000 / 60, about **8,333 per second**. - A whole fleet of 10 million connections recovering from a regional outage is twenty times the single-node burst. Each attempt costs a TCP connect, a TLS handshake, an HTTP upgrade, authentication and a registry write. TLS handshakes are CPU-heavy, so healthy nodes saturate, handshakes time out, clients retry, and the retries add more load. That feedback loop can keep the tier down long after the original cause is gone. ## A drain sequence 1. **Stop new arrivals.** Fail the node's readiness check so the load balancer sends it no new connections. Existing connections are untouched. 2. **Let natural churn help.** For a short period, clients that disconnect on their own simply land elsewhere when they return. 3. **Close the rest at a controlled rate.** Send each client an application-level "please reconnect" message carrying a randomized delay, then close. Draining 500,000 connections over 10 minutes is 500,000 / 600, about **833 closes per second**. 4. **Force-close at a deadline.** Some connections would otherwise stay open for days, so the drain has a maximum window. 5. **Roll in batches.** Drain only as many nodes at once as spare capacity can absorb. With 20 nodes, batches of 2 and a 10-minute drain, the rollout takes 10 batches x 10 minutes = **100 minutes**. Because this is slow, the gateway should stay **thin**: logic that changes weekly belongs in stateless services behind it, so gateway deploys stay rare. ## Client-side discipline Clients must not retry in lockstep: - use **exponential backoff** with a cap, so repeated failures slow down; - add **full jitter**, sleeping a random time between zero and the current cap, so clients dropped together spread out; - honour a delay the server supplies, whether in a reconnect message or a `Retry-After` header; - reset the backoff only after a connection has stayed up for a while, so a flapping server does not send everyone back to zero. ```pseudocode attempt = 0 loop: result = connect() if result is connected: opened_at = now() wait until the connection closes if now() - opened_at > STABLE_PERIOD: attempt = 0 else if result is HTTP 503 with Retry-After: sleep(Retry-After + random(0, JITTER)) attempt = attempt + 1 continue cap = min(MAX_DELAY, BASE_DELAY * 2 ^ attempt) sleep(random(0, cap)) attempt = attempt + 1 ``` Note that after any close, including a planned drain, the loop still sleeps a random time before reconnecting. ## Server-side admission control Not every client is well-behaved: old app versions and buggy retry loops exist. Each node therefore protects itself: | Control | What it does | |---|---| | Accept-rate limit at the listener | caps new TCP connections per second before any TLS work is done | | Upgrade token bucket | rejects excess upgrade requests with `503 Service Unavailable` and `Retry-After`, skipping authentication and registry writes | | Connection ceiling | refuses new connections once the node reaches its planning target | | Randomized `Retry-After` values | spread rejected clients over a window instead of one instant | A **token bucket** refills at a fixed rate and each new connection spends a token, so a node admits a steady flow and turns away the excess cheaply instead of half-completing thousands of handshakes. ## Choosing the drain speed A faster drain finishes deploys sooner but concentrates reconnects; a slower one is gentle but stretches deploys into hours. Set the rate from how many handshakes per second the rest of the fleet can absorb with headroom, and rely on admission control to catch whatever the estimate missed.

  • How do you recover when the whole fleet restarts at once, for example after a regional outage?
    Bring capacity up before accepting traffic, then admit connections at a rate the fleet can handshake, rejecting the rest with `503` and randomized `Retry-After` values. Client jitter spreads the herd further. Recovery time is roughly connections divided by admitted rate: 10 million at 50,000 per second takes about 200 seconds, which is far better than repeated collapse.
  • Why not keep old nodes running until every connection ends on its own?
    Some connections stay open for days, so a purely natural drain can keep old code running indefinitely and block the next deploy. A drain needs a maximum window: wait briefly for natural churn, close the rest at a controlled rate, and force-close whatever remains at the deadline.

Emptying a stadium one section at a time keeps the exits flowing; opening every gate at once jams the concourse that everyone needs.

saying these in an interview costs you the question

  • Restart all old nodes at once; clients simply retry.
  • Short fixed retry intervals are fine for reconnecting clients.
  • Exponential backoff alone prevents clients from retrying in sync.
  • The load balancer moves open connections to new nodes during a deploy.
  • Admission control is unnecessary as long as clients back off.