skip to content

An API gateway holds one shared HTTP connection pool (say, 100 connections) used for calls to a dozen different backend microservices. During an incident, one backend service — call it the 'search' service — starts accepting connections but never responding (its handler threads are all deadlocked). Walk through how this takes down calls to the other eleven, unrelated backend services, and how bulkheading the connection pool per backend would change the outcome.

level: seniorimportance: must knowfreq 60%

answer

  1. accept-but-hang is worse than fail-fast (ties up the resource longer)
  2. shared pool: one bad backend's stuck connections crowd out the rest
  3. partition connections per backend, proportional to traffic
  4. Envoy: separate connection pool per upstream cluster by default
  5. utilization cost in normal times vs outage containment during incident

basics

~20 s

All backends share the same pool of 100 connections. Since search accepts connections but never replies, more and more of those 100 connections get stuck talking to search and never come back to the pool. Eventually there are none left for the other eleven backends, so everything fails, not just search. Giving search its own separate slice of connections would keep the rest working.

solid answer

~50 s

With one shared pool, every outbound call — regardless of target backend — checks out a connection, uses it, and returns it. Because search accepts the TCP connection but never sends a response, connections handed to search calls are held for the full timeout duration (or indefinitely if no timeout exists) instead of being quickly returned. As gateway traffic continues, an increasing fraction of the 100 connections gets consumed by in-flight search calls that never complete, shrinking the pool available for the other eleven backends. Once the pool is fully consumed by stuck search connections, calls to healthy backends can't acquire a connection at all and fail or queue — a fully healthy set of eleven backends becomes unreachable purely due to shared-resource starvation, not any fault of their own. Partitioning the connection pool per backend (say, ~8-10 connections reserved for search specifically) means search's stuck connections can only exhaust search's own slice; the other eleven backends keep their dedicated connections and continue serving traffic normally.

go deeper

for a junior

Should grasp qualitatively that a shared pool lets one bad backend affect calls to unrelated healthy backends, even without deriving the connection-arithmetic.

for a middle

Should walk through roughly why connections get tied up (timeout duration times call rate) and state that per-backend partitioning is the fix, even if approximately.

for a senior

Should reason through the arrival-rate-times-timeout arithmetic that explains how fast a shared pool gets exhausted, articulate why 'hangs' is worse than 'refuses,' and connect the fix to real infrastructure (service mesh per-upstream pooling) as well as application-level implementation.

for a principal

Should reason about this at a fleet level — how to standardize per-upstream isolation as a platform default (e.g., via service mesh policy) rather than relying on every team to implement it correctly in application code, and how to build detection (per-upstream pool saturation alerts) that catches this before it becomes a full incident.

## Why this one is the dangerous shape of failure This scenario is the canonical bulkhead failure mode playing out at the **connection-pool layer** instead of the thread-pool layer, and walking through the mechanics precisely is worth doing because the 'accepts but never responds' detail is what makes it especially dangerous — it's slower and sneakier than an outright connection refusal. ## The shared pool, step by step Start with the shared pool as configured: 100 total HTTP connections, drawn from and returned to a single pool regardless of which of the twelve backends a given call targets. 1. **Under normal conditions**, connections are checked out briefly (for the duration of a request/response round trip) and returned promptly, so 100 connections comfortably serve bursty traffic across twelve backends because most connections are idle in the pool at any instant, available to whichever backend needs one next — this is precisely the utilization benefit of pooling that makes shared pools attractive in the first place. 2. **When search starts accepting TCP connections** (so the connection-establishment step succeeds and looks healthy from the gateway's point of view) but then never writes a response — its handler threads are deadlocked, so the request just sits — every gateway call to search now holds its checked-out connection for as long as the gateway is willing to wait, which is the read/response timeout configured on the HTTP client (or, worse, no timeout at all, in which case the connection is held forever). 3. **If search calls arrive at, say, 20 per second** and the gateway waits up to 30 seconds for a response before giving up, the steady-state number of connections tied up in stuck search calls approaches `20 x 30 = 600` — vastly more than the entire 100-connection pool — meaning within a few seconds of the deadlock beginning, all 100 connections are checked out to stuck search calls and none are available for anyone else, including the eleven backends that have nothing wrong with them. ## What the eleven healthy backends see At that point, a call to any other backend — say, the inventory service — attempts to check out a connection from the pool and finds none available; it either blocks waiting for one to free up (which won't happen until a search call times out and releases its connection, which given the arrival rate is immediately re-claimed by the next queued search call) or fails immediately with a pool-exhausted error, depending on the client's configuration. From the perspective of a caller of the gateway, or from monitoring dashboards, this looks exactly like a full gateway outage — elevated error rates and timeouts across every route — even though eleven of twelve backends never received a bad request and are operating completely normally. This is precisely why 'one service accepts connections but hangs' is more dangerous than 'one service refuses connections outright': - a **refused** connection fails fast and returns the connection to the pool almost immediately; - whereas an **accepted-but-hung** connection ties up the resource for the full timeout window, maximizing the damage per unit of bad traffic. ## What bulkheading the pool changes Bulkheading the connection pool per backend changes the outcome **by construction, not by detection**. If the 100 connections are instead partitioned — for example, roughly proportional to each backend's expected traffic share, giving search perhaps 8 connections reserved specifically for it and the remaining 92 split among the other eleven backends — then when search deadlocks: - only its own 8 connections can ever be consumed by stuck calls; - the moment those 8 are exhausted, further search calls are rejected immediately (fail-fast) rather than queuing for a connection that will never come; - and every other backend's dedicated slice of the pool remains completely untouched, because there is no shared resource left for search's failure to consume. Inventory, and the other ten backends, continue serving traffic at full capacity throughout the incident. The trade-off, as always, is that in normal times those 8 connections reserved for search sit unused whenever search's actual traffic is light, capacity that a shared pool could have lent to a momentarily busier backend — a real but much smaller cost than a full gateway outage. ## Where the infrastructure already does this This exact pattern — a service that hangs, rather than fails fast, causing shared connection-pool exhaustion that takes down unrelated traffic — is the textbook justification **Netflix** gave for building **Hystrix**, and it remains a standard root-cause pattern in postmortems for API gateways and service meshes that pool connections/threads globally rather than per-upstream. Modern service meshes like **Envoy** address it structurally by maintaining separate connection pools per upstream cluster by default, with per-cluster circuit-breaking limits (max connections, max pending requests, max requests) — which is bulkheading built into the mesh's data plane rather than something application code has to implement itself.

  • Why is 'accepts the connection but never responds' a worse failure mode for a shared pool than the backend refusing connections outright?
    A refused connection fails immediately and the caller's connection-pool resource is never actually tied up, so it's returned to the pool almost instantly and does minimal damage. An accepted-but-hung connection is held for the full response timeout (or forever, with no timeout), maximizing how long each bad call occupies a shared resource and therefore how fast the pool gets exhausted by a comparatively low rate of bad traffic.
  • If a per-backend connection pool is bulkheaded but has no timeout configured on the calls themselves, does the bulkhead still contain the damage?
    It contains the damage to the other backends' pools, since search's stuck calls can still only ever consume search's own reserved slice, but within that slice it does not help at all — search's own connections will fill up and stay exhausted indefinitely, meaning every call to search fails, which is arguably correct behavior during a genuine search outage but underscores that a bulkhead without a timeout doesn't recover on its own once the dependency comes back healthy until existing stuck connections are somehow cleared.
  • How would a service mesh like Envoy's default per-upstream-cluster connection pooling change how a team needs to think about this failure mode?
    It removes the need for application code to manually implement per-backend connection partitioning, since the mesh's data plane already isolates connection pools per upstream cluster by default, with configurable per-cluster limits acting as the bulkhead. The team's job shifts to correctly sizing and monitoring those per-cluster limits rather than building the isolation mechanism themselves.

It's like a restaurant with one shared pool of waiters serving twelve different dining rooms: if one dining room's orders get stuck in a kitchen jam and waiters just stand there waiting, soon every waiter is stuck in that one room and the other eleven rooms have no one to take their orders — even though those rooms' kitchens are working fine. Giving each dining room its own dedicated waiters means one jammed kitchen only strands its own room's staff.

saying these in an interview costs you the question

  • Thinks a shared connection pool is fine as long as individual calls have timeouts, without addressing that the timeout still allows temporary but severe cross-backend contention
  • Doesn't distinguish 'accepts but hangs' from 'refuses immediately' as different-severity failure modes
  • Proposes fixing this purely by adding more total connections to the shared pool
  • Can't explain concretely how a fully healthy backend's calls fail due to another backend's problem
  • Unaware that service meshes commonly bulkhead connection pools per upstream by default

context