A long-lived TCP connection that passes through an AWS Network Load Balancer — a pooled database session or a streaming RPC channel — dies after a few minutes of inactivity, while short request/response traffic to the same endpoint is fine. What is happening, and how do you fix it?
answer
- something in the middle forgets the flow
- only idle connections are affected
- no packet ever reaches the target
- the endpoints were never told
- probe more often than the timer
basics
~20 sThe load balancer expires idle flows: an NLB drops a TCP flow with no traffic for the idle timeout, which defaults to 350 seconds. Keep the connection warm with TCP keepalives or application-level heartbeats set well below that interval.
solid answer
~60 sNLB tracks each connection as a flow with an idle timer. If no data crosses in either direction for the idle timeout — 350 seconds by default for TCP flows, and much shorter for UDP flows — the flow entry is removed. Busy request/response traffic never sits idle long enough to notice; a pooled database connection or an idle stream does, which is why the failure looks selective. Two more things make it confusing: the drop happens in the middle box, so an endpoint that is not writing can hold a socket it believes is open until its next write fails, and each side may fail at a different moment. The fix is to keep the flow alive rather than to detect the death faster: enable TCP keepalives on the client or target with an interval comfortably under the timeout, or use the application's own heartbeat — a connection pool's idle validation, or an RPC framework's keepalive ping. Then set pool idle-eviction below the same threshold so stale entries are never reused.
go deeper
Know that connections through a load balancer can be dropped after a period of inactivity, and that keepalives or heartbeats are what stop that from happening.
Explain the flow-state model and the idle timer, why only idle connections are affected, and how to set a keepalive interval below the timeout rather than tuning the read timeout.
Diagnose from the signature — failure only after quiet, always on first use, nothing in the target's logs — and fix it in both places: keepalives on the wire and idle eviction plus validation in the connection pool.
Generalise the rule for the platform: any long-lived connection crossing infrastructure you do not control needs a heartbeat shorter than the shortest idle timer on the path, and shared client libraries should ship those defaults so no team rediscovers this.
## Why the failure looks selective A Network Load Balancer is not a proxy that terminates connections; it is a distributed flow-forwarding layer. For each connection it holds state — which target this flow goes to — and that state carries an idle timer. When no packets cross in either direction for the idle timeout, the entry is released. As of 2025 the default for TCP flows is 350 seconds, and UDP flows expire far sooner. That is exactly why a load test passes and production breaks. A service under steady traffic never lets a pooled connection sit unused for six minutes. A connection pool sized for peak, holding twenty connections while the night shift uses two, does. So does a streaming channel that is subscribed but quiet, or a worker that polls once every ten minutes. ## Why the client does not notice immediately The flow state is dropped by a middle box. Neither endpoint is told. The kernel on both sides still holds an ESTABLISHED socket, and an application that is not writing has no reason to suspect anything. The symptom therefore appears at the *next use*: a request is written, the packet reaches the load balancer, there is no flow entry for it, and the application sees an error or a stall on the first operation after the quiet period. In a connection pool the effect is nastier still — the pool hands out a connection it believes is healthy, and the very first query on it fails. Teams often mis-diagnose this as a flaky database or a bad target, because the target logs show nothing at all: the packet never reached it. Historically the drop was entirely silent, which is why "it hangs until my socket timeout fires" is the classic description. Modern behaviour is friendlier in some paths, but you should design for the connection being gone without notice. ## The fixes, in order of quality **1. Keep the flow warm.** Any traffic resets the idle timer, so send something before it expires. - **TCP keepalives.** Enable `SO_KEEPALIVE` on the socket and set the idle interval well below the load balancer's timeout — a few minutes rather than the operating-system default, which on Linux is two hours and therefore useless here. Whether you configure this per socket, in a driver setting or in a sysctl depends on your stack, but the requirement is the same: probe more often than the middle box forgets. - **Application heartbeats.** Many protocols already have one: an RPC framework's keepalive ping, a message broker's heartbeat, a connection pool's idle-validation query. Turning on the mechanism the library already ships is usually less code than socket tuning and works through every hop. **2. Make the pool distrust old connections.** Set the pool's maximum idle time below the load balancer timeout so idle connections are evicted and rebuilt rather than handed out dead, and keep a validation-on-borrow test as a backstop. This turns a user-visible error into an invisible reconnect. **3. Reconsider the topology.** If long-lived idle connections are central to the design and the traffic never leaves the VPC, ask whether it needs to traverse a load balancer at all. Sometimes the right answer is a shorter path. A fix that is *not* good: raising the client's socket read timeout. That only lengthens how long you wait before finding out the connection is dead. ## Generalising the lesson Every stateful middle box on a network path — NAT gateways, firewalls, load balancers — has an idle timeout, and they are rarely the same number. When a connection dies after a suspiciously round interval of silence, enumerate the boxes in the path and compare their timers rather than debugging the endpoints. The design rule that follows is worth saying out loud in an interview: any long-lived connection crossing infrastructure you do not control needs a heartbeat shorter than the shortest timer on the path, and a pool policy that assumes a quiet connection may already be gone.
- Why do the target's logs show nothing when this happens?Because the flow state was dropped at the load balancer, the packet that finally fails never reaches the target. The target sees no connection, no error and no close — which is exactly why teams waste time debugging the application. The evidence lives in the client's timing and in load balancer metrics, not in the target's logs.
- Would increasing the client's socket read timeout help?No — it changes only how long you wait before learning the connection is unusable. The connection is already gone. You need the flow kept alive with TCP keepalives or application heartbeats, plus a pool that evicts connections idle longer than the load balancer's timeout so a dead one is never handed out.
- How would you confirm the load balancer is the cause rather than the target?Look for the pattern: failures only after a quiet period of roughly the timeout length, always on the first operation after idleness, never mid-transfer, and no corresponding entry on the target. Then reproduce deliberately — open a connection, wait past the timeout, write — and watch whether the packet ever arrives at the target.
saying these in an interview costs you the question
- Blames the database or target application for flaky connections
- Raises the client socket timeout and calls it fixed
- Assumes a dropped flow sends a close to both endpoints
- Relies on the OS default keepalive interval of two hours
- Thinks steady traffic proves connections never expire