skip to content

Why does a TCP request/response client that writes each request as a small header, then a small body, then reads the reply see every call take about 40 ms longer than the round trip, and how do you fix it?

level: seniorimportance: must knowfreq 38%

answer

  1. two small writes, then a read
  2. sender waits for an ACK
  3. receiver waits for data to piggyback
  4. delayed-ACK timer breaks the deadlock
  5. one write per message, or TCP_NODELAY

basics

~20 s

Nagle's algorithm holds the second small write until the first is acknowledged, while the server delays that ACK because it has nothing to send yet. Each call waits out the delayed-ACK timer; send each request in one write or set TCP_NODELAY.

solid answer

~40 s

The first small write goes out immediately because nothing is outstanding. The second is small and earlier data is unacknowledged, so **Nagle's algorithm** holds it. The server's TCP has received one small segment and uses **delayed ACK**, waiting to piggyback its acknowledgement on response data, but the server application cannot respond until the body arrives. Neither side moves until the delayed-ACK timer fires. RFC 9293 and RFC 5681 only cap that delay (under 0.5 s); 40 ms is a common implementation minimum, and some stacks use up to 200 ms. The best fix is to write each request in **one send** (buffer it, or use a gather write); setting `TCP_NODELAY` also removes the hold. Disabling delayed ACK on the server is a local, non-portable workaround.

go deeper

for a junior

Recall that two TCP features cause this: the sender holds small data waiting for an ACK, and the receiver holds its ACK waiting for data to send.

for a middle

Walk through the six steps of the write-write-read deadlock and state the RFC bound on delayed ACK versus the implementation timer that sets the observed stall.

for a senior

Diagnose from the symptom: a constant stall near the delayed-ACK timer, independent of payload and far above the RTT, and fix it at the client write path rather than tuning the server stack.

for a principal

Push the fix into shared client libraries (whole-message writes plus TCP_NODELAY) so every service inherits it, instead of per-host kernel tuning that only moves the problem.

## Two mechanisms, each reasonable on its own **Nagle's algorithm** (RFC 896; RFC 9293 §3.7.4) runs on the **sender**. While any sent data is unacknowledged (`SND.NXT > SND.UNA`), new small data is buffered until that data is acknowledged or a full-sized segment can be sent. **Delayed acknowledgement** (RFC 1122 §4.2.3.2; RFC 9293 §3.8.6.3; RFC 5681 §4.2) runs on the **receiver**. Rather than acknowledge every segment at once, a TCP receiver may wait, hoping to acknowledge two segments at once or to **piggyback** the ACK on data it is about to send. The specifications bound the wait: - a TCP endpoint **SHOULD** implement delayed ACK (SHLD-18), - the delay **MUST be less than 0.5 seconds** (MUST-40; RFC 5681 phrases it as within 500 ms of the first unacknowledged packet), - an ACK **SHOULD** be generated for at least **every second full-sized segment** (or 2*RMSS bytes of new data). The RFCs set only the ceiling. The actual timer is an **implementation choice**: 40 ms is a common minimum, and some stacks have used values up to 200 ms. ## The stall, step by step A client sends each request as a small header write, then a small body write, then reads the reply: 1. The client writes the header. Nothing is outstanding, so Nagle sends it at once. 2. The client writes the body. The header is unacknowledged and the body is smaller than one MSS, so **Nagle holds the body**. 3. The server's TCP receives the header segment. It is a single segment, not a second full-sized one, so the receiver may **delay its ACK**, hoping to piggyback it on the reply. 4. The server application has only the header and **cannot reply** until the body arrives, so there is no data to piggyback on. 5. Both sides now wait on each other. The deadlock is broken only when the server's **delayed-ACK timer expires** and it sends a bare ACK. 6. The ACK releases the body, the server replies, and the call completes: about one extra round trip plus the delayed-ACK interval, which on a low-latency network is almost all delayed-ACK interval. The signature is a **constant extra latency** close to the peer's delayed-ACK timer, independent of payload size and far larger than the round trip on a local network. The classic trigger is this write-write-read pattern: a small request sent in one write never waits, because Nagle never holds the first small segment on a quiet connection. ## The fixes, in order of preference | Fix | Where | Effect | Caveat | |---|---|---|---| | Send each request in **one write** (application buffer or gather write) | client application | nothing is ever held behind unacknowledged data | requires control of the write path | | Set **`TCP_NODELAY`** | client socket | Nagle never holds small data | many small writes become many small packets | | Both together | client | lowest latency, few packets | a common choice for RPC clients | | Suppress delayed ACK on the server | server stack | ACK leaves at once | non-portable, per-implementation, adds ACK traffic | The first fix addresses the real cause, the application splitting one logical message across writes. `TCP_NODELAY` is the portable switch that RFC 9293 requires every stack to offer (MUST-17), and it is safe once messages are written whole. Suppressing delayed ACK is outside the protocol: any such knob belongs to one operating system, and changing it increases ACK traffic for every other flow on that socket. ## The same shape elsewhere - A single write **slightly larger than one MSS** can stall the same way: the full-sized first segment goes out, the small remainder is held by Nagle, and the receiver, holding just one full-sized segment, may delay its ACK. - RFC 9293 Appendix A.3 records the problem directly: common operating systems enable both Nagle and delayed ACK by default, request/response applications suffer from the combination, a **modification to Nagle** exists and is implemented in some systems without being added to the standard, and **many applications simply disable Nagle**. ## How to confirm it A packet trace shows the pattern plainly: the header segment, then a gap of about the delayed-ACK interval, then a bare ACK from the server, immediately followed by the body segment and the reply. The gap sits on the server's ACK, not on the network.

  • Why does sending the whole TCP request in one write avoid the stall even with Nagle's algorithm left on?
    Nagle only holds new small data while earlier data is unacknowledged. On a quiet request/response connection nothing is outstanding when the request is written, so a single small segment goes out immediately, and the server can reply at once, piggybacking its ACK on the response.
  • What do the RFCs actually require of a TCP receiver's delayed ACK?
    RFC 9293 and RFC 1122 say a receiver SHOULD delay ACKs but the delay MUST be less than 0.5 seconds, and an ACK SHOULD be sent for at least every second full-sized segment. RFC 5681 adds that out-of-order segments SHOULD be acknowledged immediately. Values such as 40 ms or 200 ms are implementation choices, not RFC numbers.
  • Can a single TCP write larger than one MSS hit the same stall?
    Yes. The first full-sized segment is sent, but the small remainder is held by Nagle until that segment is acknowledged. The receiver has only one full-sized segment, so it may delay its ACK, and the tail waits for the delayed-ACK timer.

saying these in an interview costs you the question

  • The 40 ms stall is the network round-trip time and cannot be removed.
  • The RFCs require a 200 ms delayed-ACK timer.
  • Setting the PSH flag on the second write prevents the stall.
  • Increasing the receive window will remove the delay.
  • The server is slow to process requests; the stall is in the application.