SYN cookies admitted a satellite client's long-lived API connections during a flood; a month on they are still slow. Why?
answer
- throughput is a window over a delay
- both ends must agree it in the handshake
- no memory of the client's shift value
- sixty-four kilobytes on a long path
- the cap lives as long as the connection
basics
~20 sThe stateless reply carried no window-scale option, so those connections agreed no scaling and are capped at 65,535 bytes in flight. On a 300 ms path that ceiling is under 2 Mbps, and it lasts for the life of the connection.
solid answer
~40 sWindow scaling has to be agreed in the handshake, and the server must remember the client's shift value to interpret its later windows. A server that stored nothing cannot, so it omits the option and scaling is off in both directions, capping in-flight data at 65,535 bytes. Throughput is bounded by window divided by round-trip time, so on a 300 ms satellite path that is roughly 1.7 Mbps no matter how fat the link is. Selective acknowledgement and timestamps are lost the same way, and the client's segment size survives only as a coarse value. Nothing errors, so nothing alerts. Turning the mechanism off after the flood does not repair connections already established — options are negotiated once — so a pooled connection carries the cap until something closes it.
code
text · 12 linesclient -> server SYN win=64240
opts: MSS=1460, SACK-permitted, TS, WS=7 (shift 7 => up to ~8 MB)
server -> client SYN-ACK <-- no connection record created; the reply carries its own proof
opts: MSS=1460 (client's value kept only as a coarse index)
... no WS ... no SACK-permitted ... no TS ...
client -> server ACK <-- server recomputes and verifies, then builds the socket
established: window scaling never agreed => at most 65,535 bytes in flight
300 ms satellite path => ~1.7 Mbps ceiling, whatever the link can do
and it holds for the life of this connectiongo deeper
Know that a TCP sender cannot have more data in flight than the receiver's advertised window, and that making that window large depends on an option agreed during the handshake.
Be ready to walk the chain: the reply carried no state, so scaling was never agreed, so in-flight data is capped at 65,535 bytes, so throughput is capped at that over the round-trip time. Then say which clients that hurts.
Show how you would find this after the fact when nothing failed: connection-level evidence for which sessions were admitted degraded, throughput bucketed by client distance, and the knowledge that only reconnecting clears it.
Be able to explain to an account owner why a defence that kept the service up quietly cost a customer segment throughput for a month, and commit to the measurement that stops it being invisible next time.
## Why a window becomes a speed limit A TCP sender may not have more unacknowledged data in flight than the receiver has said it can hold. That gives a hard ceiling on throughput: window divided by round-trip time. The window field in the header is sixteen bits, so without help the largest window is 65,535 bytes. On a 10 ms path that is about 52 Mbps and nobody notices. On a 300 ms satellite path it is about 1.7 Mbps. On a 150 ms trans-oceanic path it is about 3.5 Mbps. The link can be a gigabit; the connection will not go faster. The help is the window-scale option, which multiplies the advertised window by a shift factor agreed once, in the handshake. Both ends must send the option — the client in its SYN and the server in its SYN-ACK — or scaling is off in both directions for the life of the connection. There is no later renegotiation. ## Why a stateless reply cannot agree it To honour scaling, the server must do two things: send the option back, and remember the client's shift value so it can interpret every window the client advertises afterwards. The second is the problem. A server operating in stateless admission has, by construction, written nothing down, and the only channel it has to the future is the small value it puts in the reply and the peer hands back — a few dozen bits, not a state record. There is no room for the client's shift factor, so the server does not agree scaling at all. The same reasoning removes the others. Selective acknowledgement, which lets a receiver tell the sender exactly which segments are missing, is agreed in the handshake and so is dropped; without it, one loss on a long path forces cruder recovery and the sender retransmits more than it needed to, compounding the small-window ceiling. Timestamps go too, taking their round-trip estimation with them. The client's maximum segment size survives only as a coarse approximation drawn from a small fixed set of values, because that is all that fits. Some implementations claw part of this back by smuggling the scale factor and the acknowledgement flag into the timestamp option that the peer will echo, which turns the peer itself into the storage. That works only for clients that offered timestamps, so it is a partial recovery, not a fix, and you cannot assume it. ## Why the damage outlives the flood Everything above is negotiated exactly once, at the handshake. A connection keeps the parameters it was born with until it closes. If your clients open a connection pool and hold it — which is what a well-behaved API client does — then connections opened during the twenty minutes the queue was overflowing carry the 64 KB ceiling for as long as the pool lives, which can be weeks. Disabling the mechanism after the flood changes nothing for them; only reconnecting does, and reconnecting is the client's action, not yours. This is the shape of the harm this control does. It is not an outage. No handshake failed, no request errored, no error budget moved. A minority of customers — precisely the distant ones, who are usually the ones paying for reach — simply got slower, permanently, and told you about it weeks later if they told you at all. ## Proving it is the cause Because there is no failure signal, you need connection-level evidence rather than error rates. Two things settle it. First, look at whether scaling was agreed on the affected established connections and compare against connections opened outside the flood window; the degraded ones will show no scaling and no selective acknowledgement. Second, bucket client throughput by round-trip time. If the ceiling tracks 65,535 bytes divided by RTT — 1.7 Mbps at 300 ms, 3.5 Mbps at 150 ms, unbothered at 10 ms — that is arithmetic no backend regression produces. A backend regression slows everyone; this slows exactly the far ones, and slows them to a number you can predict in advance. ## What to do with the finding The fix in the moment is to get those connections recycled, which usually means asking the client to cycle its pool or forcing a close. The fix for next time is instrumentation: record when the degraded path engaged, tag the connections admitted in those windows, and watch segment-level throughput by distance rather than watching for errors. A defence whose only symptom is that some customers are slower will never be found by anything error-driven.
- Why doesn't disabling the mechanism after the flood restore those clients?Because options are negotiated once, in the handshake, and an established connection keeps the parameters it was born with. A long-lived pooled connection therefore carries the ceiling until something closes it. New connections opened after the queue drains negotiate normally, so the fix is to get the old ones recycled — which is an action on the client side, not on yours.
- Besides window scaling, which loss actually hurts on a lossy long path?Selective acknowledgement. Without it the receiver cannot tell the sender exactly which segments are missing, so one loss triggers cruder recovery and needless retransmission. On a high-delay path where every recovery round trip is expensive, that compounds the small-window ceiling. Timestamps also go, taking round-trip estimation and protection against wrapped sequence numbers with them.
- How would you show this is the cause rather than a backend regression?Bucket client throughput by round-trip time. A backend regression slows everyone roughly equally; this produces a ceiling that tracks 65,535 bytes divided by RTT, so distant clients hit a predictable number and nearby ones are untouched. Confirm by checking whether scaling was agreed on the affected established connections versus ones opened outside the flood window.
The reply is a postcard, not a filing cabinet. Everything the server will ever know about what you asked for has to fit on the postcard, and the scaling factor does not fit.
saying these in an interview costs you the question
- Blames link capacity or a backend regression for the slowness
- Says the options are renegotiated once the flood ends
- Thinks a larger receive buffer alone lifts the ceiling
- Claims the mechanism is completely transparent to clients
- Assumes only connection setup is affected, not throughput