After a workload overlay is introduced, small requests succeed but large responses hang between hosts - why?
answer
- size-dependent, not reachability
- silence rather than an error
- the wrapper's bytes count too
- link limit minus the overhead
- probe either side of the threshold
basics
~20 sAlmost certainly a packet-size mismatch. The workload interface still advertises the link's full size, so once the host adds its wrapper the frame exceeds what the link carries and is dropped. Small exchanges never fill a packet, so they survive.
solid answer
~50 sThe symptom is the diagnosis. Cross-host traffic is now wrapped in an outer host-to-host header, but the workload's interface was left at the link's full frame size, so a full-size packet becomes over-size once wrapped and is discarded. Short requests and connection setup never come close to filling a packet, so they work perfectly; a large response ramps up to full-size packets and stops dead. It presents as a hang rather than an error because the drops are silent - the small control messages that would tell the sender to send less are routinely filtered, so the sender just retransmits the same over-size packet until something times out. The fix is to set the workload interface to the link's limit minus the wrapper, or to raise the link's frame size fleet-wide where the fabric supports it. Confirm with one payload just under the computed size and one just over.
code
pseudocode · 10 lineslink_frame_limit = 1500 # what the underlying link carries
host_to_host_wrapper = 50 # outer header added on cross-host paths
workload_interface_mtu = link_frame_limit - host_to_host_wrapper # 1450
# verify between two workloads on DIFFERENT hosts
send(workload_packet_bytes = 1400) # 1400 + 50 = 1450 -> fits, arrives
send(workload_packet_bytes = 1500) # 1500 + 50 = 1550 -> over limit, dropped
# the second send is discarded silently; the sender retransmits the
# same size and the transfer stalls rather than reporting an errorgo deeper
Remember the pattern rather than the mechanism: small requests fine, large transfers hanging, right after a networking change, usually means packets are now too big for the path.
Explain the arithmetic. The link's limit minus the wrapper is what the workload interface may emit, and leaving it at the link's full size is what makes full-size packets disappear.
Demonstrate the reasoning chain: size-dependent failure rules out reachability and naming, silence rules out anything that reports, and same-host success localises it to the wrapped path. Then confirm with probes either side of the threshold.
Treat it as a fleet property. Decide whether headroom comes from shrinking every workload interface or from raising frame sizes across the fabric, and make sure any second layer that adds its own header is accounted for in the same budget.
## Read the symptom literally Two facts are doing all the work here. **Small succeeds, large fails** means the failure is a function of size, not of reachability, permission or name resolution - those fail the first packet, not the thousandth. **It hangs rather than errors** means packets are being discarded silently somewhere neither end can see. Together those two point at one mechanism: something on the path now has less room for a packet than the sender believes it has, and the sender is never being told. That something is the wrapper. An encapsulated workload network puts an outer host-to-host header in front of every cross-host packet. The link still carries the same maximum frame it always did. If the workload's interface was left believing it may emit packets that fill the link exactly, then every full-size packet becomes over-size the moment it is wrapped, and it is dropped before it goes anywhere. ## Why small traffic is completely unaffected 1. Connection setup exchanges a handful of tiny packets. They fit with room to spare. 2. A short request - a telemetry sample, a health check, a status call - fits in one small packet. It fits too. 3. A large response is the first thing that produces packets filling the available space. It is the first thing to disappear. So the service looks healthy by every cheap check anyone runs, and fails only for the one operation nobody automated. ## Why it hangs instead of failing When a packet is too large for a link, the mechanism that is supposed to tell the sender is a small control message sent back to it. Those messages are filtered at host and network boundaries far more often than people expect - sometimes by a hardening rule written years earlier by someone else. With the message suppressed, the sender has no signal at all. It sees no acknowledgement, assumes loss, retransmits **the same over-size packet**, and repeats. The transfer neither completes nor fails; it stalls until an application-level timeout fires, which is why the first report is always "it hangs" and the first (wrong) instinct is to raise the timeout. ## The arithmetic, which is the whole fix | quantity | value in this example | |---|---| | what the link carries | 1500 bytes | | what the host-to-host wrapper costs | about 50 bytes | | what the workload interface may emit | 1450 bytes | Set the workload interface to the link's limit **minus** the wrapper, and a full-size workload packet lands exactly at the link's limit once wrapped. Get the subtraction backwards - or forget it entirely - and you have reproduced the incident. The alternative fix is to give the wrapper its own headroom by raising the frame size on the underlying links, where the fabric supports larger frames end to end. That keeps the workload at a full-size packet and costs no usable bytes, but it is a fabric-wide change: **every** device and host on the path has to agree, and one that does not becomes a new silent black hole for large packets. ## Confirming it rather than guessing Send one payload comfortably below the computed size and one comfortably above, between the same pair of workloads on different hosts. The first arrives and the second vanishes: that is the signature, and it also tells you the threshold, which tells you the wrapper's real cost on this path. ## Variations with the same signature - **Only cross-host pairs fail.** Workloads on the same host never leave the host, never get wrapped, and never hit the limit. Same-host success is confirming evidence, not a contradiction. - **It appears months later.** A second layer added on top - anything that adds its own header to the same packets - eats more of the same headroom. The interface size that was correct for one wrapper is wrong for two. - **It is intermittent.** A mixed fleet, where some paths are wrapped and some are not, fails only for the replicas that happen to land on the wrapped ones. Replica churn then makes it look random.
- Why does the transfer stall rather than fail with a clear error?Because the over-size packets are discarded silently and the small control message that would report it is routinely filtered at a host or network boundary. With no signal, the sender assumes ordinary loss and retransmits the same over-size packet, over and over. Nothing on either end ever learns the reason, so the only thing that eventually ends it is an application-level timeout - which reports a timeout, not a size problem.
- Workload pairs on the same host are fine and only cross-host pairs hang. Does that weaken the diagnosis?It strengthens it. Only cross-host traffic is wrapped; traffic between two workloads on one host never leaves the host and never gets the outer header, so it never exceeds anything. A size problem introduced by encapsulation should show exactly this split. If same-host pairs failed too, the wrapper would not be the explanation and you would look elsewhere.
saying these in an interview costs you the question
- Blames the application's read timeout rather than the path.
- Says the link is simply too slow for large responses.
- Assumes the sender is always told when a packet is too large.
- Expects the failure to surface as a connection error, not a stall.
- Forgets the wrapper's bytes when sizing the workload's interface.