A nightly batch's calls to one partner endpoint fail at peak while bandwidth stays low — what has the shared gateway run out of?
answer
- concurrency ceiling, not bandwidth
- state is kept per flow
- one destination fixes three fields
- only the source port varies
- entries linger after close
basics
~20 sFree source ports for that one destination pair. An address-translating gateway holds an entry per flow, and with the destination address and port fixed only the source port varies, so simultaneous flows to a single endpoint hit a ceiling long before bandwidth does.
solid answer
~40 sThe gateway is a shared chokepoint with a concurrency ceiling, not a bandwidth one. Each flow it translates occupies one entry identified by the translated source address and port together with the destination address and port. A nightly batch hammering **one** partner endpoint fixes three of those four, so the number of simultaneous flows is bounded by the source ports available on that one translated address — tens of thousands, and fewer in practice because entries linger for an idle period after a flow closes. The signature is distinctive: failures cluster on one destination, throughput is low, retries often succeed, and everything else through the same gateway is fine. Fix it by opening fewer flows — connection reuse first — then by adding addresses or gateways.
code
pseudocode · 12 lines// the gateway keeps one entry per flow
key = (translatedSourceAddress, sourcePort, destinationAddress, destinationPort)
// the nightly batch calls ONE partner endpoint, so three fields are fixed
for each call in batch:
sourcePort = allocateFreePort(translatedSourceAddress, destinationAddress, destinationPort)
if sourcePort is none:
fail("no free source port for this destination pair") // seen by the caller as a timeout
else:
openFlow(key)
// an entry is released only after the flow closes AND an idle period elapsesgo deeper
Recall that an address-translating gateway is stateful and keeps an entry per flow, so it has a limit on simultaneous connections that has nothing to do with bandwidth.
Explain why a single destination is the hard case: the destination address and port are fixed, leaving only the source port to vary within one translated address.
Recognise the signature — partial failures, one destination, low throughput, healthy partner — and order the fixes by cost, with connection reuse first and timeout tuning last.
Treat a single shared egress device as a chokepoint with an owner and a capacity model, and split egress along the same lines as routing intent before it becomes an incident.
## The ceiling is concurrency per destination, not throughput An address-translating gateway is stateful: for every flow leaving through it, it records an entry so the reply can be mapped back to the workload inside. That entry is identified by four things — the translated source address, the source port chosen for the flow, and the destination address and port. Three of those four are **fixed** when a batch calls one partner endpoint over and over, so the only field left to vary is the source port, and the ports available on a single translated address number in the tens of thousands. That is a large number until a batch job with high concurrency and short-lived connections meets it. Worse, a closed flow does not free its entry immediately: gateways hold entries for an idle period so late packets are handled correctly, which means the effective ceiling is lower than the raw port count and depends on how fast the batch churns connections. ## The signature that identifies it - **Failures are partial and timing-shaped.** A fraction of calls fail at peak and the same calls succeed on retry a moment later. - **They cluster on one destination.** Everything else leaving through the same gateway is unaffected, because a different destination pair has its own port space. - **Throughput is low.** The bytes moved are nowhere near any bandwidth figure, which is what rules out the obvious suspects. - **The partner is healthy.** Their own metrics show the requests that arrived succeeding normally; the failures never reached them. - **It arrived with a change in concurrency, not a change in code.** Someone raised the batch's parallelism, or the run window shrank and the same work is now compressed. ## What to do about it, cheapest first 1. **Open fewer flows.** Connection reuse is the highest-leverage fix by a wide margin: a batch that keeps a small pool of connections and sends many requests over each one collapses its flow count by orders of magnitude. Most occurrences of this failure are really a client opening a new connection per request. 2. **Lower the concurrency or spread the run.** If the work does not need to be simultaneous, the ceiling stops being reachable. 3. **Give the gateway more translated addresses.** Each additional address multiplies the port space for the same destination pair. Note the consequence elsewhere: every one of those addresses is one the partner has to accept. 4. **Add gateways and route subnets at their own.** This splits the flow count across devices and removes the single chokepoint, at the cost of more egress addresses again. 5. **Spread across destination addresses** where the partner publishes several, since each distinct destination pair has its own port space. Tuning the gateway's idle timeout downward is sometimes offered as a fix. It does recover entries sooner, and it also tears down flows that were legitimately idle, so treat it as a last resort rather than a knob. ## Why it is a design smell, not just an incident A single translating gateway carrying every subnet's outbound traffic is a **chokepoint by construction**: one device's capacity is the ceiling for the whole estate's egress, and it is a device most teams never think about because it has no code and no deploy. The placement matters as much as the sizing — a gateway lives somewhere specific, and pointing many subnets at one instance concentrates both the flow count and the path. Splitting egress along the same lines you already split routing intent keeps each device's flow count proportional to what it serves. ## What it is not | looks like | actually is | |---|---| | the partner rate limiting you | their limiter returns a response; here nothing arrives | | bandwidth saturation | throughput is low, and bandwidth limits degrade rather than reject cleanly | | a filtering rule denying traffic | a rule denies consistently, not only at peak and only for a fraction of flows | | a workload-side resource leak | the same workloads are fine against every other destination | The habit worth showing: when failures are partial, destination-specific and correlated with concurrency rather than volume, think about the **state** something in the path is keeping per flow, and count what bounds it.
- Why does adding a second translated address to the gateway raise the ceiling?Because the flow identity includes the translated source address, so each additional address brings its own full range of source ports for the same destination pair. The cost is elsewhere: every extra address is one more the partner's allowlist has to carry.
- The team proposes shortening the gateway's idle timeout. What is the trade-off?Entries are recovered sooner, which does raise effective headroom, but flows that were legitimately idle get torn down and reappear to the application as connections dying mid-conversation. Reuse fewer, longer-lived connections before touching the timeout.
- How would you tell this apart from the partner rate limiting you?A limiter answers — the caller gets a response saying it was throttled, and the partner's own metrics show the request arriving. Here nothing arrives at all: the flow never got a source port, so the caller sees a timeout and the partner sees no request.
saying these in an interview costs you the question
- Blames bandwidth when throughput is obviously low
- Assumes the partner must be rate limiting the batch
- Thinks a translating gateway is stateless
- Says raising concurrency always increases completed calls
- Reaches for a shorter idle timeout before connection reuse