Why does terminating the client connection at a nearby edge location cut time to first byte for an uncacheable response?
answer
- the cost is crossings times distance
- setup happens before any useful byte
- the short hop absorbs the repeated exchanges
- one full crossing remains as a floor
- warm long connection skips the cautious start
basics
~20 sConnection setup costs several round trips before any request bytes move. Terminating nearby pays those against the short hop, while the forwarded request crosses the long distance once over a connection the edge already holds warm.
solid answer
~40 sBefore a client sends its first request byte, the transport connection and the security handshake each cost round trips - how many depends on the protocol version and whether the session resumes. Those round trips are paid against whatever the client connected to. Connect straight to a distant region and every one of them costs the full distance. Terminate at a nearby location instead and they cost a few milliseconds each, after which the location forwards the request over a connection to the region it already has established and warmed. The response still crosses the long distance once, so the region's distance and service time remain a floor. What disappears is the multiplier: the repeated crossings that happened before any useful work started.
code
pseudocode · 17 lines# illustrative figures, not any provider's published numbers
nearEdgeRtt = 8 # round trip from client to the closest small location
toRegionRtt = 150 # round trip from that location to the single region
regionServiceTime = 40
# how many exchanges setup costs depends on protocol version and resumption
setupRoundTrips = 3
# client connects straight to the distant region
direct = setupRoundTrips * toRegionRtt + toRegionRtt + regionServiceTime
# client connects to the nearby location, which forwards over a warm connection
viaEdge = setupRoundTrips * nearEdgeRtt + toRegionRtt + regionServiceTime
# direct = 3*150 + 150 + 40 = 640
# viaEdge = 3*8 + 150 + 40 = 214
# the request leg and the region's work are identical in both branchesgo deeper
Recall that a connection and its security handshake cost several round trips before the request is even sent, and that those round trips are charged at the distance to whatever the client connected to.
Explain the arithmetic: setup crossings move to the short hop while the request leg and the region's work are unchanged. State the floor that remains and why a warm forwarded connection also avoids a slow start.
Demonstrate that you split an observed latency figure into setup, crossing and region service time before attributing it, and that you can predict when proximity will not move the number at all.
Consider whether the improvement justifies the tier across the whole estate: which traffic patterns actually pay setup often, and whether the answer is proximity, a nearer region, or fixing what the region spends its time on.
## What happens before the first request byte A request does not begin with the request. Before a client can send anything useful it has to: 1. Establish a transport connection, which costs at least one exchange with the other end. 2. Complete a security handshake, which costs one or more further exchanges depending on the protocol version in use and whether an earlier session is being resumed. 3. Only then send the request and wait for the response - one more crossing. Every one of those exchanges is a **round trip**, and the cost of a round trip is set by the distance to whatever the client is talking to. That is the whole mechanism: the setup cost is not a fixed number of milliseconds, it is a fixed number of crossings multiplied by however far the far end is. ## Where the crossings get paid If the client connects directly to a single distant region, setup crossings and the request crossing are all charged at the long distance. If the client instead connects to a location a few milliseconds away that terminates the connection, the setup crossings are charged at that short distance, and only the forwarded request pays the long one. | Leg | Direct to the region | Through a nearby location | |---|---|---| | Transport setup | Long distance | Short distance | | Security handshake | Long distance, once per exchange | Short distance, once per exchange | | Request and response | Long distance, once | Long distance, once | | Region service time | Unchanged | Unchanged | There is a second, quieter effect. The connection between the location and the region is long-lived and already carrying traffic, so it is past the cautious start that a brand-new connection goes through while it discovers how fast it may send. A fresh long-distance connection spends its early moments ramping up; a warm one does not. For a small response this matters less than the handshake; for a larger one it matters more. ## What does **not** improve This is where candidates overclaim. Terminating nearby does not move the data, does not shorten the distance the request itself must travel, and does not make the region answer faster. - The request still crosses the full distance once, and so does the response. - The region's own service time is untouched. - If the region takes a long time to produce the answer, that time dominates and the saving looks small. So the honest statement is: the thin tier removes the **setup multiplier**, not the distance. A user far from the region will still see a floor of roughly one full crossing plus the region's service time, and no amount of proximity changes that floor. ## Why this matters for uncacheable work The common misreading is that a thin tier is only useful for content it can serve from its own copy. Connection termination is the counter-example: a fully personalised, never-cacheable response benefits too, because the benefit is attached to the connection rather than to the body. That is why teams put an entirely dynamic interface behind such a tier and still measure a real improvement in time to first byte. It is also why the improvement is easy to mispredict. The saving scales with how many crossings setup would have cost and how far away the region is, and it is invisible on a connection that is already open and being reused. A client making one request on a fresh connection sees the biggest win; a client making many requests down one long-lived connection sees much less, because it paid the setup once. ## Measuring rather than assuming Before claiming the improvement, split the observed time into three parts: setup, the request crossing, and the region's service time. Each has a different remedy. - **Setup dominated** - terminating nearby is the right fix and will show up immediately. - **Crossing dominated** - the only remedies are serving from a copy nearer the user, or placing a region nearer the user. - **Service-time dominated** - neither placement helps; the work in the region is the problem. A team that skips this split frequently buys proximity and finds that the number they cared about barely moved, because the region was spending most of the budget all along.
- Which client sees the smallest benefit from this, and why?One that already holds a long-lived connection and reuses it for many requests. It pays setup once, so the saving is amortised to almost nothing and every subsequent request is dominated by the crossing and the region's service time. The biggest winners are clients making a small number of requests on fresh connections.
- Does terminating the connection nearby change what the region sees?Yes - the region sees a connection from the edge location rather than from the client, typically a long-lived one shared by many clients. That is usually good for the region's connection handling, but it means client information the region needs must be carried forward deliberately rather than read off the socket.
- If time to first byte barely moves after adding the tier, what is the likely explanation?Most of the budget was never setup. Either the region's own service time dominates, in which case placement is irrelevant and the work inside the region is the target, or clients were already reusing connections and paying setup rarely. Split the measurement into setup, crossing and service time before buying anything.
saying these in an interview costs you the question
- Claims the nearby location shortens the distance to the region
- Says an uncacheable response cannot benefit from an edge tier
- Treats connection setup as negligible next to the response itself
- Assumes the saving is the same for a reused long-lived connection
- Promises a faster answer when the region's service time dominates