skip to content

Client-side timings exceed the system's own recorded durations for the same requests in a performance run. Which path segments explain the gap?

level: seniorimportance: should knowfreq 43%

answer

  1. The two clocks start at different moments
  2. Everything outside the handler is invisible
  3. A fixed gap versus a growing gap
  4. Arrived but not yet picked up
  5. Move the measurement point earlier

basics

~20 s

A client's clock and the system's own clock cover different intervals. A gap that does not change with load is fixed transport cost outside the handler. A gap that widens with load is waiting in a segment nobody times.

solid answer

~50 s

Reconcile the two numbers by asking what interval each one covers. A system's own duration usually starts when a worker picks the request up and stops when the handler returns; a client's starts when it decides to send and stops when the last byte arrives. Everything outside those handler boundaries is invisible to the system's figure. **Expected differences** are small and roughly constant per request: connection establishment and the security handshake, time on the wire both ways, the response being written out. **Differences that matter** grow with load: time the request sat waiting before any worker took it, time queued at a tier in front that records nothing, or time acquiring a slot ahead of the timed region. Split the gap by load — the part present at low load is transport, and the part that appears only under load is unmeasured waiting.

code

pseudocode · 12 lines
pseudocode
gap(load) = client_duration(load) - system_recorded_duration(load)

fixed    = gap(very_low_load)          # setup, transit, writing the response
under_load = gap(target_load) - fixed  # appeared only when load rose

if under_load is near zero:   "the handler's own figure accounts for the request"
if under_load dominates:      "waiting in a segment no timer covers"

# narrow it by moving where the clock starts
measure_from(arrival_at_front_tier)
measure_from(arrival_at_the_system)
measure_from(worker_pickup)

go deeper

for a junior

Recall that a client and the system it calls measure different intervals, so their durations for the same request never match exactly. Know that the system's figure normally covers only the work it did, not getting there and back.

for a middle

Explain which segments each figure includes and excludes, and separate the expected fixed costs — connection setup, transit, writing the response — from waiting that appears only under load. Be ready to name what a growing gap points at.

for a senior

Show the method: compute the gap at low load and at target load, treat the difference as queueing in an untimed segment, and narrow by moving the measurement point or by testing the tiers separately. Rule out mismatched windows, statistics and populations first.

for a principal

Own where the clocks start across the estate of services and tiers a team runs. Decide which boundaries must be stamped for a request's time to be accountable end to end, and weigh that standard against the recurring cost of investigations that cannot see past a handler.

## Two timers, two intervals The first move is not to explain the gap but to state precisely what each figure measures. They are almost never the same interval. | Segment of the path | Client-side figure | System's own figure | | --- | --- | --- | | Client decides to send, prepares request | included | excluded | | Connection established, security handshake | included | excluded | | Request travels to the system | included | excluded | | Request waits at a front tier | included | usually excluded | | Request waits for a worker to take it up | included | excluded | | Handler runs, including downstream calls | included | included | | Response written out and travels back | included | partly or wholly excluded | | Client reads the last byte and stops its clock | included | excluded | Read that way, the gap is not an anomaly to be explained away. It is the sum of every segment outside the handler's own boundaries, and it should always be positive. The diagnostic question is not "why is there a gap" but **"how large is it, and does it grow?"** (This assumes the measuring harness itself had spare capacity throughout, which is established separately and before anything else.) ## Differences you should expect These appear in every run and are not defects: - **Connection setup and the security handshake** on the first request over a connection, amortised away on later requests that reuse it. - **Transit in both directions**, a roughly fixed cost per request that depends on distance and payload size. - **Writing the response out.** Many handlers stop timing when they hand back a response object, before the bytes are serialised and sent, so a large response can add real time the system never counts. - **The client's own scheduling** of whatever stops its clock, small but not zero when the client is busy. Together these give a gap that is present at the lowest load level and stays roughly the same per request as load rises. That is the shape of a fixed cost. ## Differences that mean a segment is unmeasured These grow, sometimes dramatically, as load rises: - **Waiting before the clock starts.** A request that has arrived but has not yet been taken up by a worker is invisible to a timer that starts on pickup. Under overload this is often the largest single component of what the client experienced and it appears nowhere in the system's numbers. - **A front tier that records nothing.** A proxy, balancer or gateway ahead of the system has its own queue and its own limits. If nobody times it, its waiting lands entirely inside the client's figure. - **Acquiring a slot ahead of the timed region.** Where a bounded number of requests may be in progress, the time spent waiting for permission precedes the work and is often outside the timed interval. - **A hop that was forgotten**, such as a redirect, an authentication round trip or a retry the client performed silently, so the client timed two or three request lifetimes and the system timed one. ## Splitting the gap by load The most useful single technique needs no new instrumentation. Compute the gap at a very low load level and again at the loaded level: 1. The gap at low load is the fixed transport and setup cost. 2. Subtract it from the gap at load. What remains appeared only under load, so it is **queueing in a segment nobody times**. 3. If the remainder is near zero, the system's own figure is a fair account of the request and the slowness is inside the handler. 4. If the remainder dominates, no amount of profiling the handler will find the problem — the time is being spent before the handler starts. From there, narrow by moving the measurement point rather than by arguing. Record arrival at the earliest place that sees the request, not at worker pickup, and the waiting becomes directly visible. Or narrow by elimination: run the identical profile against the front tier and then against the system behind it, and the segment whose removal closes the gap is the one carrying it. ## Comparing like with like Finally, several gaps are arithmetic rather than physical, and it is worth ruling them out before anyone investigates: - **Different windows.** One side aggregated over the whole run including its start-up portion, the other over the settled part only. - **Different statistics.** A high percentile on one side compared with a mean on the other will differ by a lot and mean nothing. - **Different populations.** If one side counts every attempt and the other counts only requests that reached a handler, the two are summarising different sets of requests. A gap is a real finding only once both figures cover the same requests, over the same window, summarised the same way. Establishing that first is cheap, and it prevents an investigation into a difference that was never there.

  • The gap is large but identical at every load level tested. What does that rule in and out?
    A constant per-request gap is a fixed cost outside the handler — connection setup, distance on the wire, a large response being written out — and it rules out queueing ahead of the timed region, which would grow with load. It also rules out an overloaded tier in front. If the constant is larger than transit plausibly explains, look for a forgotten hop such as a redirect or an authentication round trip the client timed and the system counted separately.
  • How would you make the waiting before a worker picks up a request directly visible in the next run?
    Move the start of the clock earlier. Stamp the moment the request is first seen — at the earliest point on the path that can observe it — and carry that stamp into the handler, so the handler can report both the wait before it started and the work it then did. Reporting those two figures separately turns the gap from an inference into a reading and costs one timestamp per request.
  • Before investigating a gap, what arithmetic explanations should you eliminate?
    That the two figures do not describe the same thing. Check that both cover the same window rather than one including a start-up portion, that both are the same statistic rather than a high percentile against a mean, and that both summarise the same set of requests rather than one counting every attempt and the other only those a handler saw. Any of the three produces a difference with no physical cause.

saying these in an interview costs you the question

  • Assumes both figures cover the same interval
  • Blames the network without checking where each clock starts
  • Ignores time between arrival and a worker taking the request
  • Compares a high percentile against a mean
  • Forgets an untimed tier sitting in front of the system
  • Treats the response being written out as free