A dashboard says an in-memory store answered in 0.3 ms while the caller reports 40 ms for the same calls; what does each number measure?
answer
- two clocks, not one
- the server's clock starts late
- pool wait, network, queueing
- the gap is the finding
basics
~20 sThe store's number is server-side execution only, from the server starting the operation to the reply being produced. The caller's number adds waiting for a connection, both network hops and time queued before execution. That gap is where the slowdown lives.
solid answer
~50 sTwo clocks are running and they start in different places. The store reports **execution time**: the clock starts when the server begins handling a request it has already read and stops when the reply is produced, so it covers the work on the entry and essentially nothing else. The caller's stopwatch starts on the line of code that asks for the value and stops on the line that receives it, so it also contains the wait for a free connection in its pool, both network traversals, any time the request sat at the server before execution began, and any pause in the caller's own process. Neither number is wrong; quoting one without saying where it was taken is. `0.3 ms` of execution inside `40 ms` of caller time says the store is not the thing that is slow, and names the layer to look at next.
go deeper
Recall that two different clocks are in play and that the server's starts only when it begins the work. If asked which number is right, the answer is both, for different spans.
Be able to enumerate what lives between the clocks — pool wait, both network hops, waiting to be dispatched, the caller's own scheduling — and say which layer a large gap points at.
Use the gap as the first triage step before spending anything on the tier, and know that on a managed instance the provider's smoothed aggregates hide short stalls the caller's timing still sees.
Make the vantage point part of the operating contract: every latency figure teams exchange carries where it was taken, or the organisation argues for a day about two true numbers.
## Two clocks, and they start in different places Every statement that "the store is slow" is a number somebody took somewhere, and a volatile tier is routinely measured from two vantage points that barely overlap. **Server-side execution time** is what the store reports for an operation it ran. Its clock starts when the server begins handling a request it has already read off a connection, and stops when the reply has been produced. It measures work on the entry: locating the key, reading or changing the value, building the reply. It knows nothing about how the request got there or what happened to the reply afterwards. **Caller-observed time** is what the application measures around its own call: from the line of code that asks for the value to the line that receives it. Everything that happens in between lands in that number, whether the store caused it or not. On a perfectly healthy tier these two differ by a large factor, and that is expected — a microsecond-scale operation reached over a network is dominated by the network, not by the work. ## What sits in the gap 1. **Pool wait.** Before anything is sent, the caller needs a connection from its pool. If every connection is in use, the call waits inside the caller's own process. Nothing has been sent and the store does not know the request exists. 2. **The outbound hop.** Encoding the request and moving it across the network, including delay on a saturated link or a retransmitted packet. 3. **Waiting to be dispatched.** The request has arrived but execution has not begun: it may sit behind other requests, and on a server that executes one operation at a time it may sit behind a single long-running one. 4. **Execution.** The only segment server-side execution time covers. 5. **The return hop and the caller's own scheduling.** The reply travels back, and the caller's thread must be scheduled to read it. A saturated or paused caller process adds time here that has nothing to do with the store. | Segment | In server-side execution time | In caller-observed time | |---|---|---| | Waiting for a connection in the caller's pool | no | yes | | Network transit, each direction | no | yes | | Waiting at the server before execution begins | usually no | yes | | Work on the entry itself | yes | yes | | Caller thread scheduling and process pauses | no | yes | ## Why this is the first question, not the last Subtracting one number from the other localises the problem before anyone spends money. If caller-observed time is roughly execution plus a stable network baseline, the store's own work is the subject and you go looking at what it is being asked to do. If the gap is nearly the whole figure, the store is a bystander: the delay is in the connection path, the network, or the caller's own process, and adding memory or nodes to a tier that was never busy fixes nothing while costing real money and a maintenance window. The corollary is a habit: a latency figure with no vantage point is not evidence. "Zero point three milliseconds" and "forty milliseconds" are both true of the same traffic in the same second. ## What varies between stores and deployments - Not every store in this class publishes per-operation timing, and where it exists the clock does not cover identical spans: most count execution only, some note separately how long a request waited before execution began. Confirm which yours does before you subtract one number from the other. - Where the server executes operations on many threads, waiting-to-be-dispatched is usually smaller, but it does not vanish: work still queues when arrivals exceed what the server can serve. - On a managed instance you may not be able to see the process at all. You get the provider's aggregates over a fixed window, which smooths a short stall into a barely visible bump, so the caller's own timing becomes the more sensitive instrument, not the less trustworthy one. - An aggregate across callers is not any caller's experience, and percentiles collected per instance cannot be averaged into a fleet percentile. ## Saying it well The answer an interviewer is listening for is the reflex, not the arithmetic: name the clock with the number. "Zero point three milliseconds as the server measures its own execution; forty milliseconds as the caller measures the whole call, of which pool wait and network are unaccounted." That sentence turns a disagreement into a direction.
- The caller's number and the server's number agree closely. What does that tell you?That almost nothing outside execution is contributing: the pool is handing out connections immediately, the network hop is small relative to the work, and nothing is queueing at the server. It also means the work itself now dominates, so if the figure is too high the investigation belongs to what the callers are asking the store to do, not to the path.
- One application instance out of twelve reports high call times against the same store. Where do you look first?At that instance, not at the store. A single caller diverging while its peers are fine points at something local: its connection pool sized too small or leaking connections, its own process starved or paused, or its network path. A server-side cause would show up across independent callers at roughly the same time.
saying these in an interview costs you the question
- Declares the store healthy because server-side latency is flat
- Quotes caller-side timing as if it were the store's execution time
- Gives a latency number without saying who measured it
- Assumes the network contributes nothing on a local link
- Averages per-instance percentiles into a single fleet percentile