Caller-side p99 to a single-node in-memory store you can query but not log into tripled overnight; how do you separate an expensive operation, a saturated server, the network, and the caller's own queue?
answer
- split the caller's number first
- pool wait versus time on the wire
- short executions can still mean saturated
- time a trivial operation from elsewhere
- sample, never enumerate, while serving
basics
~20 sSplit the caller's timing at the pool boundary first: that convicts or clears the caller's own queue in one measurement. Then read the server's long-execution record and its throughput plateau, and time a trivial operation from a host beside the caller to isolate the path.
solid answer
~50 sWork from the cheapest, least invasive measurement outward. **First**, split caller-side timing into time waiting for a connection and time from request sent to reply received — a large pool-wait component means the queue is inside your process and no server change will help. **Second**, check the server's record of long executions and whether concurrent connections sit near the server's limit: one operation family running long points at an expensive operation, usually because an entry or a collection grew. **Third**, if executions are all short but operations served per second have flattened at a ceiling while waiting time climbs, the server is saturated — fast work arriving faster than it retires. **Fourth**, time a trivial operation from a host beside the caller and from elsewhere; if both are slow with server execution flat, the path is the subject. State which shape you assume, because on a partitioned keyspace each node needs the same treatment.
go deeper
Recall that a caller-side latency figure is a sum of segments, and that the first useful move is splitting it rather than guessing which layer owns it.
Be able to state each candidate's signature — where pool wait shows, what a long execution looks like, what saturation does to served-per-second — and which measurement separates two of them.
Run the sequence in cost order on a tier that is still serving: sample rather than enumerate, keep the traffic mirror to seconds, and refuse the restart that would empty the tier and erase the evidence.
Decide beforehand which vantage points exist at all, because a tier nobody can log into with no caller-side split leaves an on-call engineer four defensible stories and no way to choose.
## Step zero: make the caller's number separable A caller-side percentile is a sum of segments, and until it is split it can be produced by any of the candidates. The one instrumentation change worth making before anything else is to time **acquiring a connection** separately from **issuing the operation and waiting for the reply**. That single split turns an undifferentiated "p99 is 400 ms" into two numbers that point in different directions, and it is available to you without any access to the server at all — which matters here, because you can query this instance but not log into the box. ## The four candidates worth separating first | Candidate | Caller-side signature | Server-side signature | Cheapest confirming test | |---|---|---|---| | One expensive operation | slow calls concentrated in one operation family or one key prefix | long executions recorded; throughput dips around them | correlate the record's entries with an entry or collection that grew | | A saturated server | all callers slow together, in proportion to load | executions short; operations served per second flat at a ceiling; connections near the server's limit | plot served-per-second against offered load and look for the plateau | | A saturated path | slow calls scale with reply size; a trivial operation is also slow | executions short and steady; served-per-second also down | time a trivial operation from a host beside the caller and from another network position | | The caller's own queue | high pool wait; connections in use pinned at the pool's maximum | quiet server, well below its connection limit | compare pool wait against call time; compare connections in use against the pool ceiling and the server's limit | These four are the ones a tier's own telemetry can distinguish. They are not exhaustive — a noisy neighbour on the host, a container limit throttling the process, or a periodic internal maintenance step can produce a fifth shape — so treat a clean sweep of all four as a signal to widen, not as proof that nothing is wrong. ## Working the sequence 1. **Is the queue inside my process?** High pool wait with connections in use sitting exactly at the pool's maximum, while the server's concurrent connections are far below its own limit, is conclusive: your callers built the queue. The requests that get sent still execute in microseconds. Adding server capacity here changes nothing, because the server was never the constraint. 2. **Is one operation expensive?** Read the server's record of operations past its time threshold, and look at whether the slow calls concentrate in one operation family or one key prefix. The usual cause is drift rather than a deploy: something that was small became large, so an operation whose cost scales with the size of the entry or the number of entries it touches started costing what it always would have. Correlate the onset with growth, not just with releases. 3. **Is the server saturated?** This is the case with no smoking gun: every execution is short and the tier is still the constraint. The signature is arrival rate above service rate — operations served per second flat at a ceiling while offered load rises, waiting time climbing, and on stores that execute one operation at a time an execution path that is simply busy all the time. Where the server is multi-threaded, look for all workers occupied rather than one. 4. **Is it the path?** Time a trivial operation — one whose execution is unambiguously microseconds — from a host beside the caller, and from somewhere else on the network. If both are slow while the server's execution figures are flat and steady, the delay is in transit, not in the store. Bandwidth saturation shows itself first on the operations with the largest replies; loss and retransmission show up as a long tail rather than a shifted median. ## What is safe to run while it serves Diagnosis on a live volatile tier can itself be the outage. Sample rather than enumerate: never ask the server for the whole keyspace in one operation, and if you take a sampled live-traffic feed to see what is actually arriving, take it for seconds — it mirrors traffic and is a load in its own right. Prefer analysing a copy off the serving path to poking the serving one. And resist the restart: it will clear the symptom, empty the tier, and destroy the evidence, while the origin behind it absorbs the full miss traffic that the tier had been absorbing. ## How the answer shifts on other shapes On a partitioned keyspace, every step becomes per node, and "the store is slow" often means one node is — so map the affected keys to their owning node before anything else. Where a copy serves reads, a lagging copy produces stale answers rather than slow ones, which is a different complaint. On a managed instance the provider's aggregates are averaged over a window that hides a short stall entirely, so caller-side timing becomes the sensitive instrument. And if the store keeps no per-operation record at all, step 2 has to be done from the caller's side by grouping call timings by operation family.
- Why is high pool wait with a quiet server not evidence that the pool is too small?It is evidence that demand exceeded the pool, which has two causes: genuinely more concurrent calls, or each call holding its connection longer because something downstream slowed. Enlarging the pool when the second is true just moves the queue onto the server and raises its connection count. Establish first whether time per call rose.
- Both a saturated server and a saturated network show short, steady server-side executions. What tells them apart?Where the work is going. A saturated server has operations served per second flat at a ceiling while offered load rises; a saturated path shows the same served-per-second falling with no ceiling behaviour, delay scaling with reply size, and a trivial operation timed from another network position that is equally slow.
- The p99 tripled but the median barely moved. What does that shape suggest?An intermittent stall rather than a uniform tax: most calls are unaffected while a minority wait behind something. That fits a queue that forms and drains — a periodic long operation, a background copy being taken, or bursty arrivals — rather than a network path that has become uniformly slower, which would drag the median with it.
saying these in an interview costs you the question
- Adds nodes or memory before splitting the caller's timing
- Restarts the tier to clear the symptom mid-incident
- Asks the server for the whole keyspace to investigate
- Treats short server executions as proof the server is idle
- Leaves a traffic-mirroring feed running through the incident
- Diagnoses a partitioned keyspace from one node's numbers