How do you tell queueing delay from service time when a performance run recorded only end-to-end duration?
answer
- Two kinds of time in one number
- Waiting for a turn versus doing work
- Baseline the operation with nothing competing
- Vary concurrency, watch which part moves
- Rate multiplied by duration equals requests in flight
basics
~20 sService time is what one request costs with nothing competing for the resource; queueing delay is what waiting adds. Measure at a concurrency of one, then watch duration grow as concurrency rises — the growth is wait, not work.
solid answer
~50 s**Service time** is the work a request costs when nothing competes with it; **queueing delay** is the time it spends waiting for a busy resource before that work starts. An end-to-end duration is their sum, so one number cannot be split by inspection — you have to vary something and watch which part responds. Take a low-concurrency baseline first: the same operation run on its own, repeated, gives service time almost pure. Then hold the run at rising concurrency levels and record duration and completed work per second at each. Service time stays roughly flat until something shared starts contending; wait rises with the number of requests in flight, and rises without limit once arrivals meet the ceiling. The shape of the spread corroborates it: work alone gives a tight, single-peaked spread, while waiting stretches the slow end further at every step up in load.
code
pseudocode · 12 linesservice_time = median(duration of the operation run alone, repeated 20 times)
for level in [1, 2, 4, 8, 16, 32]:
hold(concurrency = level, minutes = 5)
d = median(duration measured at this level)
rate = completed_requests / measured_seconds
wait_estimate = d - service_time
in_system = rate * d # arrival rate times duration
report(level, d, rate, wait_estimate, in_system)
# duration climbing while rate is flat -> the growth is waiting
# duration and rate both climbing -> still service-dominatedgo deeper
Recall the two components by name: service time is the work itself, queueing delay is waiting for a busy resource, and an end-to-end figure is their sum. Be able to say why one recorded duration cannot be split without a second measurement.
Explain the mechanics: a single-request baseline gives service time, a ladder of concurrency levels reveals wait, and arrival rate multiplied by duration cross-checks how many requests were inside the system. Know that flat throughput with rising duration means waiting.
Show the judgement that the subtraction can lie. Demonstrate that work can genuinely get slower under overlap, that a changed operation mix invalidates the baseline, and that the two diagnoses lead to opposite remedies — more servers versus cheaper work.
Own the measurement standard rather than the individual diagnosis: decide that every layer records time-to-admit separately from time-to-serve, so the split is read rather than argued, and weigh that instrumentation cost against the hours teams lose relitigating a single summed number.
## The two kinds of time inside one number A request's end-to-end duration contains exactly two kinds of time: time spent being worked on, and time spent waiting for something that was busy with someone else. - **Service time** is the cost of doing the work once the request holds whatever it needs, with nothing competing for it. - **Queueing delay** is the interval between the request being ready to be served and actually being served. Every layer on the path has its own pair. There is a slot to acquire before a worker runs the request, a connection to acquire before an outbound call leaves, a place in line at a shared store. What was recorded is the sum of all of them, and a sum can never be split by staring at it. You separate the components by **changing one thing and watching which part of the sum moves**. ## Get a service-time baseline first The cheapest useful observation is the operation on its own: one request at a time, repeated twenty or so times, in the same environment and against the same data as the loaded run. With nothing else in flight there is nothing to wait for, so what you record is service time plus the fixed cost of getting bytes there and back. Record the middle of that sample and its spread — repeated single requests should land close together. That number is your yardstick. Under load, the amount by which a request's duration exceeds it is waiting, subject to one caveat below. ## Read the response to concurrency Now hold the run at a ladder of increasing concurrency levels — several minutes at each, measuring only the settled portion — and record duration and completed work per second at every level. The two components respond to concurrency in opposite ways, and that is what makes them separable. | Observation as concurrency rises | Queueing-dominated | Service-dominated | | --- | --- | --- | | Median duration | rises roughly with requests in flight | stays flat | | Completed work per second | flattens at a ceiling | rises with concurrency | | Spread of durations | slow end stretches further each level | stays tight | | Arrival rate multiplied by duration | tracks the number in flight | stays low | The last row is **Little's Law** used as an arithmetic check rather than a prediction: the average number of requests inside a system equals the arrival rate multiplied by the average time each one spends inside it. Two of those three quantities are almost always recorded, so the third is not a matter of opinion. When measured duration climbs while completed work per second refuses to, essentially all of the added time is queueing — the system is finishing the same amount of work and simply making each request wait longer for its turn. ## The caveat: service time itself can move Subtracting the single-request baseline is only honest if the work per request did not change. Two things break that assumption: 1. **Contention inside the work.** If requests fight over something shared — an exclusive section over one record, one refillable cache entry, one counter — the work itself genuinely takes longer when requests overlap. That is not queueing ahead of the resource; it is service time that grew. 2. **A different mix.** If the loaded run exercises a heavier blend of operations than the baseline did, the two numbers describe different work and the subtraction is meaningless. The first is worth naming explicitly rather than folding into wait, because the fix is different. Wait ahead of a resource is relieved by more servers, more slots, or fewer arrivals; work that got slower under overlap is relieved only by shrinking or removing the shared section. ## Why the distinction decides what to do next This is not a taxonomy exercise. The two diagnoses point at opposite remedies, and applying the wrong one is a common and expensive mistake: - If the time is **service**, the work must get cheaper or the resource faster. Adding parallelism multiplies the same cost across more requests and does not reduce any single request's duration. - If the time is **wait**, the resource behind the queue is often sitting with spare capacity while durations climb. Making that resource faster helps a little; increasing how many requests may be served at once, or reducing how many arrive, helps a lot. One more practical point: apply the same reasoning **per layer**, not just end to end. A request whose end-to-end time is dominated by wait may be waiting at one specific layer while every other layer is idle. Recording, for each layer, both the time to acquire the resource and the time to do the work once acquired turns the whole exercise from inference into direct reading — and it is far cheaper to add that pair of measurements before the next run than to argue about a single summed number afterwards.
- Your single-request baseline is 20 ms and under load the median is 200 ms. What would make it wrong to call the extra 180 ms queueing?The subtraction assumes the work itself did not change. If requests contend over something shared — an exclusive section on one record, a single cache entry being refilled — the work genuinely costs more when requests overlap, so part of the 180 ms is service time that grew, not time spent in line. A different operation mix under load breaks the subtraction the same way. Confirm by measuring service time inside the loaded run, not just outside it.
- Completed work per second is flat while durations climb steadily. What does that combination alone tell you?That the system is finishing the same amount of work per second and making each request wait longer for its turn, so almost all of the added time is queueing rather than work. It does not say where the queue is or what the constraint is — only that arrivals now exceed what the path can serve. The next step is to locate which layer's line is growing, by recording time-to-acquire separately from time-to-serve at each layer.
- How would you make this separation directly readable in the next run instead of inferring it?Record two figures at each layer rather than one: the time between the request becoming ready for that layer and being admitted to it, and the time from admission to completion. The first is wait, the second is service, and their sum reconciles to the layer's contribution. That turns a subtraction argument into a reading, and it costs two timestamps per layer.
A twenty-minute wait for a ten-minute haircut and a thirty-minute haircut with no queue both leave the chair after thirty minutes; only changing how many people are waiting tells you which one you sat through.
saying these in an interview costs you the question
- Treats end-to-end duration as if it were all work
- Adds capacity to a resource that is already idle
- Never measures the operation at a concurrency of one
- Assumes service time is constant across every load level
- Compares a loaded run against a baseline with a different operation mix
- Concludes from a single load level with nothing to compare