skip to content

Reading Results

What happens once a run stops: composing latency distributions, counting errors and completed work honestly, attributing a slowdown and judging it against an agreed rule.

on this pageshow

questions

22

How do you tell queueing delay from service time when a performance run recorded only end-to-end duration?

level: middleimportance: must knowfreq 58%

answer

  1. Two kinds of time in one number
  2. Waiting for a turn versus doing work
  3. Baseline the operation with nothing competing
  4. Vary concurrency, watch which part moves
  5. Rate multiplied by duration equals requests in flight

basics

~20 s

Service time is what one request costs with nothing competing for the resource; queueing delay is what waiting adds. Measure at a concurrency of one, then watch duration grow as concurrency rises — the growth is wait, not work.

solid answer

~50 s

**Service time** is the work a request costs when nothing competes with it; **queueing delay** is the time it spends waiting for a busy resource before that work starts. An end-to-end duration is their sum, so one number cannot be split by inspection — you have to vary something and watch which part responds. Take a low-concurrency baseline first: the same operation run on its own, repeated, gives service time almost pure. Then hold the run at rising concurrency levels and record duration and completed work per second at each. Service time stays roughly flat until something shared starts contending; wait rises with the number of requests in flight, and rises without limit once arrivals meet the ceiling. The shape of the spread corroborates it: work alone gives a tight, single-peaked spread, while waiting stretches the slow end further at every step up in load.

code

pseudocode · 12 lines
pseudocode
service_time = median(duration of the operation run alone, repeated 20 times)

for level in [1, 2, 4, 8, 16, 32]:
    hold(concurrency = level, minutes = 5)
    d    = median(duration measured at this level)
    rate = completed_requests / measured_seconds
    wait_estimate = d - service_time
    in_system     = rate * d          # arrival rate times duration
    report(level, d, rate, wait_estimate, in_system)

# duration climbing while rate is flat  -> the growth is waiting
# duration and rate both climbing        -> still service-dominated

go deeper

for a junior

Recall the two components by name: service time is the work itself, queueing delay is waiting for a busy resource, and an end-to-end figure is their sum. Be able to say why one recorded duration cannot be split without a second measurement.

for a middle

Explain the mechanics: a single-request baseline gives service time, a ladder of concurrency levels reveals wait, and arrival rate multiplied by duration cross-checks how many requests were inside the system. Know that flat throughput with rising duration means waiting.

for a senior

Show the judgement that the subtraction can lie. Demonstrate that work can genuinely get slower under overlap, that a changed operation mix invalidates the baseline, and that the two diagnoses lead to opposite remedies — more servers versus cheaper work.

for a principal

Own the measurement standard rather than the individual diagnosis: decide that every layer records time-to-admit separately from time-to-serve, so the split is read rather than argued, and weigh that instrumentation cost against the hours teams lose relitigating a single summed number.

## The two kinds of time inside one number A request's end-to-end duration contains exactly two kinds of time: time spent being worked on, and time spent waiting for something that was busy with someone else. - **Service time** is the cost of doing the work once the request holds whatever it needs, with nothing competing for it. - **Queueing delay** is the interval between the request being ready to be served and actually being served. Every layer on the path has its own pair. There is a slot to acquire before a worker runs the request, a connection to acquire before an outbound call leaves, a place in line at a shared store. What was recorded is the sum of all of them, and a sum can never be split by staring at it. You separate the components by **changing one thing and watching which part of the sum moves**. ## Get a service-time baseline first The cheapest useful observation is the operation on its own: one request at a time, repeated twenty or so times, in the same environment and against the same data as the loaded run. With nothing else in flight there is nothing to wait for, so what you record is service time plus the fixed cost of getting bytes there and back. Record the middle of that sample and its spread — repeated single requests should land close together. That number is your yardstick. Under load, the amount by which a request's duration exceeds it is waiting, subject to one caveat below. ## Read the response to concurrency Now hold the run at a ladder of increasing concurrency levels — several minutes at each, measuring only the settled portion — and record duration and completed work per second at every level. The two components respond to concurrency in opposite ways, and that is what makes them separable. | Observation as concurrency rises | Queueing-dominated | Service-dominated | | --- | --- | --- | | Median duration | rises roughly with requests in flight | stays flat | | Completed work per second | flattens at a ceiling | rises with concurrency | | Spread of durations | slow end stretches further each level | stays tight | | Arrival rate multiplied by duration | tracks the number in flight | stays low | The last row is **Little's Law** used as an arithmetic check rather than a prediction: the average number of requests inside a system equals the arrival rate multiplied by the average time each one spends inside it. Two of those three quantities are almost always recorded, so the third is not a matter of opinion. When measured duration climbs while completed work per second refuses to, essentially all of the added time is queueing — the system is finishing the same amount of work and simply making each request wait longer for its turn. ## The caveat: service time itself can move Subtracting the single-request baseline is only honest if the work per request did not change. Two things break that assumption: 1. **Contention inside the work.** If requests fight over something shared — an exclusive section over one record, one refillable cache entry, one counter — the work itself genuinely takes longer when requests overlap. That is not queueing ahead of the resource; it is service time that grew. 2. **A different mix.** If the loaded run exercises a heavier blend of operations than the baseline did, the two numbers describe different work and the subtraction is meaningless. The first is worth naming explicitly rather than folding into wait, because the fix is different. Wait ahead of a resource is relieved by more servers, more slots, or fewer arrivals; work that got slower under overlap is relieved only by shrinking or removing the shared section. ## Why the distinction decides what to do next This is not a taxonomy exercise. The two diagnoses point at opposite remedies, and applying the wrong one is a common and expensive mistake: - If the time is **service**, the work must get cheaper or the resource faster. Adding parallelism multiplies the same cost across more requests and does not reduce any single request's duration. - If the time is **wait**, the resource behind the queue is often sitting with spare capacity while durations climb. Making that resource faster helps a little; increasing how many requests may be served at once, or reducing how many arrive, helps a lot. One more practical point: apply the same reasoning **per layer**, not just end to end. A request whose end-to-end time is dominated by wait may be waiting at one specific layer while every other layer is idle. Recording, for each layer, both the time to acquire the resource and the time to do the work once acquired turns the whole exercise from inference into direct reading — and it is far cheaper to add that pair of measurements before the next run than to argue about a single summed number afterwards.

  • Your single-request baseline is 20 ms and under load the median is 200 ms. What would make it wrong to call the extra 180 ms queueing?
    The subtraction assumes the work itself did not change. If requests contend over something shared — an exclusive section on one record, a single cache entry being refilled — the work genuinely costs more when requests overlap, so part of the 180 ms is service time that grew, not time spent in line. A different operation mix under load breaks the subtraction the same way. Confirm by measuring service time inside the loaded run, not just outside it.
  • Completed work per second is flat while durations climb steadily. What does that combination alone tell you?
    That the system is finishing the same amount of work per second and making each request wait longer for its turn, so almost all of the added time is queueing rather than work. It does not say where the queue is or what the constraint is — only that arrivals now exceed what the path can serve. The next step is to locate which layer's line is growing, by recording time-to-acquire separately from time-to-serve at each layer.
  • How would you make this separation directly readable in the next run instead of inferring it?
    Record two figures at each layer rather than one: the time between the request becoming ready for that layer and being admitted to it, and the time from admission to completion. The first is wait, the second is service, and their sum reconciles to the layer's contribution. That turns a subtraction argument into a reading, and it costs two timestamps per layer.

A twenty-minute wait for a ten-minute haircut and a thirty-minute haircut with no queue both leave the chair after thirty minutes; only changing how many people are waiting tells you which one you sat through.

saying these in an interview costs you the question

  • Treats end-to-end duration as if it were all work
  • Adds capacity to a resource that is already idle
  • Never measures the operation at a concurrency of one
  • Assumes service time is constant across every load level
  • Compares a loaded run against a baseline with a different operation mix
  • Concludes from a single load level with nothing to compare
open as a page

In a performance run, what does a gap between offered and achieved throughput tell you?

level: middleimportance: must knowfreq 62%

basics

~20 s

Offered throughput is the demand the run applied; achieved throughput is the work that actually completed and was counted. A gap means requests were refused, expired or never finished, so the completed-work figure describes only part of the intended load.

open as a page

Before a performance run starts, what must its pass rule fix for the run's verdict to mean anything?

level: middleimportance: must knowfreq 70%

basics

~20 s

Fix five things before the run starts: which measurement, at which point in its distribution, over which window, under which workload, and read from which vantage point. Anything left open is chosen afterwards to suit the numbers.

open as a page

Why can a performance run's per-interval 95th-percentile times not be averaged into one figure for the whole run?

level: middleimportance: must knowfreq 62%

basics

~20 s

Percentiles are ranks over a set of samples, not quantities that add or average. The mean of per-interval figures weights a quiet minute like a busy one and throws away each interval's shape. Merge raw samples or bucket counts instead.

open as a page

What has to be held constant for two performance runs to be comparable at all?

level: middleimportance: must knowfreq 62%

basics

~20 s

Comparable performance runs hold the same build, the same dataset size and shape, the same environment and its neighbours, the same warm state and the same applied workload. When one of those differs, a change in the numbers cannot be attributed.

open as a page

In a performance run, why is a reply that arrives with a success status not necessarily a success?

level: juniorimportance: should knowfreq 55%

basics

~20 s

A success status only says something answered the request. The body may be a rendered error page, a truncated document, or an empty result where a data record was expected, so a run reading only the status overstates how much work really succeeded.

open as a page

How do you confirm a fixed pool of connections or workers, not the resource behind it, caps a performance run's throughput?

level: middleimportance: should knowfreq 50%

basics

~20 s

A fixed pool caps throughput near its slot count divided by mean service time, so throughput flattens while duration rises with offered concurrency. Change the slot count: if the ceiling moves in proportion, the pool was the cap.

open as a page

A performance run passed on average response time although many requests were far slower. How do you write the pass rule so that cannot happen?

level: middleimportance: should knowfreq 52%

basics

~20 s

Attach the bound to a stated point in the distribution rather than a central measure, pair it with an outer cap on the worst reported interval, and say how many intervals may breach before the run fails.

open as a page

In a performance run mixing several operation types, how can one slow operation vanish from the aggregate 95th percentile?

level: middleimportance: should knowfreq 55%

basics

~20 s

An aggregate percentile ranks every request together, so weight follows request share, not importance. An operation carrying under 5% of requests can be entirely above the 95% cut and never move it. Report a distribution per operation type.

open as a page

Before you call a difference between two performance runs a regression, what must you measure first?

level: middleimportance: should knowfreq 52%

basics

~10 s

Measure the run-to-run spread first: repeat one unchanged configuration several times and record how much the compared figure varies on its own. A difference smaller than that spread is not evidence of anything.

open as a page

Client-side timings exceed the system's own recorded durations for the same requests in a performance run. Which path segments explain the gap?

level: seniorimportance: should knowfreq 43%

basics

~20 s

A client's clock and the system's own clock cover different intervals. A gap that does not change with load is fixed transport cost outside the handler. A gap that widens with load is waiting in a segment nobody times.

open as a page

How do you attribute a performance run's slow response tail to one step on the request path when each step's own average looks acceptable?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Per-step averages hide a step that is slow occasionally or runs many times per request. Compare each step's share of total time in the slowest requests against median ones, then confirm by neutralising the suspect.

open as a page

In a performance run, how do you count requests that expired, were reset, or were never sent?

level: seniorimportance: should knowfreq 47%

basics

~20 s

Count them as outcomes, not as missing data. Each stays inside the issued total, lands in a named failure category, and the report states what time was recorded for it. Deleting them describes only the requests the system managed to serve.

open as a page

A performance run finished with no pass rule agreed beforehand. What can its report honestly conclude?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Description only: what was applied, where timings were taken, and where the numbers sat. Not a pass or a fail - a bound chosen once results are visible is fitted to them. Its real output is the next run's rule.

open as a page

Why can one latency pass rule not judge a long steady hold, a sudden step in arrival rate and a deliberate overload run?

level: seniorimportance: should knowfreq 34%

basics

~20 s

Each shape asks a different question. A long hold must still meet the bound in its final intervals; an abrupt increase in arrival rate is judged on recovery time; a run past the intended rate is judged on refusal, not latency.

open as a page

A performance run's response times form two distinct clusters. What does one 95th-percentile figure get wrong?

level: seniorimportance: should knowfreq 38%

basics

~20 s

It implies one population with one centre. With two clusters the figure lands in whichever cluster holds the rank, describes only that one, and moves with the proportion between them rather than with either cluster's own speed.

open as a page

A request crosses four services, each with a known 99th-percentile latency. Why is the end-to-end 99th percentile not their sum?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Summing per-hop 99th percentiles prices a request in which all four hops were simultaneously in their own slowest one percent, which is rare — so the sum usually overstates. Correlated hops can make it understate. Measure whole-path time per request instead.

open as a page

How do you set a regression threshold for performance runs from measured noise rather than a round number?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Derive it from the measured run-to-run spread: set the alarm level a multiple above the spread of repeats, per metric, then check that size against what change actually matters. A round percentage chosen by habit is either noise or nothing.

open as a page

When is adopting a new performance reference run legitimate rather than moving the goalposts?

level: principalimportance: should knowfreq 38%

basics

~20 s

It is legitimate when the shift in level has an identified, intentional cause, was reproduced across repeats, and still meets the obligation the team owes. It is a moved goalpost when the reason is that meeting the old figure became inconvenient.

open as a page

A slowdown reproduces only when requests overlap. How do you design the follow-up performance run that confirms it?

level: seniorimportance: nice to knowfreq 34%

basics

~20 s

Vary overlap alone. Run the same arrival rate with few busy clients and then with many idle ones: if duration differs, overlap is the cause. Repeat with requests spread across distinct records to locate what is shared.

open as a page

In a performance run, how much of each reply can you afford to verify in flight?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

Only checks cheap enough to fit inside the applying side's own budget: a length floor, a required field, a marker. Deep parsing per reply steals the capacity that generates demand, so sample it instead and report what share was verified.

open as a page

When a 99th percentile is computed from bucketed latency counts, what error do the bucket boundaries introduce?

level: seniorimportance: nice to knowfreq 22%

basics

~20 s

The read-out locates the bucket holding the ranked request, not its time, so the figure carries that bucket's width as uncertainty. Interpolating inside assumes a spread the slow end does not have, and an unbounded top bucket reports nothing at all.

open as a page