skip to content

A ticket site's on-sale dashboard shows a healthy mean response time while buyers report waits - what measurement would actually prove the system is responsive?

level: seniorimportance: must knowfreq 62%

answer

  1. a bound, not a central tendency
  2. which request, which percentile, which rate
  3. the dropped requests were the slow ones
  4. start the timer when the request was issued
  5. plot the percentile against arrival rate

basics

~20 s

A high percentile of one named request, measured over every outcome including timeouts and refusals, at a stated arrival rate, across the spike window. Responsiveness is a bounded tail under load, not a healthy average.

solid answer

~50 s

Responsiveness claims a bound, so the evidence has to be a bound, not a central tendency. I would name the request - the seat-reservation call - and report a high percentile such as the 99th, measured client-side from the moment the request was issued, with timed-out and refused requests included in the distribution rather than dropped from it. The claim also needs the arrival rate and the window it holds at, because response time is a function of load: the same code answers in milliseconds at one rate and queues without bound above the rate it can serve. The real artefact is a load curve - that percentile plotted against arrival rate - which shows both the bound and the load at which it breaks. A mean over a large population barely moves when a small minority waits for seconds, which is precisely the minority that is complaining.

code

pseudocode · 17 lines
pseudocode
// what the healthy mean was measuring
on response(r):
    if r.status is OK:
        record(r.timeInsideHandler)

// what a responsiveness claim needs
on requestIssued(id, t0):
    started[id] = t0

on outcome(id, kind):                 // OK, degraded, refused, timed out
    t0 = started.remove(id)
    record(now() - t0, kind)          // every outcome enters the distribution

report:
    percentile(99, all records)
    at arrivalRate = 12000 per second
    over window = on-sale minute 0 to minute 10

go deeper

for a junior

Know that an average response time can look healthy while a minority waits seconds, and that percentiles exist to describe that minority. Naming the 99th percentile as the thing to look at is the expected answer here.

for a middle

Explain why the mean is insensitive to a small slow group, why timed-out requests must stay in the distribution, and why a latency figure without an arrival rate beside it cannot be checked.

for a senior

Produce the artefact: a load curve for a named request, measured over all outcomes from when the request was issued, with the knee where the bound breaks identified. Say what you would do above that knee.

for a principal

Decide what bound the organisation defends and at what load, make that the acceptance criterion every service is measured against, and require any change sold as improving performance to show the curve rather than a single figure.

## What responsiveness actually claims Responsiveness is the goal of the four system properties, and it is the only one stated in terms of somebody outside the system. The claim has three parts, and an answer that supplies fewer than three is not checkable: 1. **Which request.** "The system is fast" is not a claim. "The seat-reservation request" is. 2. **Which bound, at which percentile.** A bound is a statement about the worst case that is still acceptable, so it is carried by a high percentile - commonly the 99th or the 99.9th - not by the mean or the median. 3. **At which arrival rate, over which window.** Response time is a function of load, so a bound with no load attached cannot be falsified. ## Why the mean hid the problem A mean over a large population is nearly insensitive to a badly served minority. Take the scenario: a 120 ms mean over the on-sale window. If one request in a hundred waits ten seconds, that minority contributes about `0.01 x (10 - 0.12)` seconds to the mean, which is roughly a tenth of a second - so the dashboard shows something like 220 ms and nobody looks twice. Meanwhile, at a million requests through the window, ten thousand people waited ten seconds. Two effects make the gap worse than the arithmetic suggests: - **Distributions under load are usually bimodal.** Most requests are served immediately and a minority waits behind a queue. A mean sits between the two modes and describes neither. - **Tail exposure compounds per user.** If one page needs `n = 20` backend calls to render and each independently has a 1-in-100 chance of landing in the slow tail, the page avoids the tail entirely with probability `0.99^20 = 0.82`. So about one page view in five hits a tail that the per-call statistic reports as one per cent. ## The two measurement traps that flatter a collapsing system - **Excluding requests that did not complete.** A request cut off at a deadline has an end time - the deadline - and belongs in the distribution at that value or above. If only successful responses are recorded, the system gets *faster* on the dashboard as it starts shedding load, because the slowest requests are exactly the ones being removed from the sample. Record the outcome alongside the duration, and report the percentile over all of them. - **Measuring from the wrong instant.** If the timer starts when a worker picks the request up, everything spent queueing before that point is invisible, and queueing is the whole story during a spike. The same distortion appears in load generation: a generator that waits for each response before sending the next one stops issuing requests exactly while the system is slow, so it under-samples the worst period. Issue on a schedule and measure from the moment the request was due. ## Why the arrival rate belongs in the claim Little's Law relates the mean number of requests in the system, the mean arrival rate and the mean time each spends there: `L = arrival rate x time in system`. Rearranged, it says the throughput a bounded concurrency can sustain is that concurrency divided by the service time. A path that can hold 200 requests at once, each taking 100 ms, tops out near 2,000 per second. Offer it 20,000 per second and there is no steady state: the queue grows for as long as the excess lasts and the wait grows with it. No code change inside the path alters that ceiling. This is why "our 99th percentile is 300 ms" is an incomplete sentence. At what rate? ## The artefact that actually settles it | Artefact | What it shows | Why it is not enough alone | |---|---|---| | Mean response time | central tendency of the sample | insensitive to the minority that is complaining | | A single percentile figure | one point on a curve | no load attached, so it cannot be falsified | | Success-only distribution | how fast the survivors were | improves as the system sheds load | | Load curve: percentile against arrival rate, all outcomes | the bound and the rate at which it breaks | this is the evidence; keep the window and the outcome mix beside it | A load curve also answers the question behind the question. Interviewers are not testing whether you know what a percentile is; they are testing whether you know that a responsiveness claim is a conditional statement, and that the condition is load. ## Reading the result When the curve bends sharply at some arrival rate, that knee is the system's honest capacity, and the bound is only claimable below it. Above the knee, the useful question stops being "how fast" and becomes "what does the system do with work it cannot serve" - which is where the other three properties start doing their job.

  • Why does a load generator that waits for each response before sending the next one understate the tail?
    Because it stops issuing requests exactly while the system is slow. Every second a response is late is a second the generator did not offer new load, so the worst period is sampled least. Issuing on a fixed schedule and timing from when each request was due keeps the slow window represented in the distribution.
  • The 99th percentile of every individual call is inside the bound, but users still report slow pages. How?
    Fan-out. If one page makes twenty backend calls and each independently has a one per cent chance of landing in the tail, the page escapes it with probability 0.99^20, about 0.82 - so roughly one page view in five is slow. Measure the user-facing operation end to end, not only its parts.
  • Does responsiveness have to hold during a failure, or only under load?
    Both. The property claims a bound under the load and the failures the system says it handles, which is why resilience is one of the means. A request that depends on a broken component should still answer inside the bound - with a degraded or explicit failure answer - rather than hang until the caller's deadline expires.

An average water depth of one metre is no comfort to the person crossing at the three-metre point. The average describes the river; the tail describes the crossing.

saying these in an interview costs you the question

  • Quoting a mean response time as proof of responsiveness
  • Measuring latency only over requests that succeeded
  • Stating a latency bound with no arrival rate attached
  • Assuming a healthy average means nobody waited long
  • Starting the timer after the request left the queue