During a ramp load test you plot achieved throughput and p99 latency against the offered request rate. How do you identify the knee of the latency curve, and what number do you take away as the service's usable limit?
answer
- throughput flat, latency climbing
- queueing, not CPU
- one over one minus utilisation
- publish the SLO-bound rate, not the peak
basics
~20 sThe knee is where achieved throughput stops tracking the offered rate and p99 latency starts climbing super-linearly — queues are no longer draining. The usable limit is the highest sustained rate that still meets the latency target, which sits below the knee, not at peak throughput.
solid answer
~50 sI ramp the offered rate in steps and watch three curves together: achieved throughput, p99 latency, and error rate. Below the knee, throughput tracks the offered rate and latency stays roughly flat, because requests mostly find the server idle. At the knee, throughput flattens while latency starts rising super-linearly — arrivals are now waiting behind other requests, and queueing delay dominates service time. Past it, throughput can actually *fall* as timeouts and client retries burn capacity on work nobody will use. The number I publish is not the peak throughput: it is the highest rate at which p99 still meets the SLO over a sustained flat hold, re-verified by a separate steady-state run at that rate. One methodology caveat — a closed-loop generator with a fixed pool of virtual users cannot produce the cliff, because slow responses throttle its own send rate. Use an open, arrival-rate model to see the true shape.
code
python · 9 lines# M/M/1 idealisation: how response time explodes as utilisation rises
service_ms = 20.0
for rho in (0.5, 0.7, 0.8, 0.9, 0.95):
resp = service_ms / (1.0 - rho)
print(f"utilisation {rho:.0%} -> {resp:6.1f} ms ({resp / service_ms:4.1f}x service time)")
# Little's Law: requests in flight = arrival rate x response time
rps, resp_s = 800, 0.250
print(f"{rps * resp_s:.0f} requests in flight at {rps} rps and {resp_s * 1000:.0f} ms")go deeper
Know that latency stays roughly flat until a system nears its limit and then rises sharply, and that the safe operating rate is below that turning point rather than at maximum throughput.
Explain the mechanism: past the knee, time is spent queueing rather than working, and mean response time scales roughly as service time over one minus utilisation. Be able to describe how you would step a ramp and confirm the result with a flat hold.
Show you can defend a specific published number — what mix you tested, how long you held it, what the timeout budget was, and how retries and instance-level variance shift the cliff in production compared to the lab.
Own the argument about where the organisation chooses to sit on that curve: the latency cost of running closer to the knee versus the infrastructure cost of headroom, and which classes of service are allowed which answer.
## The three curves and what each one says A ramp test steps the offered load upward and records, at each step, the achieved throughput, the latency distribution, and the error rate. Read together they tell one story: - **Below the knee**: achieved throughput equals offered rate. Latency is roughly flat and close to the service time — the work a request needs when nothing is in its way. Requests mostly arrive to an idle server. - **At the knee**: achieved throughput flattens — the system cannot go faster. Latency turns upward sharply, because time is now spent *waiting* rather than *working*. - **Past the knee**: latency explodes, and achieved throughput frequently declines. Clients time out and retry, so the server spends capacity on requests whose callers have already given up. Goodput — useful completed work — falls even as offered load rises. ## Why the knee is where it is This is a queueing result, not a property of your code. In the simplest idealisation, a single server with random arrivals and utilisation ρ has a mean response time of `service_time / (1 - ρ)`. That denominator is the whole story: ```python service_ms = 20.0 for rho in (0.5, 0.7, 0.8, 0.9, 0.95): print(rho, round(service_ms / (1 - rho), 1)) # 0.5 -> 40ms (2x) # 0.8 -> 100ms (5x) # 0.9 -> 200ms (10x) # 0.95 -> 400ms (20x) ``` Going from 80% to 90% utilisation costs you 10% more throughput and doubles response time. That is the knee: a region where a small increase in load buys a large increase in latency. Real systems have multiple queues, batching, and non-random arrivals, so the exact numbers differ — but the shape, and the fact that the last 10% of capacity is unaffordably expensive in latency, always holds. Little's Law gives the companion view: concurrency = arrival rate × response time. At a fixed arrival rate, rising latency means more requests in flight simultaneously, which consumes threads, connections, and memory — which is why past the knee systems tend to fall over rather than merely get slow. ## Reading the knee in practice 1. **Step, don't sweep.** Hold each step long enough to reach steady state (warm-up, cache fill, pool growth, autoscaling) and discard the transition. A continuous smooth ramp mixes transient and steady-state behaviour and shifts the apparent knee to the right. 2. **Plot latency against achieved throughput as well as against offered rate.** Past saturation these two diverge, and the divergence point is itself the clearest marker of the knee. 3. **Watch error rate and the timeout budget.** Very often the true limit is reached when the p99 crosses the caller's timeout, not when the server falls over — beyond that point every additional millisecond is a failed request. 4. **Confirm with a flat run.** A rate that looked fine for a two-minute step may not survive thirty minutes at the same rate, because slow-growing queues and background work catch up. The published number should come from a held test, not from a step in a ramp. ## Open loop versus closed loop, and coordinated omission This is the most common methodology error in the whole subject. A **closed-loop** generator runs N virtual users, each of which sends a request, waits for the response, then sends the next. When the service slows down, every user slows down with it — the generator automatically reduces its own offered rate. The consequence is that you cannot drive the system past the knee: it self-throttles, latency rises gently, and the cliff never appears. It also produces the *coordinated omission* problem, where the long requests that would have been sent during a stall are simply never sent, so the recorded percentiles understate real user-visible latency badly. An **open-loop** or arrival-rate model sends at a specified rate regardless of how the service is responding, which is how real internet traffic behaves. Most modern load tools support it explicitly: ```javascript export const options = { scenarios: { ramp: { executor: 'ramping-arrival-rate', startRate: 50, timeUnit: '1s', preAllocatedVUs: 500, maxVUs: 2000, stages: [ { duration: '5m', target: 200 }, { duration: '20m', target: 200 }, { duration: '10m', target: 600 }, ], }, }, }; ``` If you must use a closed model, watch generator-side queueing and check that the generator itself is not the bottleneck — a saturated load generator produces a fake plateau that looks exactly like a real one. ## What number you actually take away Not the peak. Peak throughput is measured at latency nobody would accept and with no margin for a bad instance, a noisy neighbour, or a traffic spike. The deliverable is: > "This service sustains R requests/second per instance with p99 under T ms, verified over a 30-minute hold with production-like request mix." That R, divided into forecast peak, is what sizing decisions are made from — and it is deliberately below the knee, because operating in the knee means normal jitter becomes a latency incident. ## Where people go wrong Quoting the maximum RPS the generator ever recorded. Reading only mean latency, which stays flat well past the point where the p99 has already left the building. Testing one endpoint and generalising to the whole service. And ignoring that beyond the knee, retries make the offered rate higher than the client's nominal rate — so the cliff in production arrives sooner than the test suggested.
- Why is peak throughput the wrong number to hand to whoever sizes the fleet?Because it is measured at latency no user would accept and with zero margin. At the knee, normal variation — a slow instance, a GC pause, a retry burst — pushes you over the cliff. The usable figure is the sustained rate that still meets the latency target, and provisioning then sits below that so that ordinary jitter is absorbed rather than amplified.
- Your mean latency looks flat right through the ramp while users complain. What is happening?The mean is dominated by the fast majority. Queueing hits the tail first: p99 and p99.9 climb steeply while the mean barely moves, because only a minority of requests ever wait behind a long queue. Always plot the high percentiles against offered rate — the knee is visible there long before it shows up in the average.
- How do client retries change where the cliff appears in production compared to your test?They move it earlier. Once latency crosses the caller's timeout, each logical request generates two or three actual requests, so the offered rate the server sees exceeds the nominal client rate exactly when it is least able to cope. If your test generator does not retry the way real clients do, your measured limit is optimistic.
It is the same curve as a motorway. Adding cars increases the number passing a point until the road is near capacity; after that, throughput stops improving and travel time explodes, and once it jams, flow actually falls below what it was.
saying these in an interview costs you the question
- Reports the highest RPS the generator ever produced as capacity
- Judges the knee from mean latency instead of high percentiles
- Uses a fixed pool of virtual users and never sees the cliff
- Thinks the knee appears only when CPU nears 100%
- Treats one two-minute ramp step as a sustainable rate