skip to content

How do you tell whether the load generator, not the system under test, is the bottleneck in a run?

level: seniorimportance: should knowfreq 44%

answer

  1. The rig is a system too
  2. Offered against achieved, plotted over time
  3. Two hosts instead of one, same workload
  4. Ports, descriptors and interface limits
  5. Idle service at its supposed ceiling

basics

~20 s

Compare achieved load against offered load and instrument the generator hosts: processor use, memory pressure, port and descriptor exhaustion, network saturation. If the achieved rate falls short while the system under test sits idle, the rig is the limit.

solid answer

~50 s

Treat the generator as a system under test in its own right. Chart **offered versus achieved** rate over time: a persistent shortfall means requests were not sent, and that has to be explained before any latency number is read. Instrument the generator hosts — processor saturation, garbage or allocation pressure, ephemeral port and file-descriptor exhaustion, network interface throughput, and any internal queue of pending work. Compare generator-side latency with timings recorded inside the system under test; a large gap that grows with load points at the generator or the path to it, not the service. The two decisive experiments are to add a second generator host and see whether total throughput rises roughly proportionally, and to hold the workload while halving it per host. A run whose generator is saturated reports a ceiling that belongs to the test rig, and quietly reintroduces closed-loop measurement bias.

code

pseudocode · 9 lines
pseudocode
run = execute(workload)

shortfall   = 1 - (run.achieved_rate / run.offered_rate)
issue_delay = percentile(run.scheduling_delay, 99)   // due time to actual send

if shortfall > 0.02 or issue_delay > 50.ms or run.rig_cpu_peak > 0.80:
    invalidate(run, reason: "generator could not sustain the offered workload")
else:
    publish(run.latency_percentiles, rig_headroom: 1 - run.rig_cpu_peak)

go deeper

for a junior

Be ready to say that the tool generating the load can run out of capacity itself, and that the first check is whether the run actually sent the requests it was configured to send.

for a middle

Explain the concrete limits — processor use on the rig, memory pressure, ephemeral ports and descriptors, interface throughput — and the comparison between generator-side and service-side timings that localises the queue.

for a senior

Demonstrate the diagnostic experiments: splitting the same workload across two rig hosts, scaling rig and load together to test linearity, and recognising the tell-tale of a service sitting idle at its supposed ceiling.

for a principal

Own rig capacity as a standing engineering concern: sizing and isolating the rig, making utilisation a published part of every result, and setting the conditions under which a run is invalidated rather than debated.

### The rig is part of the experiment A performance result is a claim about the system under test, and that claim is only valid if the generator was capable of producing the workload it was asked for. Generators are ordinary software running on ordinary machines and they saturate like anything else. When they do, the run reports a plateau — and the plateau is the rig's, not the service's. Worse, the failure is quiet: most generators do not shout that they fell behind. ### The signals, roughly in order of usefulness **Offered versus achieved rate.** The single most valuable chart. If the plan called for 460 requests per second and the run delivered 383, nothing downstream can be interpreted until that gap is explained. In a population-driven run the equivalent check is achieved iteration rate against what the population and think time predict. **Generator host resources.** Processor saturation across the generator's cores is the common one, especially when each simulated user costs a thread, or when response bodies are parsed, correlated or checksummed. Memory pressure and allocation churn slow the issuing loop. Watch the generator's own scheduling delay if it exposes one — the time between when a request was due and when it was actually issued is the most direct evidence there is. **Connection and port limits.** A host has a finite range of ephemeral ports and a finite descriptor limit, and short-lived connections hold ports in a waiting state after close. A run that climbs smoothly and then flattens at a suspiciously round concurrency, with connection errors appearing, is usually hitting one of those ceilings rather than the service's. **Network path.** The interface, the virtual network, or an intermediary between rig and service can saturate long before the service does — particularly with large payloads. Compare bytes per second against the interface's rated capacity, and remember that a proxy or gateway on the path is also a shared resource. **Timing gap.** Compare latency measured at the generator with latency measured inside the service. Some difference is real network and framing cost. A difference that grows with load is a queue somewhere between the two, and if the service's own timings stay flat while generator-side latency climbs, the queue is on the rig side. ### The two experiments that settle it **Scale the rig, not the load.** Split the same workload across two generator hosts. If throughput and latency stay the same, the rig was not the limit. If the result improves, it was. **Scale the load with the rig.** Add a second host and raise the offered load proportionally. If total achieved throughput rises roughly linearly, you have not yet found the service's ceiling; you were looking at the rig's. A third, cheaper sanity check: point the generator at a trivial endpoint on the same service that does almost no work. The rate it achieves against that endpoint is an upper bound on what your rig can produce at all. ### A worked example An insurance quote engine is exercised nightly for 6 hours. Over successive releases the report shows the same ceiling, 291 quotes per second, with a p99 that looks stable. The service's own dashboards show processor use around 38% and a request queue that never grows — a combination that should already prompt suspicion, because a service at its limit usually shows *something* saturating. Instrumenting the rig shows the generator host pinned at 97% processor use, spending most of it parsing and validating the quote response payload for every request. Splitting the identical workload across two rig hosts moves the ceiling to 574 quotes per second and the service finally starts to show its own pressure. The post-mortem also surfaced an off-by-one boundary in the rig's connection-pool sizing: the pool was created with one fewer connection than the configured concurrency, so one simulated user per host was always waiting on a connection. Small on its own, but it is a reminder that rig defects distort the shape of the result, not merely its magnitude. ### Why it matters beyond the wrong number A saturated generator does not just under-report throughput; it changes the character of the measurement. Requests wait in the rig before they are sent, which is precisely the condition under which latency measured from actual send time understates real waiting, and an arrival-rate-driven workload silently degenerates into a self-throttling one. So a rig capacity check is not housekeeping — it is a precondition for the validity of every latency statistic in the report. Publish rig utilisation alongside the results, and state the headroom the rig had, so a reader can see the result was not the rig speaking.

  • How much headroom should the generator have?
    Enough that its own resource use is uninteresting — a common working rule is to keep the rig's busiest resource well below saturation, verified rather than assumed, and to re-verify whenever payloads or per-request processing change. The specific threshold matters less than publishing the number: a report that states rig utilisation lets a reader judge validity instead of trusting it.
  • What makes a generator expensive per request, and how do you make it cheaper?
    Parsing and validating full response bodies, correlating values out of them, per-request logging of every sample, a thread per simulated user, and encryption handshakes repeated instead of reused. Cheaper alternatives: assert on status and a small extracted field rather than the whole body, aggregate latency into a compact histogram instead of writing every sample, reuse connections, and prefer a non-blocking issuing model when concurrency is high.
  • Your rig and the service share a network path with production traffic. What does that do to your results?
    It couples the experiment to something you do not control, so runs stop being comparable and a bad result may reflect a busy path rather than the service. Either place the rig where the path is dedicated and measured, or record path utilisation during every run and treat contention as an invalidating condition. Silent sharing is the version that produces confident, wrong conclusions.

saying these in an interview costs you the question

  • Assuming a plateau in throughput is always the service's ceiling
  • Never instrumenting the generator hosts at all
  • Ignoring a gap between offered and achieved rate
  • Blaming the service when it sits idle at the supposed limit
  • Adding simulated users on a saturated rig to push load higher
  • Publishing results without stating the rig's headroom

context