skip to content

How do you prove a Triton batching config change actually improved throughput?

level: seniorimportance: should knowfreq 38%

answer

  1. a curve, not a single number
  2. concurrency sweep before and after
  3. queue time versus compute time
  4. requests over executions equals batch size
  5. Model Analyzer searches, perf_analyzer verifies

basics

~20 s

Sweep load with perf_analyzer at increasing concurrency and compare throughput against p95 latency before and after. Then confirm the mechanism: Triton's execution count versus successful request count gives the achieved average batch size, and the queue-time component shows what the wait actually cost.

solid answer

~40 s

Load-test with `perf_analyzer -m <model> --concurrency-range 1:16:1 --percentile=95`, which reports throughput plus a latency breakdown including server queue time and compute input/infer/output time. That breakdown is the diagnosis: if raising `max_queue_delay_microseconds` added queue time without moving throughput, the change bought nothing. Corroborate from the server's own metrics — dividing `nv_inference_request_success` by `nv_inference_exec_count` gives the mean batch size the scheduler actually formed, which is the number the config was meant to change. For choosing among many candidate configs rather than validating one, `model-analyzer profile` generates config variants across instance counts and batching settings, measures each, and reports which satisfy a latency constraint. One currency note: `perf_analyzer` and `genai-perf` now ship from the `triton-inference-server/perf_analyzer` repository, not the server repo.

code

bash · 9 lines
bash
perf_analyzer -m resnet50 -u localhost:8001 -i grpc \
  --concurrency-range 1:16:1 --percentile=95

curl -s localhost:8002/metrics | grep -E 'nv_inference_(exec_count|request_success)'

model-analyzer profile \
  --model-repository /models \
  --profile-models resnet50 \
  --output-model-repository-path /tmp/ma-output

go deeper

for a junior

Know that perf_analyzer is the tool that loads a Triton model and reports throughput and latency, and that you compare before and after rather than trusting the config change.

for a middle

Run a concurrency sweep rather than a single point, and read the latency breakdown — queue time versus compute time — instead of only the throughput figure.

for a senior

Tie the result to a mechanism: use the request-success over execution-count ratio to show the achieved batch size moved, isolate one variable per run, and confirm on production metrics rather than synthetic load alone.

for a principal

Decide when a config search is worth the GPU hours at all — Model Analyzer sweeps at onboarding and hardware changes — and define what the organisation records as evidence before a serving config ships.

## Why a single before/after number is not proof Batching settings move two quantities in opposite directions. A change that raises throughput 30% while doubling p95 latency is not obviously an improvement — whether it is depends on what you promised. So a valid comparison is never one number: it is a **curve**, throughput plotted against latency as offered load rises, measured the same way on both configurations. And even a better curve is not yet an explanation. You also want the mechanism: did the scheduler actually form larger batches? ## perf_analyzer `perf_analyzer` is the standard load generator for Triton. Two load modes matter: - **Concurrency mode** (`--concurrency-range 1:16:1`) keeps N requests outstanding at all times. This is a closed loop: if the server slows down, the client sends more slowly. It maps the throughput/latency frontier well. - **Request-rate mode** (`--request-rate-range`) sends at a fixed offered rate regardless of how the server is coping. This is an open loop and is the honest way to see behaviour under overload, because queues actually grow. Useful flags: `-m` for the model, `-u` and `-i grpc` for the endpoint and protocol, `-b` for the *client-side* batch size, `--percentile=95` to report a tail rather than the average, `--shape` and `--input-data` to supply realistic inputs. That last one matters more than people expect: random zeros are fine for a fixed-shape vision model and misleading for anything shape-sensitive. What you read from the output is not just "inferences/sec". The report breaks latency into client send/receive, network, **server queue**, and compute input, compute infer and compute output. That decomposition is the whole point: - Queue time up, compute time flat → requests are waiting on the scheduler or for a free instance. Either the delay knob is too generous, or you need more instances. - Compute infer up with batch size up → you are converting latency into throughput as intended. - Compute input/output large → you are spending time on tensor copies, which more instances (overlapping copy with compute) help more than bigger batches do. ## The metric that names the mechanism Triton's Prometheus endpoint exposes per-model counters. Two of them answer "did batching happen?": - `nv_inference_request_success` — number of successful inference *requests*. - `nv_inference_exec_count` — number of model *executions*. Requests divided by executions is the average batch size the scheduler achieved. If you set a 5 ms queue delay and the ratio is still 1.02, the batcher is not finding requests to merge and the delay is pure added latency. This single ratio settles more batching arguments than any amount of reasoning about the config file. The duration counters — `nv_inference_queue_duration_us`, `nv_inference_compute_input_duration_us`, `nv_inference_compute_infer_duration_us`, `nv_inference_compute_output_duration_us` — give the same decomposition as perf_analyzer but from real production traffic instead of synthetic load, which is where you should ultimately confirm the change. ## Model Analyzer, for search rather than verification When the question is "which of these dozens of configurations is best" rather than "did this one change help", hand-running perf_analyzer does not scale. `model-analyzer profile` takes the model repository and a model name, generates variant configurations across instance counts and batching settings, runs measurements against each, and produces reports ranking them — typically under a latency constraint you supply, so you get the highest-throughput configuration that still meets the bound. It writes the generated variant configs to an output repository so the winner can be promoted into your real repository. It is not free: a full sweep runs many measurement windows and takes real GPU time, so run it once when a model is onboarded or when hardware changes, not on every deploy. ## Method notes that decide whether the result is trustworthy - **Warm up.** First executions include lazy initialisation and allocator warm-up. Discard them. - **Isolate.** Measure on a GPU that is not shared with another model, or you are measuring your neighbour. - **Move one knob at a time.** Instance count and queue delay both change throughput; changed together, you learn nothing about either. - **Run the client off the server.** A perf_analyzer process on the same host competes for CPU with the backend's own threads and can bottleneck before the GPU does. - **Use realistic payload sizes.** Compute input/output time scales with tensor size, and that is exactly the term batching does not amortise. ## Where the tools live now `perf_analyzer` and `genai-perf` used to be found alongside the server and, earlier still, in the client repository. They now ship from `triton-inference-server/perf_analyzer`. Pointing a colleague at the server repository for them is a stale instruction, and it is a small but real currency signal in an interview.

  • When would you use perf_analyzer's request-rate mode instead of concurrency mode?
    Concurrency mode is a closed loop — the client only sends a replacement when a response returns — so the server can never be pushed past what it can absorb, and queues stay bounded. Request-rate mode sends at a fixed offered rate regardless, which is the only way to observe genuine overload: growing queue duration, rejections from queue policy, and the point where latency stops being a function of load.
  • Your throughput improved but the requests-per-execution ratio did not move. What else could explain the gain?
    Something other than batching changed. More instances overlapping input and output copies with compute will raise throughput at constant batch size, as will a warmed allocator, a different client protocol, or simply having measured before warm-up last time. Attribute the change to a mechanism before you keep it, because a gain you cannot explain usually will not survive a different traffic mix.
  • Why is running perf_analyzer on the same host as the server a problem?
    The load generator needs CPU to serialize inputs, manage connections and parse responses, and it takes that CPU from the same pool the backend's own threads and copy engines use. You end up measuring a client bottleneck and concluding the server saturated earlier than it does. Run the client on a separate machine on the same network, or at minimum verify client CPU is not pinned.

saying these in an interview costs you the question

  • Comparing one latency number before and after
  • Never checking whether batch size actually changed
  • Load testing without any warm-up period
  • Changing instance count and queue delay together
  • Looking for perf_analyzer in the Triton server repo

context