skip to content

Raising server batch size from 8 to 64 triples throughput but slows each user — how do you choose?

level: seniorimportance: should knowfreq 52%

answer

  1. A frontier, not a best value
  2. Target first, curve second
  3. Find the knee, keep headroom
  4. Throughput that meets the SLO
  5. Tail latency breaks before throughput does

basics

~20 s

Fix the user-facing latency target first, then sweep concurrency and pick the largest batch size whose p95 first-token and per-token latencies still meet it. Optimise for throughput that satisfies the target, not raw tokens per second, and re-check the tail rather than the mean.

solid answer

~50 s

Batch size is a dial on a throughput-versus-latency frontier, not a value with a correct answer. Larger batches amortise each expensive weight read across more sequences, so aggregate tokens per second climb steeply at first and then flatten; meanwhile every sequence shares each step with more neighbours, so per-user token rate degrades and queueing pushes the first-token tail out. The method is to state the product target first — say 400 ms p95 first-token latency and a 30 tokens-per-second floor for a chat product — then run a load sweep at 1, 8, 32 and 64 concurrent requests and plot achieved throughput against p95 latency. Take the knee: the largest concurrency that still clears the target with headroom. What you are maximising is goodput, the traffic served *within* the SLO, and past the knee goodput falls even as raw throughput rises.

go deeper

for a junior

Know that serving more requests at once raises total tokens per second but makes each individual response stream more slowly, so the setting is a trade-off rather than a maximum.

for a middle

Explain why the curve has a knee — amortised weight reads early, contention and queueing later — and describe a concrete sweep that records tail first-token and per-token latency separately.

for a senior

Demonstrate the whole method: target before measurement, production-shaped prompt mix, goodput as the objective, headroom for bursts, and knowing that queue-driven tail latency and cached-state memory fail before throughput plateaus.

for a principal

Own the framing that the operating point encodes a business trade between serving cost and user experience, that different traffic classes deserve different points, and that the knee must be re-derived whenever the model or prompt mix changes.

## The frontier is the whole point There is no batch size that is simply best. Batching more requests together makes each generation step do more useful work per byte of weights read, so aggregate throughput rises sharply as concurrency goes from 1 to 8 to 32. It also makes each individual step take longer and forces more requests to wait their turn, so per-user token rate falls and the first-token tail lengthens. Throughput and latency sit on a frontier, and choosing a batch size is choosing a point on it. The engineering question is never "what is the fastest setting" but "which point on this curve does this product need". ## Start from the target, not the curve Measure nothing until the target exists, because a curve with no target on it cannot be read. For an interactive assistant a defensible target looks like: p95 first-token latency under 400 ms, p95 per-token latency under about 33 ms so streams sustain 30 tokens per second, defined at a stated arrival rate. That number is a product decision — a chat product where a 2-second first token reads as broken has a genuinely different requirement from an internal tool where it does not. Once the target exists, the quantity you are maximising is **goodput**: requests per second served *while meeting* the objective. Raw tokens per second is the wrong maximisation because it keeps rising after the point where the traffic it serves no longer satisfies anybody. A configuration that produces three times the tokens while 40% of them arrive too late to matter is worse than the one it replaced. ## Running the sweep Use a load generator that reproduces production shape, not a synthetic uniform prompt. Prompt-length distribution matters more than almost anything else, because long prompts inflate first-token latency and consume the memory that would otherwise hold more concurrent sequences; a sweep run on 100-token prompts will recommend a batch size that collapses on real traffic. Sweep the concurrency level — 1, 4, 8, 16, 32, 64 — holding the model, hardware and prompt mix fixed. At each point record aggregate throughput, p50/p95/p99 first-token latency and p50/p95 per-token latency, and record them separately: the two latency metrics degrade for different reasons and at different rates. Plot throughput on one axis and p95 latency on the other. The result is the classic hockey stick: an early region where throughput rises and latency barely moves, a knee, and a region past the knee where latency rises steeply for small throughput gains. Pick the largest concurrency inside the target with headroom for bursts, not the exact boundary point. Traffic is not smooth; sizing to the edge means every arrival spike violates the objective. ## Sanity-check with Little's Law Concurrency in the system equals arrival rate times time in system. If you serve 20 requests per second and each spends 6 seconds being generated, roughly 120 requests are resident at any moment. That number tells you whether the batch size you picked is even reachable at your traffic, and whether you need more replicas rather than a bigger batch. It also exposes the trap in long outputs: doubling average output length doubles residency and therefore doubles the concurrency the fleet must sustain at the same request rate. ## What breaks first past the knee Two failures show up before throughput stops improving. Queueing blows out the first-token tail — p50 stays respectable while p99 goes to multiple seconds, which is exactly the pattern that generates support tickets while dashboards look healthy. And memory for per-sequence cached state runs out, at which point the server must either reject, queue, or evict in-flight work, and evicted work is re-done later, so effective throughput falls even as the configuration claims a larger batch. ## Operate it, don't set it once The knee moves. A new model, a longer default system prompt, a change in the prompt-length mix, or a shift in average output length all relocate it. Treat the sweep as a periodic exercise and keep the latency objectives on a dashboard sliced by route, so a regression is attributed rather than argued about. And separate traffic classes rather than compromising: an interactive route and a background route pinned to one shared setting force each to pay for the other's requirements. ## The judgment being tested An interviewer asking this wants to see that you refuse the framing of a single best value, that you name a target before you name a number, that you tune on tails rather than means, and that you know which resource actually runs out first. Answering with a specific batch size and no measurement plan is the failure mode.

  • How do you decide the latency target itself rather than just measuring the curve?
    It is a product decision informed by the interaction. For a streamed chat surface the binding number is first-token latency plus a token rate that outruns reading, because everything after that is invisible to the user. For a synchronous backend call it is total completion time against the caller's own timeout. Pick the number from user behaviour, then find out what it costs in capacity.
  • What does Little's Law tell you when sizing this?
    Requests resident in the system equal arrival rate times time in system. At 20 requests per second with 6 seconds of generation each, about 120 are in flight at once. That tells you whether your chosen concurrency is even attainable, and it exposes that longer average outputs raise required concurrency at unchanged request rates.
  • Which failure appears first when you push concurrency past the knee?
    Usually the first-token tail, driven by queueing — p50 stays fine while p99 goes to seconds. Close behind is memory exhaustion for per-sequence cached state, which forces the server to queue, reject or evict in-flight sequences; evicted work is regenerated later, so effective throughput drops even though the configured batch size is larger.

saying these in an interview costs you the question

  • Picks the batch size with the highest tokens per second and ignores the tail
  • Claims batching is free because the accelerator was idle anyway
  • Tunes on mean latency and misses the p99 regression
  • Assumes throughput keeps scaling linearly with concurrency
  • Applies one latency target to every route in the product

context