skip to content

Why does speculative decoding stop paying off as an inference server's batch size grows?

level: seniorimportance: should knowfreq 52%

answer

  1. spare capacity is the fuel
  2. busy GPUs have none to spare
  3. rejected drafts steal real work
  4. acceptance differs across the batch
  5. gate speculation above a batch threshold

basics

~20 s

At small batch a decode step has idle arithmetic capacity, so scoring extra drafted positions is nearly free. As concurrency rises the GPU's matrix units fill up, and the extra positions — most of the rejected ones pure waste — start costing real time and stealing throughput from other requests.

solid answer

~60 s

Speculation converts spare arithmetic into fewer sequential steps. That trade is only free while arithmetic is spare. With one or a handful of sequences in flight, a decode step is dominated by moving weights and the extra drafted positions ride along for almost nothing. As the running batch grows, the server is already doing plenty of arithmetic per weight read, so verifying k+1 positions per sequence multiplies the real work by roughly the draft length, and everything rejected is discarded compute that other requests needed. Three server-side effects compound it: sequences in a batch accept different numbers of drafts, so the iteration is shaped by the longest draft and the rest is padding waste; the drafter's weights and per-sequence KV cache shrink the pool that limits concurrency; and draft steps serialise ahead of every verification. The practical answers are to gate speculation on load — many engines can skip it above a batch-size threshold — or to split traffic into a latency pool that speculates and a bulk pool that does not.

go deeper

for a junior

Know that guessing ahead helps when the GPU is not busy, and stops helping once many users are being served at the same time.

for a middle

Explain that verification rides free only while arithmetic capacity is idle, and that rejected drafted tokens are real computed work which becomes expensive once the GPU is full.

for a senior

Name the server-side effects: wasted verification competing with queued requests, heterogeneous acceptance inside a batch, drafter memory shrinking admitted concurrency, and serial draft passes. Then give the operational response — gate on load or split the pools.

for a principal

Own the fleet-level shape: speculation is a latency purchase, not a capacity purchase, so decide per traffic class and hold the decision to a goodput measurement under real concurrency rather than a benchmark speedup.

## The regime, stated plainly Speculation is a bet that the GPU has arithmetic capacity going unused during decode. At batch size 1 that bet is overwhelmingly correct: the step's cost is dominated by streaming weights, and adding drafted positions barely moves wall-clock. As the running batch grows, the server is naturally doing more arithmetic per weight read, and the unused capacity disappears. Past that point every drafted position is real work, and the roughly (k+1)x multiplier on positions-per-sequence turns into a roughly proportional cost. That is the core of the answer, but the interviewer usually wants the serving-specific consequences, which are more interesting than the regime argument itself. ## Effect 1: rejected work is stolen from other requests At low load, wasted verification FLOPs cost nobody anything — the alternative was idle silicon. At high load the GPU is the shared resource every queued request is waiting for, so a rejected drafted token is capacity taken from a real user. A 60%-acceptance setup at draft length 4 discards a large share of its verification work; at low load that is invisible, at saturation it is a direct throughput tax. This is why aggregate tokens per second on a busy fleet usually *falls* when speculation is enabled, even though per-request latency looks better in a single-stream test. ## Effect 2: batches do not accept uniformly Every sequence in a batch drafts its own candidates and accepts a different number of them. The iteration must evaluate all drafted positions for all sequences, so the work is set by the full draft width, while the useful output is set by each sequence's own acceptance. Some engines mitigate this by dropping speculation for sequences whose recent acceptance is poor, but the general shape holds: heterogeneous acceptance inside a batch means padding waste that grows with batch size. ## Effect 3: the memory bill hits concurrency A draft model occupies weights and, for every in-flight sequence, its own KV cache. Both come out of the same GPU memory that would otherwise hold request KV. Fewer concurrent sequences fit, so under load requests spend longer in the queue. It is entirely possible to enable speculation, improve inter-token latency, and *worsen* end-to-end p95 because queueing grew. When someone reports that, the first thing to compare is admitted concurrency before and after. ## Effect 4: draft steps serialise With a draft model, the k small forward passes happen before verification and are on the critical path for the whole batch. As batch size grows those draft passes also grow in cost, so the overhead does not stay constant — it scales with the batch you were hoping to amortise it over. ## What to do about it **Gate on load.** The cleanest answer is dynamic: skip speculation when the running batch exceeds a threshold, and use it when the server is quiet. This is a first-class feature in serving engines — vLLM's speculative configuration, for example, exposes a batch-size threshold above which speculation is bypassed. Set the threshold empirically at the point your measured speedup crosses 1.0. **Split the fleet.** If you have both interactive and bulk traffic, run two pools: a speculation-enabled pool sized for low concurrency and strict latency, and a plain pool tuned for batch throughput. Routing by workload is usually cleaner than trying to make one pool serve both, because the two want opposite scheduler settings. **Measure goodput, not speedup.** The number that decides this is how many requests per second you serve *within* your latency target, at production concurrency. A single-stream speedup measurement will always flatter speculation, because single-stream is exactly the regime it is designed for. **Re-check after any capacity change.** Adding replicas, changing GPU SKU, or a traffic shift all move where the crossover sits. The threshold is a measured property of a specific deployment, not a constant. ## The one-line version Speculation buys latency with spare compute. When there is no spare compute — because your users are already using it — there is nothing left to buy it with, and continuing to speculate makes the fleet slower for everyone.

  • Would you rather disable speculation under load or reduce the draft length?
    Reducing draft length is the gentler lever and often enough: it cuts wasted verification work while keeping some of the win. Disabling is the right call once measured speedup crosses below 1.0, because past that point you are strictly worse off. In practice set both — a smaller draft length in the mid range and a hard cut-off threshold above it — and calibrate each from a concurrency sweep.
  • Why can enabling speculation improve inter-token latency but worsen end-to-end p95?
    Because the drafter's weights and per-sequence KV cache consume GPU memory that the request KV pool needs, so the replica admits fewer concurrent sequences. Under load requests then queue longer, and queue time is part of end-to-end latency while inter-token latency only measures the decoding of admitted requests. The fix is a smaller drafter, a shorter draft length, or accepting that this pool should not speculate.
  • How do you pick the batch-size threshold at which to stop speculating?
    Measure it. Sweep concurrency with speculation on and off, holding everything else fixed, and plot goodput — requests served inside the latency target — for both. The crossover point is your threshold, with a little margin for traffic variance. It is deployment-specific: a different GPU, a different draft length or a shift in traffic mix moves it, so re-run the sweep after capacity or model changes.

saying these in an interview costs you the question

  • Assumes the single-stream speedup holds at production concurrency
  • Says speculation raises a saturated server's throughput
  • Ignores that rejected drafts consume shared GPU capacity
  • Forgets the drafter's memory reduces admitted concurrency
  • Treats the batch-size cut-off as a universal constant

context