skip to content

Speculative decoding drafts 5 tokens but latency barely moved — how do you diagnose and tune it?

level: seniorimportance: must knowfreq 55%

answer

  1. two numbers: acceptance and draft cost
  2. acceptance compounds across positions
  3. the k-th token lands a^k of the time
  4. optimum draft length is interior
  5. check whether the KV pool shrank

basics

~20 s

Measure the acceptance rate and the per-position acceptance curve first, then the draft cost as a fraction of a target step. Low acceptance means a mismatched or weak drafter; high acceptance with no win means the drafter is too expensive or the draft length overshoots.

solid answer

~50 s

Two numbers decide everything: acceptance rate (accepted drafted tokens over drafted tokens) and draft cost relative to one target step. Expected tokens per verification pass is roughly 1 + a + a^2 + ... + a^k for per-token acceptance a, while cost is roughly 1 + c*k target-steps. Because acceptance compounds, the k-th drafted token only lands a^k of the time — at 70% acceptance the fifth token is worth about 17% and the eighth about 6%, yet you pay for every one of them. That makes the optimum draft length interior, usually small, and typically lower than people set it. So: if acceptance is low, the drafter is the problem — wrong domain, too small, or a lookup proposer on non-copying traffic. If acceptance is high but latency is flat, the drafter is too expensive, or the extra VRAM shrank the KV pool and pushed requests into the queue. Sweep k under real traffic rather than trusting a default.

code

python · 13 lines
python
def expected_tokens(alpha, k):
    """Mean tokens emitted per verification pass at uniform acceptance alpha."""
    return sum(alpha ** i for i in range(k + 1))


def speedup(alpha, k, draft_cost):
    """draft_cost = cost of one draft step, in target-step units."""
    return expected_tokens(alpha, k) / (1.0 + draft_cost * k)


for k in (1, 3, 5, 8):
    print(k, round(speedup(0.7, k, 0.15), 2))
# 1 1.48 / 3 1.75 / 5 1.68 / 8 1.45  -> the optimum is interior

go deeper

for a junior

Know that the useful measurement is how many guessed tokens the big model keeps, and that guessing more tokens is not automatically better.

for a middle

Explain why acceptance compounds — the k-th drafted token only counts if all earlier ones were accepted — and why that makes the best draft length small rather than as large as possible.

for a senior

Split the diagnosis cleanly into low acceptance (drafter or workload mismatch) versus high acceptance with no win (draft cost, overshoot draft length, or a shrunken KV pool), and sweep under production concurrency rather than single-stream.

for a principal

Treat the setting as perishable. Acceptance depends on traffic mix and on both checkpoints, so own a re-measurement policy tied to model releases and route changes rather than freezing a number someone found once on an idle box.

## Instrument before you tune Speculation has an unusually clean diagnostic surface, and a senior answer starts by naming the measurements rather than the knobs: - **Acceptance rate** — accepted drafted tokens divided by drafted tokens. Serving engines expose the accepted and drafted counters; the ratio is the headline number. - **Per-position acceptance** — the probability that the first, second, third... drafted token is accepted. This curve is what tells you the right draft length, and it always decays. - **Mean tokens emitted per verification pass** — the direct measure of how many sequential steps you saved. - **Draft cost** — the share of iteration wall-clock spent producing drafts rather than verifying them. - **The end-to-end metric you actually care about**, measured against real traffic, not a single-stream benchmark. ## The arithmetic that explains the disappointment Model per-token acceptance as a constant a. The expected number of tokens emitted per verification pass is 1 + a + a^2 + ... + a^k (the trailing 1 is the bonus token when everything is accepted). The cost of that pass is roughly one target step for verification plus c*k, where c is the cost of one draft step relative to one target step. Two consequences fall out immediately. **Acceptance compounds, so long drafts stop paying.** The k-th drafted token only survives if all k-1 before it survived, so its contribution is a^k. At a = 0.7 the fifth token contributes 0.17 tokens and the eighth 0.06 — while their draft cost is paid on *every* pass, accepted or not. The expected-speedup curve therefore rises, peaks, and falls. At a = 0.7 with c = 0.15 the peak sits around k = 2-3, and k = 8 is meaningfully worse than k = 3. Teams that set the draft length to 8 "to get more tokens" have usually walked past the peak. **A big drafter eats its own gains.** If c = 0.3, drafting 5 tokens costs 1.5 target steps before verification. Even at perfect acceptance you would need to emit 2.5 tokens per pass just to break even. This is the case that looks most confusing in the field: acceptance is excellent — 85% or better — and latency has not moved. The drafter is simply not cheap enough. ## Working the two failure branches **Branch A: acceptance is low (say under 50%).** - *Domain mismatch.* The drafter was trained on different data than your traffic. A drafter distilled from, or fine-tuned alongside, the target's own domain accepts far better. - *Wrong proposer for the workload.* A context-lookup proposer on open-ended generation has nothing to copy. Check whether your traffic actually repeats prompt material. - *Configuration mismatch.* Sampling settings that differ between drafter and target push proposals into the target's tail, where they get rejected. - *Prompt shape.* A drafter that is fine on short chat turns can degrade on long structured contexts. Segment acceptance by route before concluding anything. **Branch B: acceptance is high but latency is flat.** - *Draft cost.* Measure c directly. Shrink the drafter, quantize it, or move to a proposer with no serial model steps at all. - *Draft length past the peak.* Sweep k downward; it very often improves both latency and throughput at once. - *The KV pool shrank.* Drafter weights and drafter KV came out of the same GPU memory the request cache lives in. If concurrency dropped, requests are now waiting in the queue, and queueing time swamps the per-token gain. Look at whether admitted concurrency fell when speculation was enabled. - *The server is not in the regime where speculation helps.* If the replica is running a large batch, the verification work is no longer close to free. ## How to run the sweep Sweep draft length under representative concurrency, not at batch size 1, and read the end-to-end latency distribution rather than the mean. Hold everything else fixed and change one variable at a time, because draft length interacts with both the memory budget and the batch the scheduler forms. Record acceptance alongside latency in every run; a sweep that only records latency cannot tell you *why* a setting won, and the next model release will make you redo the whole thing. ## The trap to name out loud Speculation tuned on an idle box, measured single-stream, will show a beautiful speedup that evaporates under production load. Always validate the setting at the concurrency you actually run at.

  • Acceptance is 85% but latency did not improve at all. What is your first hypothesis?
    That the drafter is too expensive relative to a target step. At 85% acceptance you should be emitting three or more tokens per verification pass, so the win is being spent somewhere — most likely k serial draft forward passes from an oversized proposer. Measure the share of iteration time spent drafting; if it is large, shrink or quantize the drafter, or cut the draft length before touching anything else.
  • Why is per-position acceptance more useful than the aggregate acceptance rate when choosing draft length?
    Because the aggregate hides the decay. Two setups can both report 60% overall while one accepts the first four tokens reliably and the fifth never, and the other accepts each position at a flat rate. The per-position curve tells you exactly where the marginal drafted token stops earning its cost, which is the definition of the right k. Cut the draft length at the point the curve falls below your break-even contribution.
  • Enabling speculation dropped your p95 even though per-token latency improved. What happened?
    Almost certainly a capacity effect rather than a decoding effect. The drafter's weights and its per-sequence KV cache came out of the same GPU memory the request KV pool uses, so the replica now admits fewer concurrent sequences. Requests spend longer queued, and queueing time dominates the per-token gain. Compare admitted concurrency and queue wait before and after, and consider a smaller drafter or fewer speculative tokens.

saying these in an interview costs you the question

  • Raises the draft length to get more tokens per pass
  • Reports acceptance rate without the per-position curve
  • Ignores the draft model's own serial forward passes
  • Tunes speculation single-stream and ships the setting
  • Forgets that drafter memory shrinks the request KV pool

context