skip to content

What does TGI's --speculate flag enable, and when does it stop paying off?

level: seniorimportance: should knowfreq 36%

answer

  1. Extra tokens proposed per step
  2. Heads if the checkpoint has them, else n-grams
  3. Trades idle compute for fewer steps
  4. Fades as the batch fills the GPU
  5. Copying-shaped workloads accept best

basics

~20 s

TGI's --speculate sets how many extra tokens are proposed per step and verified in one forward pass - from Medusa heads when the checkpoint has them, otherwise by n-gram lookup in the existing context. It helps a lightly loaded server and fades as concurrency fills the GPU.

solid answer

~50 s

`--speculate N` turns on speculative generation in TGI. The source of the proposals depends on the checkpoint: a Medusa-style model carries extra prediction heads and N is how many of them are used; on an ordinary checkpoint TGI falls back to n-gram speculation, proposing continuations copied from text already in the context. Either way the base model verifies the proposal in a single forward pass, so a step can commit several tokens instead of one. The economics are the whole point. Speculation trades spare GPU compute for fewer sequential steps, which is a great deal when decoding is memory-bandwidth-bound and the GPU is idling between reads - a lightly loaded server, a single interactive stream. Under heavy concurrency the batch already saturates compute, so the extra verification work competes with real requests and throughput can drop. n-gram speculation also only pays where output repeats input: code edits, RAG summarisation, structured rewrites. On free-form prose acceptance is poor.

code

bash · 4 lines
bash
docker run --gpus all --shm-size 1g -p 8080:80 \
  ghcr.io/huggingface/text-generation-inference:3.3.5 \
  --model-id bigcode/starcoder2-7b \
  --speculate 3

go deeper

for a junior

Know that TGI can generate more than one token per step by guessing ahead and checking the guesses, and that it is switched on at server start rather than per request.

for a middle

Explain that proposals come from Medusa heads when the checkpoint has them and from n-gram lookup in the context otherwise, and that the base model verifies them in one forward pass.

for a senior

Argue the economics: idle compute traded for fewer sequential steps, gains shrinking as the batch saturates the GPU, and n-gram acceptance depending on how much output echoes input. Say you would measure at production concurrency.

for a principal

Treat it as a per-deployment decision tied to workload shape and load profile: enable it on interactive or copy-heavy endpoints, leave it off on saturated batch endpoints, and be willing to split deployments rather than accept one server-wide setting.

## The flag `--speculate` takes a number: how many tokens to propose beyond the one the model would normally produce this step. It can also be supplied through the launcher's environment variable form, and it is a server-level setting - you choose it when you start the container, not per request. What provides the proposals depends on what you are serving: - **Medusa-style checkpoints** ship extra prediction heads trained on top of the base model. Each head guesses a token at a future offset, so `--speculate` is effectively how many heads you use. - **Any other checkpoint** falls back to **n-gram speculation**: TGI looks for the recent token sequence elsewhere in the request's own context and proposes whatever followed it there. No draft model, no extra weights, no extra VRAM. The base model then runs one forward pass that both verifies the proposed tokens and produces the next one. Accepted proposals commit immediately; the first rejection truncates the rest. ## Why it helps - and why the help is conditional Single-stream decode is bandwidth-bound: each step streams the model weights through the GPU to produce one token, leaving arithmetic units largely idle. Verifying four candidate tokens in that same pass costs almost nothing extra in memory traffic, so if the guesses are right you got several tokens for one weight read. That is the entire trick - it converts idle compute into fewer sequential steps. The condition is that the compute was actually idle. As concurrency rises and the running batch grows, each forward pass already has plenty of arithmetic to do; the weights are being amortised across many sequences and the GPU is no longer waiting on memory. Now the speculative verification is real added work competing with real requests, and net throughput can fall below the unspeculated baseline. This is why speculation is a latency feature for lightly-loaded or interactive deployments rather than a throughput feature for saturated ones - and why you must benchmark it at your *actual* load, not at concurrency one. ## Where n-gram speculation earns its keep n-gram speculation has no draft model, so its guesses are only good when the output genuinely echoes the input. That is common in exactly the workloads people self-host for: - **Code editing / refactoring** - most of the emitted file is copied from the file you pasted. - **RAG answering and summarisation** - phrases and names are lifted from retrieved passages. - **Structured rewrites and translations of templated text** - large spans are verbatim. On free-form creative generation, where nothing in the prompt predicts the continuation, proposals are rejected constantly and you have paid verification cost for nothing. ## Choosing N Bigger N is not monotonically better. Each additional proposed token costs verification work whether or not it is accepted, and the probability that the whole run of guesses survives falls with length. Small values are the usual sweet spot; the honest method is to sweep N against your own workload and watch end-to-end latency and tokens per second together, not just one of them. Watch two things while you sweep: inter-token latency (should improve if speculation is helping) and total throughput at your target concurrency (must not regress). If they move in opposite directions, you are seeing the compute-contention effect and should either lower N or turn speculation off for that deployment. ## Operational cautions - Medusa requires a checkpoint that actually has the heads. Passing `--speculate` at a plain checkpoint gets you the n-gram path, not Medusa. - Speculation changes the shape of each forward pass, which interacts with the memory budget the server profiled at startup - re-check headroom after enabling it rather than assuming the previous batch sizing still holds. - It is a server-wide switch. If one route on the endpoint is code-completion (great acceptance) and another is chat (poor acceptance), one setting serves both; that is an argument for separate deployments. ## What interviewers listen for A weak answer says 'it makes generation faster'. A strong one names the resource being traded (idle compute for fewer sequential steps), names the regime where the trade stops working (a saturated batch), and names the workloads where n-gram proposals actually land. Being willing to say 'I would turn it off under high concurrency' is what marks the answer as measured rather than recited.

  • Why does n-gram speculation need no draft model or extra VRAM?
    Because the proposals come from text already in the request's context: TGI matches the recent token sequence against earlier positions and reuses whatever followed. There are no second-model weights to load and no separate cache to maintain, which is why it can be switched on for any checkpoint. The price is that acceptance depends entirely on the output echoing the input.
  • You enable it and tokens-per-second gets worse under production load. What happened?
    The GPU was not idle. Under a full batch each forward pass already has ample arithmetic work, so verifying speculative tokens competes with real sequences instead of filling dead time. Either lower the speculate value or disable it for that deployment, and re-benchmark at your real concurrency rather than at one stream - the concurrency-one measurement is where speculation always looks good.
  • How would you decide the speculate value for a code-completion endpoint?
    Sweep it - a small range, one value at a time - and measure inter-token latency and total throughput together at the concurrency you actually serve. Code completion is the favourable case because so much output is copied from the open file, so acceptance is high and a slightly larger value may pay. Stop increasing when throughput at target load starts to regress, even if single-stream latency is still improving.

saying these in an interview costs you the question

  • Claiming it always increases throughput
  • Assuming it requires a separate draft model
  • Benchmarking only at concurrency one
  • Expecting n-gram gains on free-form prose
  • Thinking --speculate gives Medusa on any checkpoint

context