When would you use n-gram prompt-lookup speculation instead of a separate draft model?
answer
- what does the output copy?
- lookup needs no weights
- tokenizer must match for a draft model
- heads are trained per checkpoint
- novel prose kills lookup acceptance
basics
~20 sUse prompt lookup when the output largely copies the input — summarisation, document QA, code edits, structured rewrites. It costs no extra weights or VRAM and drafts almost instantly, but on open-ended generation with nothing to copy its acceptance rate collapses to near zero.
solid answer
~60 sThere are three families of proposer, and they differ in what they cost you at deploy time. A **separate draft model** is general-purpose but needs the same tokenizer and vocabulary as the target, occupies VRAM for its weights and its own KV cache, and drafts autoregressively — k small forward passes on the critical path. **N-gram / prompt-lookup** drafting has no model at all: it matches the last few generated tokens against earlier text in the context and copies whatever followed. Drafting is a string search, so it is effectively free, and acceptance is excellent whenever the response quotes the prompt — RAG answers, summarisation, diff-style code editing. On open-ended prose it proposes almost nothing useful. **Self-speculation heads** (Medusa, EAGLE) attach extra prediction heads to the target itself, so there is no second model to schedule, but they must be trained against each specific checkpoint. In practice: try prompt lookup first on copy-heavy workloads because it is free, and reach for a model or trained heads when generation is genuinely novel.
code
python · 11 linesdef prompt_lookup(tokens, n=3, k=4):
"""Draft k tokens by reusing the continuation of the last matching n-gram."""
pattern = tokens[-n:]
for i in range(len(tokens) - n - 1, -1, -1):
if tokens[i:i + n] == pattern:
return tokens[i + n:i + n + k]
return []
ctx = "the invoice total is 4 2 9 . the invoice total".split()
print(prompt_lookup(ctx, n=3, k=4)) # ['is', '4', '2', '9']go deeper
Know that the guess can come from a small model, from copying text already in the prompt, or from extra heads on the big model, and that copying works when the answer repeats the input.
Compare the three proposers on concrete costs: extra weights and KV for a draft model, a shared-vocabulary requirement, zero cost but workload-dependent acceptance for lookup, and a training step for heads.
Match proposer to traffic. Argue for enabling lookup on retrieval and editing routes while leaving open-ended chat alone, and size a draft model by measured acceptance and draft cost rather than by parameter count alone.
Weigh the lifecycle cost, not just the speedup: trained heads couple speculation to every checkpoint release, a draft model becomes a second artifact to version and evaluate, and lookup is the only option that survives a model swap for free.
## Three ways to produce a draft Speculation needs a proposer that is much cheaper than the target. How you get one is the real design decision, and the three families have very different deployment profiles. ### 1. A separate draft model A small model from the same family — the classic pairing is a 1B-class drafter for a 70B-class target. Requirements and costs: - **Shared tokenizer and vocabulary.** The acceptance test compares probabilities over the same token ids, so a drafter with a different vocabulary needs a mapping layer and is usually not worth it. - **VRAM.** The drafter's weights sit on the same GPU as the target, and it maintains its own KV cache for every in-flight sequence. Both come out of the pool that would otherwise hold request KV, so speculation reduces the concurrency the replica can support. - **Serial draft cost.** Producing k tokens means k small forward passes, all on the critical path. If the drafter is a tenth the size of the target, drafting 5 tokens costs roughly half a target step before verification even begins — which puts a hard ceiling on the achievable speedup. - **Generality.** It works for any workload, which is why it is the default choice for open-ended chat. ### 2. N-gram / prompt-lookup drafting No model at all. The engine takes the last n generated tokens, searches backwards through the context (prompt plus generated text) for an earlier occurrence, and proposes the k tokens that followed that occurrence. Properties: - **Zero weights, zero extra KV, negligible draft latency.** It is a substring search over an integer array. - **Acceptance is workload-shaped, not model-shaped.** When the answer reproduces material that is already in the context, matches are long and acceptance is very high. Summarisation, extractive document QA, retrieval-grounded answers that quote sources, code editing where most of a file is echoed with small changes, and structured reformatting are the sweet spots. - **It falls flat on novel text.** Free-form creative writing, translation into a language absent from the prompt, or short-prompt chat give it nothing to copy, and acceptance drops toward zero. The scheme then degrades to plain decoding plus a small verification overhead — bad but not catastrophic. - **Tuning knobs are the n-gram width and the draft length.** Wider match patterns raise precision and lower hit rate; narrower patterns match more often and get rejected more. ### 3. Self-speculation heads Medusa adds several extra prediction heads on top of the target's final hidden state, each trained to predict a different future position, and verifies multiple candidate branches at once through a tree-shaped attention mask. EAGLE instead runs a small autoregressive head at the *feature* level — predicting the next hidden state and reusing the target's own output layer — which raises acceptance relative to token-level drafting; its later revisions add dynamic draft trees and richer feature fusion. Properties: - **No second model to load, schedule or version.** The heads are tiny relative to the target. - **They must be trained per checkpoint.** Change the base model, or fine-tune it, and the heads must be retrained or at least re-validated. That is a real MLOps cost that a lookup drafter does not have. - **Verification is more complex** because tree-structured candidates need special attention masks and cache handling — this is engine work, not something you configure casually. ## How to choose Start from the workload, not the technique. If a large fraction of the response is copied from the context, prompt lookup is the highest return per unit of effort in this whole tree: no weights, no VRAM, no training, and it can be turned on and off freely. If generation is genuinely novel and you have a suitable small sibling model, a draft model is the general answer, and you should size it by measuring, not by intuition — a too-large drafter wipes out its own gains. If you serve one stable checkpoint at high value and can afford a training step, trained heads usually beat a separate model on acceptance per unit of draft cost, at the price of a per-checkpoint dependency. Engines expose these as a choice of speculation *method*. vLLM 0.27, for example, takes all speculation settings as a single `--speculative-config` object naming the method and the number of speculative tokens, rather than the separate speculative flags that earlier versions used; other servers expose their own equivalents. Whatever the spelling, the method and the draft length are the two decisions that matter.
- What acceptance rate would you expect from prompt lookup on open-ended chat with a short prompt?Close to zero. With no repeated material in the context, the n-gram search either finds no match or finds a coincidental one whose continuation the target rejects immediately. You then pay verification overhead for roughly one token per pass. That asymmetry is why prompt lookup is usually enabled per traffic pool — copy-heavy routes get it, open-ended chat does not.
- Why can't you pair an arbitrary small model with any target as its drafter?The acceptance test compares the two models' probabilities over the same token ids, so they must share a tokenizer and vocabulary. A mismatch requires a mapping layer that adds cost and dilutes acceptance. Beyond that, the drafter should be trained on similar data — an out-of-domain drafter proposes plausible-but-wrong tokens and its acceptance rate collapses even though the vocabulary lines up.
- What operational cost do trained speculation heads add that prompt lookup does not?A per-checkpoint dependency. The heads are fit against a specific base model's hidden states, so any base-model change — a new release, or your own fine-tune — invalidates them and requires retraining and re-validation before the new checkpoint can ship with speculation. Prompt lookup has no such coupling and survives model swaps untouched.
saying these in an interview costs you the question
- Assumes a draft model always beats prompt lookup
- Pairs a drafter with a different tokenizer
- Thinks prompt lookup needs its own GPU memory for weights
- Believes Medusa or EAGLE heads work on any checkpoint untrained
- Ignores that draft-model steps are themselves sequential