When is speculative decoding the wrong way to buy inference latency for a fleet?
answer
- what is the fleet actually limited by?
- latency instrument, not a capacity one
- it preserves the output distribution
- weigh the second artifact's lifecycle
- judge it on goodput under load
basics
~20 sWhen the constraint is cost or capacity rather than per-request latency. Speculation raises total GPU work and consumes memory that limits concurrency, so a throughput-bound or budget-bound fleet is usually better served by more replicas, a smaller or quantized model, or better batching.
solid answer
~60 sSpeculation buys one thing: fewer sequential steps per response, for a single request. It does not raise a fleet's throughput ceiling — it lowers it, because rejected drafts are computed and discarded, and the drafter's weights and KV shrink the pool that governs concurrency. So the question to ask is what your fleet is actually limited by. If it is cost per million tokens or requests per GPU, the levers are a smaller model, quantization, better batching, and more replicas. If it is a strict inter-token or end-to-end latency target on low-concurrency interactive traffic, speculation is one of the few levers that helps without changing the model's output distribution — which is its unique property. Then weigh the lifecycle cost: a draft model is a second artifact to version, evaluate and roll out; trained speculation heads must be refit for every base-model change; a context-lookup proposer costs nothing and survives checkpoint swaps. I would default to enabling it only on the pools whose traffic is latency-critical and copy-heavy, gate it on load, and hold it to a goodput measurement.
go deeper
Know that speculation makes one response arrive faster but does not let a busy server handle more users, so it is not a general cost fix.
Be able to separate a latency problem from a throughput or queueing problem, and say that speculation only helps the first while consuming memory and compute that the others need.
Argue the deployment shape: enable per traffic pool, gate on load, prefer a free lookup proposer where the workload copies context, and validate on goodput at production concurrency rather than a benchmark speedup.
Own the tradeoff end to end — the unique value is preserving the output distribution, the hidden costs are capacity and a second versioned artifact, and the decision needs a review trigger tied to model releases so a stale setting does not become a permanent tax.
## Ask what the fleet is limited by Every latency technique answers a different constraint, and speculation answers a narrow one. Sort the situation first: - **Cost-bound** — you are trying to reduce dollars per million tokens. Speculation makes this worse: more FLOPs per emitted token and less concurrency per replica. Serve a smaller model, quantize, or improve batching. - **Capacity-bound** — you cannot admit enough concurrent sequences. Speculation makes this worse too, because the drafter takes memory from the request cache. Free memory or add replicas. - **Queue-bound** — requests wait before they start. That is a scheduling and scaling problem, and speeding up decoding for the ones already running does not touch it. - **Latency-bound on the decode path** — users are waiting on tokens arriving too slowly, at modest concurrency. This is the case speculation was built for, and it is the only one where it is the right first answer. The misdiagnosis to avoid is treating speculation as a general performance feature. On a fleet already running near saturation it is a throughput tax dressed as an optimisation. ## What speculation uniquely offers Its distinguishing property is that it preserves the model's output distribution. Every other lever for lowering decode latency changes what the model says — a smaller model, an aggressive quantization, a shorter context, truncated output — and therefore requires a quality re-evaluation and a product conversation. Speculation does not. That makes it the right tool when quality is contractual or heavily evaluated and you cannot afford a behavioural change, and it is worth saying so explicitly in an interview: you are paying compute to keep the output law fixed while shortening the critical path. ## The lifecycle bill The speedup is the visible cost; the maintenance is the hidden one, and it differs sharply by proposer: - **A separate draft model** is a second artifact in the registry, with its own version, its own deployment, its own memory footprint, and a compatibility constraint (shared tokenizer and vocabulary) that survives only as long as nobody upgrades the target carelessly. - **Trained speculation heads** are the strongest per-unit-cost proposer and the most coupled. They are fit to a specific checkpoint's internals, so every base-model release or fine-tune invalidates them. If you ship model updates frequently, that is a recurring training and validation step on your release critical path. - **Context-lookup drafting** carries no lifecycle cost at all — no weights, no training, no version coupling. When it fits the workload, it is nearly always the right place to start, precisely because it can be turned off without consequence. A principal-level answer weighs these against the release cadence of the team, not just against a benchmark. ## How to structure the decision 1. **Segment traffic.** Interactive chat and bulk batch jobs want opposite settings. Run them as separate pools with separate configurations rather than compromising one server for both. 2. **Enable per pool, gated on load.** Speculate on the latency pool, skip it above a measured concurrency threshold, and leave the bulk pool alone. 3. **Match proposer to workload.** Retrieval-grounded, summarisation and code-editing routes get lookup drafting for free. Genuinely novel generation needs a model or heads, and only if the latency case justifies the artifact. 4. **Hold it to goodput.** The acceptance criterion is requests served within the latency target at production concurrency — not a single-stream speedup number, which always flatters speculation. 5. **Set a review trigger.** Acceptance depends on traffic mix and on both checkpoints. Tie re-measurement to model releases and significant routing changes, or the setting silently decays into a tax. ## The honest summary Speculative decoding is a targeted latency instrument with a real capacity cost and a real maintenance cost, whose unique merit is that it does not change what the model says. Deployed where the constraint is decode latency at modest concurrency, it is excellent. Deployed fleet-wide because it sounded like free performance, it quietly reduces the number of users each GPU can serve.
- Your cost per million tokens is the problem, not latency. Where does speculation rank among your options?Last, or negative. It increases FLOPs per emitted token and shrinks the KV pool, so it reduces the requests each GPU can serve. The levers that actually address cost are a smaller or distilled model, weight quantization with validated quality, better batching and scheduling, and right-sizing the GPU SKU. Speculation should probably be turned off on a cost-bound pool, not tuned.
- What makes speculation different from simply serving a smaller model to hit a latency target?A smaller model changes what the system says and requires a full quality re-evaluation and, often, a product decision. Speculation preserves the target model's output distribution exactly under the standard verification rule, so it lowers latency without a behavioural change. That is its unique merit and the reason it is worth its compute overhead where quality is contractual or heavily evaluated.
- How does your release cadence affect which proposer you choose?Strongly. Trained speculation heads give the best acceptance per unit of draft cost but are fit to one checkpoint, so a team shipping model updates often adds a retraining and validation step to every release. A separate draft model is looser but still a versioned artifact with a tokenizer compatibility constraint. Context-lookup drafting has no coupling at all and survives checkpoint swaps untouched.
saying these in an interview costs you the question
- Enables speculation fleet-wide as free performance
- Expects speculation to reduce cost per million tokens
- Ignores that the drafter reduces admitted concurrency
- Forgets trained heads must be refit for each checkpoint
- Judges adoption on a single-stream speedup number