skip to content

Why is speculative decoding called lossless, and what exactly does it preserve?

level: middleimportance: should knowfreq 48%

answer

  1. distribution preserved, not the string
  2. accept with min(1, p/q)
  3. rejection draws from the leftover mass
  4. bad drafter costs speed, not quality
  5. some head-based schemes relax the rule

basics

~20 s

The verification rule accepts a drafted token with probability capped by the ratio of the target's probability to the drafter's, and resamples from a corrected distribution otherwise. The result is distributed exactly as the target model alone — the drafter changes speed, not output quality.

solid answer

~50 s

Losslessness is a claim about the output **distribution**, not about byte-for-byte reproducibility. Verification uses modified rejection sampling: for a drafted token x proposed with probability q(x) where the target assigns p(x), accept it with probability min(1, p(x)/q(x)); on rejection, sample the replacement from the normalised residual distribution max(0, p - q). That test is constructed so the accepted token is drawn from p exactly, whatever q was. Under greedy decoding the equivalent rule is simpler: accept while the drafted token equals the target's argmax. So a weak drafter costs you acceptance rate — speed — and never quality. Two caveats worth naming: at non-zero temperature you still get different text run to run, because equality of distributions is not equality of samples; and not every scheme uses this test — Medusa-style heads default to a relaxed typical-acceptance threshold that trades the exact guarantee for higher acceptance.

code

python · 19 lines
python
import random


def spec_sample(p, q, draft_token):
    """Modified rejection sampling: the returned token is distributed as p."""
    if random.random() < min(1.0, p[draft_token] / q[draft_token]):
        return draft_token
    residual = {t: max(0.0, p[t] - q[t]) for t in p}
    r = random.random() * sum(residual.values())
    for t, w in residual.items():
        r -= w
        if r <= 0:
            return t
    return max(p, key=p.get)


p = {"a": 0.6, "b": 0.3, "c": 0.1}   # target distribution
q = {"a": 0.5, "b": 0.4, "c": 0.1}   # drafter over-proposes "b"
print(spec_sample(p, q, "b"))

go deeper

for a junior

Know that the small model only proposes and the big model decides, so the answer you get is the big model's answer. Saying the drafter cannot make the output worse is the point to land.

for a middle

State the acceptance rule concretely — accept with probability min(1, p/q), otherwise resample from the leftover mass — and explain that this reconstructs the target's distribution for any proposer.

for a senior

Draw the line between distributional equality and reproducibility, and mention the two real caveats: batched verification perturbs floating-point results, and typical-acceptance schemes deliberately trade the exact guarantee for acceptance rate.

for a principal

Own why the guarantee matters organisationally: it makes speculation a per-pool operational setting rather than a model release, so drafter changes need latency re-measurement, not a full quality re-certification. Insist on knowing which verifier a vendor means by lossless.

## What "lossless" is claiming Candidates often hear "lossless" and assume the server will return the identical string it would have returned without speculation. That is not the claim. The claim is distributional: the token emitted at each step is a sample from exactly the same distribution the target model would have sampled from on its own. The drafter influences *which* candidates get tested, never the law they are drawn from. ## The acceptance test Let p be the target model's next-token distribution at a position and q the drafter's. The drafter samples a candidate x from q. Verification then: - accepts x with probability min(1, p(x) / q(x)); - otherwise rejects it and samples a replacement from the residual distribution, proportional to max(0, p(t) - q(t)) over all tokens t. The intuition is a probability-mass argument. Where the drafter over-proposes a token (q(x) > p(x)), the ratio is below one and some of those proposals are rejected, shaving the excess mass down to p(x). Where the drafter under-proposes (q(x) <= p(x)), every proposal is accepted, and the shortfall is made up by the residual branch — which by construction holds exactly the mass the drafter failed to cover. Summing the two paths gives p(x) for every token. This works for *any* q, including a drafter that is badly wrong: the worse q is, the more often you land in the residual branch, which costs acceptance rate but not correctness. Greedy decoding is the degenerate case. With temperature 0, p is a point mass on the argmax, so the test reduces to "accept while the drafted token is the token the target would have chosen", and the corrected token is simply that argmax. ## Where the guarantee stops **It is not determinism.** At temperature above zero you are sampling, so a speculative run and a non-speculative run produce different text. Even with a fixed seed the random draws are consumed in a different order and quantity. Distributional equality is not sample equality, and an interviewer probing this is checking that you understand the difference. **It is not bitwise numerics.** Verification evaluates several positions in one batched pass rather than one position per pass. Different batch shapes select different kernels and different floating-point reduction orders, so logits can differ in the last bits. That is enough to flip a near-tie argmax occasionally, which is why even greedy output is "the same distribution" rather than a guarantee of identical strings against a non-speculative baseline. **It is not universal across schemes.** Draft-model speculation and n-gram / prompt-lookup speculation both plug into the same rejection-sampling verifier, so both inherit the guarantee. Medusa-style extra heads, as published, default to a *typical acceptance* rule instead: a drafted token is accepted if its target probability clears a threshold derived from the distribution's entropy, which accepts far more tokens but no longer reproduces p exactly. If a vendor advertises "lossless speculation", the question to ask is which verification rule is in use. ## Why this matters in production The guarantee is what makes speculation an *operational* knob rather than a model change. If output quality were coupled to the drafter, enabling speculation would require re-running your evaluation suite, and every drafter swap would be a model release. Because it is not, you can enable, disable, resize, or swap the drafter per replica and per traffic pool without re-certifying quality — you only need to re-measure latency and acceptance. That also shapes the right way to validate a rollout. Do not diff strings against a non-speculative baseline; on a sampling endpoint they will differ for reasons that have nothing to do with speculation. Instead confirm the verification rule in use, then compare aggregate quality metrics across a decent sample, and watch acceptance rate and latency as the real signals. If aggregate quality moves at all under a rejection-sampling verifier, suspect a bug — a KV-cache rollback error after a partial rejection is the classic cause, and it corrupts context rather than merely biasing sampling. ## The short version to say out loud A weak drafter makes you slower, not dumber. The only ways speculation degrades output are a relaxed acceptance rule chosen deliberately for speed, or an implementation bug in state rollback.

  • If it is lossless, why does the same prompt still come back with different text when speculation is on?
    Because losslessness is equality of distributions, not of samples. At any temperature above zero you are drawing from that distribution, and a speculative run consumes random draws in a different order and quantity. Even greedy runs can differ in rare near-ties, because verifying several positions in one batched pass changes kernel selection and floating-point reduction order enough to move the last bits of a logit.
  • Does a poor draft model degrade output quality?
    No — it degrades acceptance rate. Every token the drafter gets wrong is caught by the acceptance test and replaced by a sample from the target's own corrected distribution, so the emitted text remains a sample from the target model. What you lose is tokens per verification pass, and therefore the latency win, while still paying the draft cost and the wasted verification FLOPs.
  • How would you validate a speculative rollout given that guarantee?
    Not by diffing strings against a non-speculative baseline, which will differ for sampling reasons alone. Confirm the verifier is the rejection-sampling rule rather than a relaxed threshold, then compare aggregate quality metrics over a reasonable sample while watching acceptance rate and latency. Any real aggregate quality shift under the exact rule points at an implementation bug, most often KV-cache rollback after a partial rejection.

saying these in an interview costs you the question

  • Claims speculation returns byte-identical text to a non-speculative run
  • Says a weak draft model lowers answer quality
  • Thinks the rejected position just takes the target's argmax
  • Assumes every speculation scheme uses the exact acceptance rule
  • Believes losslessness means deterministic output

context