skip to content

What belongs in an SLO for a streaming LLM endpoint, and why not just e2e p95?

level: principalimportance: should knowfreq 44%

answer

  1. length contaminates end-to-end time
  2. two components, bounded separately
  3. cap the output so e2e is derived
  4. a target with no load level is unfalsifiable
  5. measure where the caller stands

basics

~20 s

End-to-end latency scales with answer length, which the model and caller choose, so an e2e p95 target is mostly an SLO on how long answers are. Bound time to first token and per-output-token latency instead, each at a stated load level and length regime.

solid answer

~60 s

An end-to-end p95 target on a streaming endpoint is a trap: a request generating 1,500 tokens legitimately takes ten times longer than one generating 150, so the metric moves whenever prompt style or model verbosity changes, and it breaches for reasons no one can act on. Bound the two components users actually feel instead — **TTFT** (p95 and p99) and **per-output-token or inter-token latency** (p95) — and add a **max output token cap** so end-to-end time is bounded by construction. Then state the conditions the target holds under, because an inference-server latency figure is a function of load: the load level it holds at, the input/output length regime, and the measurement point (client-side, including queue time). Give each traffic class its own SLO — interactive chat needs a TTFT bound; an offline batch job needs a completion deadline and no TTFT bound at all. Tie the error budget to goodput, and define what happens when demand exceeds SLO-bounded capacity: shed load with a fast rejection rather than queue unboundedly.

go deeper

for a junior

Know that streaming endpoints are judged on how fast the first token appears and how steadily the rest streams, not on total response time, because total time depends on how long the answer is.

for a middle

Explain that end-to-end latency scales with output length, and that bounding first-token latency, per-token latency and the max output token count gives a stable target with a derived worst-case total.

for a senior

Attach the missing qualifiers — load level, prompt length regime, client-side measurement point — and show you would validate the thresholds against a measured saturation sweep rather than asserting them.

for a principal

Own the SLO as a contract: traffic classes with separate targets, an error budget denominated in goodput, an explicit overload policy, and the cost conversation that follows when the promised capacity turns out to be unaffordable.

## Why end-to-end p95 is the wrong single target On a conventional service, end-to-end latency is a clean SLI because the work per request is roughly constant. On a streaming LLM endpoint the work per request is chosen at generation time. A request that emits 1,500 tokens takes roughly ten times the wall-clock of one that emits 150 — and the model decided that, not your infrastructure. Three consequences follow. The metric moves when a prompt is reworded to elicit longer answers, which is a product change, not a reliability event. It breaches during incidents your team cannot act on, so on-call learns to ignore it. And it can look healthy while the experience is terrible: a server that streams the first token after 12 seconds and then finishes fast has a fine e2e number and an unusable product. Worse, it also *hides* good behaviour. A model that got more thorough — longer, better answers — registers as a latency regression. ## The components to bound instead **Time to first token (p95 and p99).** This is the perceived responsiveness of the product and it is the number that queueing degrades. Bound both percentiles: p95 protects the typical experience and p99 protects against the queue tail, which is where saturation shows up first. **Per-output-token latency, or inter-token latency (p95).** The streaming cadence after the first token. It is length-independent by construction, which is exactly what e2e lacks, and it maps to a comprehensible product requirement: streaming should outpace reading speed. **A maximum output token cap.** Enforce it server-side rather than trusting callers. With TTFT bounded, per-token latency bounded, and output length capped, worst-case end-to-end time is bounded arithmetically — you get the e2e guarantee as a derived property instead of as an unstable direct target. If a product genuinely needs a wall-clock guarantee — a voice agent that must answer within a turn-taking window, a synchronous API with a hard caller timeout — then state it as a deadline for a specified maximum output length, not as a percentile over unconstrained traffic. ## The qualifiers without which a target is meaningless Latency on an inference server is a function of load, not a property of the model. Every target must therefore carry: 1. **The load level it holds at.** "TTFT p95 < 500 ms at up to 20 requests/s of production-shaped traffic." Without this, the target is unfalsifiable and no capacity decision follows from it. 2. **The length regime.** Prefill dominates TTFT, so a target that holds for 500-token prompts will not hold for 32k-token ones. State the input length band, and define a separate class for long-context traffic if you serve it. 3. **The measurement point.** Measure client-side, at the edge that faces the caller, so queue time and network are inside the number. A server-internal metric that starts the clock at admission excludes exactly the delay that saturation causes — which makes it the most dangerously reassuring number available. 4. **The window and error budget.** A rolling 28-day window with an explicit budget, denominated in goodput (requests meeting the full conjunction of bounds), so tuning and capacity decisions have a currency. ## Per traffic class, not one global number One SLO across all traffic forces the most demanding class to dictate provisioning for everyone. Split by what the caller is doing: - **Interactive chat.** Strict TTFT, strict per-token cadence, modest output cap. - **Non-streaming synchronous API calls** (classification, extraction, structured output). No TTFT bound is meaningful because nothing renders progressively; bound end-to-end at a capped output length instead — here e2e *is* the right SLI, because length is bounded and small. - **Background and batch work.** No latency SLO at all. A completion deadline and a throughput floor. That split also tells you how to route and prioritise, and it prevents batch traffic from consuming the headroom that interactive traffic's SLO depends on. ## What happens at the edge of the budget An SLO is a promise, so it needs a stated behaviour when demand exceeds SLO-bounded capacity. Unbounded queueing is the default and the worst option: it converts an overload into an SLO breach for every waiting caller, including ones that would otherwise have been served comfortably. Better options, chosen deliberately: bound the queue and reject fast so callers can retry or degrade; prioritise by traffic class so interactive requests jump batch work; or fall back to a smaller, cheaper model for overflow. Whichever you pick, decide it in advance and write it into the SLO document — the failure mode is part of the promise. ## What an interviewer is listening for That you separate first-token latency from streaming cadence; that you know e2e is contaminated by output length; that you attach a load level and a measurement point to every number; that you differentiate traffic classes; and that you have an answer for overload other than "it queues". The weak answer is a single p95 number with no load level attached.

  • When is end-to-end latency actually the right SLI for an LLM call?
    When nothing streams and output length is bounded and small — classification, extraction, routing, or structured-output calls with a tight max token cap. There is no progressive rendering, so TTFT is meaningless to the user, and the capped length removes the contamination that makes e2e unstable on open-ended chat.
  • Why measure TTFT at the client edge rather than inside the engine?
    Because queue time is the delay saturation produces, and a server-internal timer that starts at admission excludes it entirely. That metric stays flat while callers experience seconds of waiting — the most reassuring possible number during an incident. Measuring at the edge that faces the caller includes admission wait and network, which is what the promise is about.
  • How do you set the actual threshold numbers rather than copying someone else's?
    Work backwards from the product. For a chat UI, first token should land inside the window where a user still perceives the system as responsive, and streaming should comfortably outpace reading speed. Then check that number against a measured sweep: if your SLO-bounded capacity at that threshold implies an unaffordable replica count, the conversation is a product one — loosen the target, cap lengths, or use a smaller model.
  • What should the SLO say about behaviour under overload?
    That queueing is bounded and excess load is shed or downgraded, explicitly. An unbounded queue spreads one overload across every waiting caller and turns a capacity shortfall into a total latency breach. Naming the response in advance — fast rejection, class-based prioritisation, or fallback to a smaller model — makes the promise honest about its own limits.

saying these in an interview costs you the question

  • A single end-to-end p95 target across all LLM traffic
  • Quoting a latency target with no load level attached
  • Measuring first-token latency inside the engine, after admission
  • One SLO shared by interactive chat and batch scoring jobs
  • Leaving overload behaviour undefined so the queue grows without bound

context