Which TGI /metrics series separate queueing time from GPU inference time?
answer
- Prometheus endpoint on the serving port
- Split waiting from computing
- Two gauges: queue size, batch size
- GPU utilisation is a useless saturation signal
- Overload has its own status code
basics
~20 sTGI exposes Prometheus metrics on /metrics, where tgi_request_queue_duration records time spent waiting before the GPU touched a request and tgi_request_inference_duration records the model work itself. Splitting slow end-to-end latency across those two is the first triage step.
solid answer
~50 sTGI serves a Prometheus endpoint at `/metrics` on the same port as the inference API. The series that matter for triage are the per-request timers - `tgi_request_queue_duration` for time waiting to be admitted to a batch, `tgi_request_inference_duration` for time doing model work, and `tgi_request_duration` for the end-to-end figure - plus `tgi_request_mean_time_per_token_duration` for inter-token pace and the gauges `tgi_queue_size` and `tgi_batch_current_size` for what the router is doing right now. The triage rule follows directly. If end-to-end latency is bad and queue duration dominates, you are admission- or capacity-limited: add replicas, or admit more eagerly. If inference duration dominates, the model step itself is slow - wrong precision, too-large batch, prompts longer than you assumed. Those gauges are also the right autoscaling signal: queue depth and TTFT track saturation on an LLM server, whereas GPU utilisation stays pinned near 100% and tells you nothing.
code
bash · 1 linecurl -s 127.0.0.1:8080/metrics | grep -E 'tgi_queue_size|tgi_batch_current_size|tgi_request_queue_duration'go deeper
Know that TGI exposes Prometheus metrics at /metrics on the same port as the API, and that latency has two parts - waiting and computing - that are recorded separately.
Name the queue-duration and inference-duration series and explain what each direction implies: waiting means capacity or admission, computing means prompt length, precision or batch size.
Walk the triage tree from a real incident, cross-check the queue-size and batch-size gauges, and explain the 429 overloaded contract as deliberate load shedding rather than a defect.
Set the SLO and the scaling signal: pick queue depth or TTFT over GPU utilisation, accept that weight-loading cold starts make the reaction slow, and define what the endpoint promises callers under shed load.
## Why an LLM server needs its own signals On a stateless web service, a p95 latency chart plus CPU utilisation gets you a long way. Neither works here. GPU utilisation on a busy inference server sits near 100% almost all the time - a single decode step keeps the device nominally busy - so it cannot tell a comfortable server from a drowning one. And end-to-end latency lumps together two completely different problems with completely different fixes. TGI's Prometheus endpoint at `/metrics` exists to separate them. ## The series to know **Per-request timers** - `tgi_request_queue_duration` - how long the request sat before the router admitted it into a batch. - `tgi_request_inference_duration` - how long the model actually spent on it. - `tgi_request_validation_duration` - time in request validation, before any scheduling. - `tgi_request_duration` - the end-to-end figure the client experiences. - `tgi_request_mean_time_per_token_duration` - the inter-token pace, which is what a streaming user perceives as 'speed' after the first token arrives. **Batch and queue gauges** - `tgi_queue_size` - requests waiting right now. - `tgi_batch_current_size` - requests decoding right now. **Counters and size distributions** such as `tgi_request_count` and `tgi_request_generated_tokens` round out the picture, letting you turn latency into per-token economics. ## The triage decision tree Start from bad end-to-end latency and split it: 1. **Queue duration dominates.** Requests are waiting, not computing. Either the server is genuinely at capacity - add replicas - or the admission policy is too lazy and queued work is not being folded into the running batch quickly enough. Cross-check `tgi_queue_size` against `tgi_batch_current_size`: a long queue beside a small batch points at admission or at a memory budget that will not let the batch grow. 2. **Inference duration dominates.** The model step itself is slow. Look at prompt lengths (prefill cost scales with them), at precision and kernel path, at whether the batch has grown so large that per-request decode has slowed, and at whether output lengths are longer than you assumed. 3. **Both look fine but users complain.** Check `tgi_request_mean_time_per_token_duration`. Frequent admission pauses show up here as inter-token jitter even when averages look healthy. ## Overload behaviour TGI's router enforces `--max-concurrent-requests`. Past that limit it does not silently queue forever - it rejects with HTTP **429** and an error type of `overloaded`. This matters for two reasons. First, it is the correct load-shedding contract: a fast rejection lets a caller retry elsewhere or degrade, where an unbounded queue converts overload into timeouts everywhere. Second, clients must actually handle it - retry with backoff and jitter, and treat a sustained 429 rate as a scaling signal rather than a bug. An overload rejection is distinct from a validation rejection: a prompt that exceeds the configured input length fails validation with a 4xx that no amount of retrying will fix, and it never reaches the queue at all. ## Choosing the autoscaling signal When you wire these into a scaling policy, the useful signals are queue depth and time-to-first-token, because both track saturation in the direction users feel. GPU utilisation does not. Whatever mechanism you scale with, the constraint that dominates an LLM deployment is that adding a replica means loading tens of gigabytes of weights, so the reaction is measured in minutes - scale on a leading indicator, not a lagging one. ## Streaming complicates error handling One practical wrinkle worth carrying into an interview: once a streamed response has started, the HTTP status is already 200. A failure part-way through therefore cannot come back as a status code - it arrives as an event in the stream carrying an error payload. Clients that only check the response status will silently treat a truncated generation as a complete one. Metrics show these as failures; naive clients do not. ## What interviewers listen for Name real series rather than saying 'we'd monitor latency'. Then show the split - queue versus inference - and say what each direction implies. Finishing with the 429 contract and the observation that GPU utilisation is a useless saturation signal here is what separates someone who has run one of these from someone who has read the README.
- A TGI server starts returning HTTP 429 with an overloaded error type. What should the client do, and what should you do?The client should back off with jitter and retry, or shed to a fallback - a 429 here means the router hit its concurrency limit, not that the request was malformed. Operationally, a sustained 429 rate is your capacity signal: add replicas, or raise the concurrency limit only if queue duration shows there is genuinely slack. Raising the limit without slack just converts fast rejections into slow timeouts.
- Your client streams responses and sees no HTTP errors, yet some answers are truncated. Where is the failure hiding?In the stream body. Once streaming begins the response status is already 200, so a mid-generation failure arrives as an event carrying an error payload rather than as a status code. A client that checks only the status treats the truncation as a normal short answer. Inspect every event for an error field, and reconcile against your own token counts.
- Why is GPU utilisation a poor autoscaling signal for a TGI deployment?Because a single decode step keeps the device nominally busy, so utilisation reads near 100% whether the server is comfortable or drowning. It has almost no dynamic range in the region you care about. Queue depth and time-to-first-token move monotonically with saturation and map onto what users feel, which makes them the right inputs to a scaling policy.
saying these in an interview costs you the question
- Autoscaling an LLM server on GPU utilisation
- Only tracking end-to-end latency, never queue time
- Treating 429 overloaded as a client bug
- Assuming a streamed 200 response cannot fail
- Raising the concurrency limit to make 429s go away