skip to content

Batching and Performance Features

You will learn the TGI knobs that shape throughput and latency, plus the built-ins teams actually use — streaming, grammar-constrained JSON, and speculation. Interviewers use TGI-vs-vLLM comparisons to see whether you evaluate engines on mechanics rather than popularity.

on this pageshow

questions

6

In TGI, what does --max-batch-prefill-tokens control, and what breaks if it is too high?

level: middleimportance: must knowfreq 58%

answer

  1. Governs one phase, not the whole request
  2. Prompt tokens per forward pass
  3. Activation peak, not KV cache
  4. Too high means CUDA OOM, not slowness
  5. Derived from --max-input-tokens if unset

basics

~20 s

TGI's --max-batch-prefill-tokens caps how many prompt tokens the router packs into one prefill step. Raising it improves prompt-processing throughput but raises peak activation memory, so setting it too high produces CUDA out-of-memory crashes under concurrent long prompts.

solid answer

~50 s

TGI's router runs two kinds of work: **prefill**, which processes whole prompts, and **decode**, which emits one token per sequence per step. `--max-batch-prefill-tokens` is the token budget for a single prefill step - the router will not group more prompt tokens than that into one forward pass. Left unset, TGI derives it from `--max-input-tokens` plus a small margin, so the longest prompt you allow can still be covered. It is fundamentally a *memory* knob, not just a speed knob. Prefill attends over the whole prompt at once, so activation memory scales with the tokens in that batch, and that allocation is separate from the KV cache. Raising it lets more prompts prefill together (fewer prefill rounds, better throughput); lowering it protects headroom but makes prompts queue longer. Don't confuse it with `--max-batch-total-tokens`, which bounds prompt plus generated tokens across every request in the running batch and is inferred automatically on flash/paged-attention models.

code

bash · 6 lines
bash
docker run --gpus all --shm-size 1g -p 8080:80 \
  ghcr.io/huggingface/text-generation-inference:3.3.5 \
  --model-id meta-llama/Llama-3.1-8B-Instruct \
  --max-input-tokens 4096 \
  --max-total-tokens 8192 \
  --max-batch-prefill-tokens 8192

go deeper

for a junior

Know that TGI processes the prompt (prefill) and then generates tokens one at a time (decode), and that this flag is about the prompt phase. Being able to name the phase is already most of the answer at this level.

for a middle

Explain that the flag caps prompt tokens per prefill forward pass, that the cost of raising it is peak activation memory, and that it is distinct from --max-input-tokens and --max-batch-total-tokens.

for a senior

Show how you would pick a value from a real prompt-length distribution and concurrency target, and explain why the failure mode under a traffic mix is an OOM crash rather than gradual slowdown.

for a principal

Frame it as part of a single VRAM budget: prefill headroom, KV cache capacity and weights compete for one pool, so an aggressive prefill budget silently reduces how many concurrent sessions a replica can hold, which is a cost-per-token decision.

## Two phases, two budgets Every request in TGI goes through two distinct phases. **Prefill** runs one forward pass over the entire prompt, producing the key/value tensors that land in the cache plus the first output token. **Decode** then emits one token per sequence per step, reading the cache instead of recomputing it. Prefill is compute-heavy and bursty; decode is memory-bandwidth-heavy and steady. Because the two phases stress different resources, TGI's launcher gives them different budgets: - `--max-input-tokens` - the longest prompt a *single* request may send. Exceed it and the router rejects the request during validation, before any GPU work happens. - `--max-batch-prefill-tokens` - the total prompt tokens the router will pack into *one* prefill step across all requests it batches together. - `--max-batch-total-tokens` - the total tokens (prompt + generated) that all requests in the running batch may occupy at once. On flash/paged-attention models TGI profiles free memory at startup and infers this value, logging the number it chose. - `--max-total-tokens` - the per-request ceiling on prompt + generation. A common interview stumble is treating the last three as the same thing. They are not: one is per prefill step, one is per running batch, one is per request. ## Why it is a memory knob During prefill, attention is computed over the full prompt length simultaneously. The intermediate activations for that pass - projections, attention outputs, MLP intermediates - are proportional to the number of tokens in the batch multiplied by hidden size and layer count. Those tensors are transient, but they must all be resident at once, on top of the model weights and the KV cache. That is why a server that has been happily serving 512-token prompts can OOM the moment several 8k-token prompts arrive together. Nothing about the weights changed and the KV cache may have plenty of free blocks; the crash comes from the prefill activation peak. `--max-batch-prefill-tokens` is the throttle on that peak. It puts a hard ceiling on how large that transient allocation can get, no matter how many long prompts happen to arrive in the same instant. ## Choosing a value Start from your traffic. If your p99 prompt is 2k tokens and you want roughly four prompts to prefill together, a budget around 8k is a reasonable starting point. Then measure: load the server with realistic prompt lengths at your target concurrency and watch memory headroom and time-to-first-token together. The tradeoff runs in both directions: - **Too low**: prefill happens in many small steps. Each step has fixed overhead, so GPU utilisation during prompt processing drops and prompts spend longer queueing before their first token. On prompt-heavy workloads (RAG with long retrieved context, document summarisation) this shows up directly as poor TTFT under load. - **Too high**: you have effectively promised the GPU it can allocate a very large transient buffer. The failure is not graceful degradation, it is a crash - and because prefill batching depends on arrival patterns, it is the kind of crash that only appears under production traffic mixes, not in a single-request smoke test. The value also interacts with the automatically-inferred `--max-batch-total-tokens`: memory you reserve for large prefill peaks is memory not available for KV cache, so an aggressive prefill budget quietly shrinks how many concurrent sequences the server can hold. ## Newer TGI and prefill chunking Recent TGI releases split a long prompt's prefill across several steps rather than insisting the whole prompt fit one pass. This softens the interaction between prompt length and this budget, and it shortens the pause that a big prefill imposes on requests already decoding. It does not remove the knob: the per-step token budget is still what bounds the activation peak. ## What to say in an interview Name the phase it governs (prefill, not decode), name the resource it protects (transient activation memory, not KV cache), and describe the tradeoff as throughput-versus-headroom. Then say how you would set it - measure at realistic prompt lengths and concurrency rather than copying a number from a blog post. Mentioning that a too-high value fails as an OOM crash rather than as slowness is the detail that signals you have actually operated one of these servers.

  • Why does TGI infer --max-batch-total-tokens for you instead of asking you to set it?
    On flash/paged-attention models TGI can measure how much memory is actually left for the KV cache after weights and prefill headroom, then convert that into a token count. Hand-picking the number means guessing at the same measurement, and guessing high means the server accepts more concurrent sequences than the cache can hold. If you do override it, you are asserting you know the cache capacity better than the profiler did.
  • A request with a 6000-token prompt is rejected before it ever reaches the GPU. Which flag is responsible?
    `--max-input-tokens`, which the router enforces during request validation. That is a per-request ceiling and it fails fast with a validation error rather than a crash. `--max-batch-prefill-tokens` is a batching budget: it never rejects a request, it only limits how many prompts prefill together in one step.
  • Your server OOMs only during traffic spikes, never in load tests with fixed 512-token prompts. What do you look at first?
    The prompt-length distribution during the spike. Fixed short prompts never build a large prefill batch, so the activation peak stays small; a spike that mixes in long prompts can put several thousand prompt tokens into one prefill step. Lower `--max-batch-prefill-tokens` to cap the peak, then re-run the load test with the real length distribution rather than a constant.

saying these in an interview costs you the question

  • Calling it a KV cache size limit
  • Confusing it with the per-request --max-total-tokens
  • Assuming a higher value only costs latency, never memory
  • Setting it to a huge number to 'maximise throughput'
  • Believing it rejects long prompts

context

open as a page

How does TGI's grammar parameter force a response to match a JSON schema?

level: middleimportance: should knowfreq 45%

basics

~20 s

TGI accepts a grammar object in a /generate request's parameters, either type json with a JSON Schema or type regex with a pattern. At each decoding step the server masks out any token that could not continue a valid match, so the output is guaranteed to parse.

open as a page

Which TGI /metrics series separate queueing time from GPU inference time?

level: seniorimportance: should knowfreq 40%

basics

~20 s

TGI exposes Prometheus metrics on /metrics, where tgi_request_queue_duration records time spent waiting before the GPU touched a request and tgi_request_inference_duration records the model work itself. Splitting slow end-to-end latency across those two is the first triage step.

open as a page

What does TGI's --speculate flag enable, and when does it stop paying off?

level: seniorimportance: should knowfreq 36%

basics

~20 s

TGI's --speculate sets how many extra tokens are proposed per step and verified in one forward pass - from Medusa heads when the checkpoint has them, otherwise by n-gram lookup in the existing context. It helps a lightly loaded server and fades as concurrency fills the GPU.

open as a page

How does TGI's --waiting-served-ratio decide when to pause decoding for queued requests?

level: seniorimportance: should knowfreq 34%

basics

~20 s

TGI's --waiting-served-ratio is the ratio of waiting requests to running requests at which the router will interrupt decoding, prefill the queued requests, and fold them into the running batch. A lower value admits newcomers sooner and protects their time-to-first-token.

open as a page

When would you choose TGI over vLLM for a production LLM endpoint?

level: principalimportance: should knowfreq 44%

basics

~20 s

Choose on mechanics and fit, not popularity: TGI is the low-friction option for teams already on the Hugging Face stack, with built-in grammar guidance and a Rust router. vLLM moves faster on new architectures and quantization formats. Both do continuous batching and paged KV cache.

open as a page