skip to content

In TGI, what does --max-batch-prefill-tokens control, and what breaks if it is too high?

level: middleimportance: must knowfreq 58%

answer

  1. Governs one phase, not the whole request
  2. Prompt tokens per forward pass
  3. Activation peak, not KV cache
  4. Too high means CUDA OOM, not slowness
  5. Derived from --max-input-tokens if unset

basics

~20 s

TGI's --max-batch-prefill-tokens caps how many prompt tokens the router packs into one prefill step. Raising it improves prompt-processing throughput but raises peak activation memory, so setting it too high produces CUDA out-of-memory crashes under concurrent long prompts.

solid answer

~50 s

TGI's router runs two kinds of work: **prefill**, which processes whole prompts, and **decode**, which emits one token per sequence per step. `--max-batch-prefill-tokens` is the token budget for a single prefill step - the router will not group more prompt tokens than that into one forward pass. Left unset, TGI derives it from `--max-input-tokens` plus a small margin, so the longest prompt you allow can still be covered. It is fundamentally a *memory* knob, not just a speed knob. Prefill attends over the whole prompt at once, so activation memory scales with the tokens in that batch, and that allocation is separate from the KV cache. Raising it lets more prompts prefill together (fewer prefill rounds, better throughput); lowering it protects headroom but makes prompts queue longer. Don't confuse it with `--max-batch-total-tokens`, which bounds prompt plus generated tokens across every request in the running batch and is inferred automatically on flash/paged-attention models.

code

bash · 6 lines
bash
docker run --gpus all --shm-size 1g -p 8080:80 \
  ghcr.io/huggingface/text-generation-inference:3.3.5 \
  --model-id meta-llama/Llama-3.1-8B-Instruct \
  --max-input-tokens 4096 \
  --max-total-tokens 8192 \
  --max-batch-prefill-tokens 8192

go deeper

for a junior

Know that TGI processes the prompt (prefill) and then generates tokens one at a time (decode), and that this flag is about the prompt phase. Being able to name the phase is already most of the answer at this level.

for a middle

Explain that the flag caps prompt tokens per prefill forward pass, that the cost of raising it is peak activation memory, and that it is distinct from --max-input-tokens and --max-batch-total-tokens.

for a senior

Show how you would pick a value from a real prompt-length distribution and concurrency target, and explain why the failure mode under a traffic mix is an OOM crash rather than gradual slowdown.

for a principal

Frame it as part of a single VRAM budget: prefill headroom, KV cache capacity and weights compete for one pool, so an aggressive prefill budget silently reduces how many concurrent sessions a replica can hold, which is a cost-per-token decision.

## Two phases, two budgets Every request in TGI goes through two distinct phases. **Prefill** runs one forward pass over the entire prompt, producing the key/value tensors that land in the cache plus the first output token. **Decode** then emits one token per sequence per step, reading the cache instead of recomputing it. Prefill is compute-heavy and bursty; decode is memory-bandwidth-heavy and steady. Because the two phases stress different resources, TGI's launcher gives them different budgets: - `--max-input-tokens` - the longest prompt a *single* request may send. Exceed it and the router rejects the request during validation, before any GPU work happens. - `--max-batch-prefill-tokens` - the total prompt tokens the router will pack into *one* prefill step across all requests it batches together. - `--max-batch-total-tokens` - the total tokens (prompt + generated) that all requests in the running batch may occupy at once. On flash/paged-attention models TGI profiles free memory at startup and infers this value, logging the number it chose. - `--max-total-tokens` - the per-request ceiling on prompt + generation. A common interview stumble is treating the last three as the same thing. They are not: one is per prefill step, one is per running batch, one is per request. ## Why it is a memory knob During prefill, attention is computed over the full prompt length simultaneously. The intermediate activations for that pass - projections, attention outputs, MLP intermediates - are proportional to the number of tokens in the batch multiplied by hidden size and layer count. Those tensors are transient, but they must all be resident at once, on top of the model weights and the KV cache. That is why a server that has been happily serving 512-token prompts can OOM the moment several 8k-token prompts arrive together. Nothing about the weights changed and the KV cache may have plenty of free blocks; the crash comes from the prefill activation peak. `--max-batch-prefill-tokens` is the throttle on that peak. It puts a hard ceiling on how large that transient allocation can get, no matter how many long prompts happen to arrive in the same instant. ## Choosing a value Start from your traffic. If your p99 prompt is 2k tokens and you want roughly four prompts to prefill together, a budget around 8k is a reasonable starting point. Then measure: load the server with realistic prompt lengths at your target concurrency and watch memory headroom and time-to-first-token together. The tradeoff runs in both directions: - **Too low**: prefill happens in many small steps. Each step has fixed overhead, so GPU utilisation during prompt processing drops and prompts spend longer queueing before their first token. On prompt-heavy workloads (RAG with long retrieved context, document summarisation) this shows up directly as poor TTFT under load. - **Too high**: you have effectively promised the GPU it can allocate a very large transient buffer. The failure is not graceful degradation, it is a crash - and because prefill batching depends on arrival patterns, it is the kind of crash that only appears under production traffic mixes, not in a single-request smoke test. The value also interacts with the automatically-inferred `--max-batch-total-tokens`: memory you reserve for large prefill peaks is memory not available for KV cache, so an aggressive prefill budget quietly shrinks how many concurrent sequences the server can hold. ## Newer TGI and prefill chunking Recent TGI releases split a long prompt's prefill across several steps rather than insisting the whole prompt fit one pass. This softens the interaction between prompt length and this budget, and it shortens the pause that a big prefill imposes on requests already decoding. It does not remove the knob: the per-step token budget is still what bounds the activation peak. ## What to say in an interview Name the phase it governs (prefill, not decode), name the resource it protects (transient activation memory, not KV cache), and describe the tradeoff as throughput-versus-headroom. Then say how you would set it - measure at realistic prompt lengths and concurrency rather than copying a number from a blog post. Mentioning that a too-high value fails as an OOM crash rather than as slowness is the detail that signals you have actually operated one of these servers.

  • Why does TGI infer --max-batch-total-tokens for you instead of asking you to set it?
    On flash/paged-attention models TGI can measure how much memory is actually left for the KV cache after weights and prefill headroom, then convert that into a token count. Hand-picking the number means guessing at the same measurement, and guessing high means the server accepts more concurrent sequences than the cache can hold. If you do override it, you are asserting you know the cache capacity better than the profiler did.
  • A request with a 6000-token prompt is rejected before it ever reaches the GPU. Which flag is responsible?
    `--max-input-tokens`, which the router enforces during request validation. That is a per-request ceiling and it fails fast with a validation error rather than a crash. `--max-batch-prefill-tokens` is a batching budget: it never rejects a request, it only limits how many prompts prefill together in one step.
  • Your server OOMs only during traffic spikes, never in load tests with fixed 512-token prompts. What do you look at first?
    The prompt-length distribution during the spike. Fixed short prompts never build a large prefill batch, so the activation peak stays small; a spike that mixes in long prompts can put several thousand prompt tokens into one prefill step. Lower `--max-batch-prefill-tokens` to cap the peak, then re-run the load test with the real length distribution rather than a constant.

saying these in an interview costs you the question

  • Calling it a KV cache size limit
  • Confusing it with the per-request --max-total-tokens
  • Assuming a higher value only costs latency, never memory
  • Setting it to a huge number to 'maximise throughput'
  • Believing it rejects long prompts

context