skip to content

In TGI, what do --max-input-tokens and --max-total-tokens each cap?

level: middleimportance: should knowfreq 52%

answer

  1. one bounds the prompt, one bounds the whole slot
  2. the difference is your generation budget
  3. 422, not a shortened prompt
  4. unset means derived at warmup
  5. longer slots mean fewer concurrent ones

basics

~20 s

--max-input-tokens caps the prompt length TGI will accept for one request; --max-total-tokens caps prompt plus generated tokens together. The router validates both before queueing and rejects an oversized request with a 422 validation error rather than truncating it.

solid answer

~50 s

They are two limits on the same request. `--max-input-tokens` bounds the tokenized prompt; `--max-total-tokens` bounds the whole conversation slot — input plus everything the model generates — so the effective generation budget for a request is `max_total_tokens - input_length`, and `max_input_tokens` must be strictly less than `max_total_tokens`. Enforcement is at the **router**, before the request reaches a shard. TGI does not silently truncate a long prompt; it returns a 422 with a validation error naming the limit. That is a deliberate design choice — a truncated prompt would produce a plausible wrong answer, while a rejection makes the client's mistake visible. In TGI 3.x these values are inferred from the model config and the memory the server can actually allocate when you do not set them, so read `/info` to see what the server settled on. Setting `--max-total-tokens` explicitly is how you trade per-request context against how many requests fit concurrently. `--max-input-tokens` replaced the older `--max-input-length` spelling.

code

bash · 6 lines
bash
docker run --gpus all --shm-size 1g -p 8080:80 \
  -v "$PWD/tgi-data:/data" \
  ghcr.io/huggingface/text-generation-inference:3.3.5 \
  --model-id HuggingFaceH4/zephyr-7b-beta \
  --max-input-tokens 3072 \
  --max-total-tokens 4096

go deeper

for a junior

Know that one flag limits the prompt and the other limits prompt plus generated output together, and that an over-long request is rejected with an error rather than quietly shortened.

for a middle

Explain that the router validates before queueing and returns a 422 validation error, that the generation budget is total minus actual input, and that unset values are derived at warmup so /info is where you read the truth.

for a senior

Demonstrate sizing from measured traffic: p99 prompt length plus the answer length the product needs, with the concurrency cost of a longer slot made explicit, and the limits surfaced to callers rather than discovered by bisection.

for a principal

Own the limit as a product contract across the fleet: one advertised context budget per deployment tier, chosen so every GPU SKU in the pool can honour it, with changes treated as versioned API changes rather than a flag someone bumps.

## Two limits, one request Every request TGI serves occupies a slot sized in tokens. Two launcher flags bound that slot: - **`--max-input-tokens`** — the maximum length of the tokenized prompt. - **`--max-total-tokens`** — the maximum of prompt plus generated tokens combined. The relationship is arithmetic, not independent: for a given request, the most tokens it may generate is `max_total_tokens` minus its actual input length. A 4,000-token prompt against `--max-total-tokens 4096` leaves 96 tokens of output, regardless of what `max_new_tokens` the caller asked for. The launcher requires `max_input_tokens` to be strictly less than `max_total_tokens` — a configuration where they are equal would admit prompts with zero room to answer, so it is refused at startup rather than at 3am. Both are also bounded above by the model itself. You cannot configure a slot longer than the checkpoint's positional limit; asking for more is a startup failure, not a silent clamp. ## Rejection, not truncation The router — the component in front of the shards — tokenizes and validates each request before it is queued. An over-length request comes back as **422 Unprocessable Entity** with a JSON body carrying an `error` message and `error_type` of `validation`. It never occupies GPU time. This is worth defending in an interview, because the naive expectation is truncation. Truncating a prompt is the worse behaviour: the model answers confidently from a mutilated context and the caller has no signal that anything was dropped. An explicit 422 pushes the decision back to the client, which is the only layer that knows whether to summarize, drop old turns, or split the work. Application code in front of TGI should therefore count tokens with the same tokenizer and manage its own context budget, treating 422 as a bug in its budgeting rather than a condition to retry. Note the distinction from the overload path: a request that is *valid* but arrives when the server has no capacity is a queueing/backpressure concern and surfaces differently (TGI signals an overloaded server with 429). Validation is about the shape of the request; 429 is about the state of the server. ## What happens when you leave them unset TGI 3.x does not force you to specify these. When unset, the launcher derives them: it reads the model's configuration for the architectural ceiling and, during the warmup phase, measures how much memory it can actually claim on the device, then settles on values that fit. This is convenient and it is also a trap for reproducibility — the same image and the same model id can settle on different limits on an 80 GB card than on a 24 GB one, or when another process is already holding memory on the same GPU. So: read `/info` after every deploy. It reports the limits the server actually chose. "It works on the A100 and rejects requests on the L4" is nearly always this, and it is invisible unless you look. ## Setting them deliberately The reason to pin these is that they are a capacity dial. Every admitted request reserves cache space proportional to its slot length, so raising `--max-total-tokens` makes each request more expensive and reduces how many can run concurrently on the same card. Lowering it packs more conversations in. The right value comes from your traffic, not from the model's maximum: - Measure the real distribution of prompt lengths — the p95 and the tail, not the mean. - Set `--max-input-tokens` just above the p99 you intend to support, so genuinely oversized requests fail fast and loudly. - Set `--max-total-tokens` to that plus the longest answer the product actually needs. A team that sets `--max-total-tokens` to the model's full advertised context "because we paid for it" and then serves 500-token chats has quietly bought a much smaller concurrency ceiling for no user benefit. Conversely, a team that sets it too tight discovers it through a stream of 422s from exactly the users with the longest, most valuable documents. ## Making the limits visible to callers Because the limits are server-side and per-deployment, they belong in whatever contract your clients read. Surface the configured input and total limits in your API docs or a capabilities endpoint of your own, and give the 422 a client-facing message that names the number. A caller that has to discover the limit by bisection is a caller that will hard-code the wrong one. ## Changing them These are launcher flags, so changing them means restarting the server: there is no runtime knob. On a rolling deploy that is a full model reload per replica, with the cold-start cost that implies. Plan limit changes as deployments, and validate the new value on the same GPU SKU you run in production — a limit that fits on one card may not fit on another.

  • A caller sends a 5,000-token prompt to a server configured with --max-input-tokens 3072. What comes back?
    A 422 with a JSON body whose `error_type` is `validation` and whose message names the limit that was exceeded. The request is rejected at the router before any GPU work happens, so it costs nothing but a tokenization pass. Nothing is truncated — the caller is expected to shorten or split the prompt, because only the caller knows which part of its context is safe to drop.
  • You did not set either flag. How do you find out what limits the server is enforcing?
    Query `/info`, which reports the values the server settled on along with the model id, revision and dtype. This matters because TGI derives the limits from the model config and the memory it can actually allocate at warmup, so the same image and model can land on different limits across GPU SKUs — or on the same SKU when another process already holds memory. Checking `/info` after a deploy turns that into a visible fact.
  • Why would you deliberately configure --max-total-tokens well below the model's maximum context?
    Because each admitted request reserves cache space proportional to its slot, so a long slot cuts how many requests run concurrently on the same card. If your traffic is short chat turns, a large slot buys nothing for users while lowering the concurrency ceiling and raising cost per request. Size the slot from the observed prompt-length distribution plus the longest answer the product needs, not from the model's advertised context.
  • Can you change these limits without a restart?
    No — they are launcher flags read at process start, so a change is a redeploy and a full model reload per replica, with the cold-start cost that carries. That is an argument for choosing them from measured traffic rather than tuning them reactively, and for validating a new value on the same GPU SKU production runs, since a slot that fits on one card may not fit on another.

saying these in an interview costs you the question

  • Expecting TGI to truncate a long prompt
  • Thinking max-total-tokens is only the output limit
  • Setting both flags to the same value
  • Treating a 422 validation error as a retryable overload
  • Assuming unset means the model's full context is available

context