Your TensorRT-LLM engine rejects a 32k-token request; max_seq_len was 8192. Now what?
answer
- shape ranges live in the compiled profiles
- serving config can only clamp downward
- headroom costs cache and activation memory
- size from the length distribution
- rebuild should be a pipeline run
basics
~20 sRebuild. max_input_len, max_seq_len, max_batch_size and max_num_tokens are compile-time ceilings serialized into the engine's optimization profiles; the serving layer can only configure values at or below them. No runtime setting raises a limit the engine was not built for.
solid answer
~50 sThose limits are declared to `trtllm-build` and baked into the engine's optimization profiles, so the only fix is a new build with a larger `--max_seq_len` (and `--max_input_len`, if long prompts rather than long outputs are the problem). Triton's `triton_max_batch_size` and `max_num_tokens` can be set *below* what the engine supports, never above — the serving layer clamps down, it cannot extend. The real question an interviewer is after is how you choose the ceilings in the first place. Headroom is not free: a longer maximum sequence enlarges the KV-cache footprint per sequence and the activation memory a step must reserve, so an engine built for 128k when your traffic is 4k buys fewer concurrent sequences on the same card. The practical approach is to size from the actual prompt/output length distribution with a margin, keep the build in CI so a rebuild is a pipeline run rather than an incident, and use `--multiple_profiles` when one engine must serve genuinely varied shapes.
go deeper
Remember that the maximum batch size and sequence length are chosen when the engine is built, so a request beyond them needs a new engine rather than a config change.
Explain optimization profiles as the reason the limit is structural, and state that serving-layer values can only be set at or below the compiled ceilings.
Show that you size the ceilings from the real prompt and output length distribution, know that headroom costs KV cache and activation memory, and have a rebuild pipeline plus a defined behaviour for the requests that fail meanwhile.
Decide the fleet shape: one engine sized for the tail, or a common-case fleet plus a long-context fleet with routing. Own that tradeoff in cost per token and in how many artifacts the org has to build and validate.
## Why the limit is hard A TensorRT engine contains optimization profiles: declared min/opt/max shape ranges that the builder used to pick kernels and to plan memory. A request outside the compiled range has no valid execution path, so it is rejected rather than handled slowly. This is the sharpest behavioural difference between a compiled engine and a Python-level server, where the same limit is usually a configuration value the process reads at startup. The four ceilings that bite: - **`--max_input_len`** — longest prompt admitted. - **`--max_seq_len`** — longest prompt-plus-generation. - **`--max_batch_size`** — most concurrent sequences. - **`--max_num_tokens`** — most tokens processed in one forward pass, which bounds how much prefill can be packed into an iteration. ## Diagnosing which one you hit The error text points at a shape, but the mapping is worth knowing. A long prompt that fails immediately is `max_input_len` or `max_num_tokens`. A request that starts generating and then stops short is bumping `max_seq_len` — the prompt fit, the sum did not. Concurrency that plateaus below what your GPU memory should allow is `max_batch_size`, or the serving-layer value you configured under it. Rejections that arrive under load but not in isolation are usually queue bounds, not engine ceilings, and no rebuild will help those. ## Why not just build everything huge The temptation is to build with the largest numbers the model supports and never think again. Three costs argue against it. **KV-cache math.** The bytes a sequence can consume scale with its maximum length. The runtime reserves and plans against the compiled maximum, so an engine built for very long contexts leaves less room for concurrency at the same VRAM budget. If your median request is 2k tokens, building for 128k means paying for a tail you rarely serve. **Activation memory.** A larger per-pass token budget means larger intermediate tensors, which the builder must plan workspace for. That workspace is memory not available as cache. **Tactic selection.** The builder tunes against the shapes you declared. A profile spanning an enormous range gives it a harder optimization problem than a narrow one, and the kernels it settles on may be less good for the shapes you actually serve. `--multiple_profiles` exists for exactly this tension: build several profiles so different shape regimes get separately tuned kernels, at the cost of a longer build and a bigger engine. ## Sizing from data, not from the model card The defensible method is to take the request-length distribution from your logs — prompt tokens and output tokens separately — and build for a high percentile plus margin, not for the model's theoretical maximum context. Then decide explicitly what happens to the tail: reject it, route it to a second deployment built with longer limits, or truncate it upstream. "Route long requests to a long-context deployment" is a perfectly good production answer, and it is often cheaper than making every replica carry long-context overhead. ## Make the rebuild boring The reason this question tests seniority is that the correct answer — "rebuild" — is only acceptable if rebuilding is routine. That means the build lives in CI, parameterised by the ceilings; the artifact is versioned with those ceilings in its key; accuracy validation runs against the new engine; and deploying it is a normal rollout with the old artifact still available. Teams that build by hand on a login node experience a limit change as an outage; teams with a pipeline experience it as a ticket. ## What to say about the request that failed There is also an immediate-response half to this answer. While the rebuild is in flight, you decide what the client sees: a clear error naming the limit, truncation at a documented boundary, or a redirect to a deployment that can serve it. Silently truncating a user's 32k prompt to 8k and answering from the first quarter of it is the worst of the three, because it produces a confidently wrong answer instead of an honest failure. ## The interview shape Weak answers reach for a runtime knob that does not exist. Adequate answers say "rebuild". Strong answers say rebuild, explain why the ceiling is structural, explain why building for the maximum is not free, and describe how the ceilings were chosen from traffic data in the first place — plus what the client sees in the meantime.
- Traffic is mostly 2k tokens but 3% of requests are 100k. How do you serve both?Two deployments. Build the main fleet for the common regime so it keeps high concurrency per GPU, and build a smaller long-context fleet with the larger ceilings, accepting its lower concurrency. Route by measured prompt length at the gateway. Making every replica carry 100k-capable overhead to serve 3% of traffic is the expensive alternative.
- Requests start failing under load but the same request succeeds when the server is idle. Is that a build ceiling?No — engine ceilings are deterministic per request, not load-dependent. Load-dependent rejections come from the serving layer: a bounded queue rejecting when full, a queue-delay limit expiring, or KV-cache pressure with a scheduler policy that refuses to admit work it cannot finish. Rebuilding fixes none of those; the fix is queue sizing, cache fraction or capacity.
- What does --multiple_profiles buy, and what does it cost?It builds several optimization profiles so different shape regimes get separately tuned kernels, which helps when one engine must serve both short interactive prompts and long documents. The cost is a longer build and a larger engine, plus more surface to validate — so it is worth it when the shape spread is genuinely wide, not as a default.
saying these in an interview costs you the question
- Looks for a runtime flag to raise the sequence limit
- Sets the Triton batch size above the engine's built ceiling
- Builds for maximum context by default, ignoring the cache cost
- Blames engine ceilings for load-dependent rejections
- Silently truncates oversized prompts instead of failing clearly