How do you tune an LLM server's per-iteration token budget and concurrent-sequence cap?
answer
- two budgets, tokens and sequences
- step duration is everyone's token latency
- the cap is not memory
- measure with real length distributions
- change one budget at a time
basics
~20 sSize the per-iteration token budget by how long a step may take, since step duration is every user's inter-token latency, and size the concurrent-sequence cap by how much KV cache you actually have. Then measure at your target load rather than guessing.
solid answer
~50 sTwo scheduler budgets bound a continuously batched server. The **per-iteration token budget** — `max_num_batched_tokens` in vLLM — caps how many token positions go into one forward pass; with chunked prefill on it also sets the prefill chunk size. Raising it ingests prompts faster and improves first-token latency, but lengthens each step, and step duration *is* the inter-token latency every running sequence sees. The **concurrent-sequence cap** — `max_num_seqs` in vLLM, roughly `--max-concurrent-requests` in TGI — bounds how many sequences may decode at once. Raising it buys throughput only while KV-cache memory can actually hold that many sequences at your real context lengths; past that the true limiter is memory, and the scheduler starts refusing admission or evicting running work. The method is: pick the SLO you are defending, set the budget that governs it, then run a load sweep at production prompt and output length distributions and read the curve.
code
bash · 3 linesvllm serve meta-llama/Llama-3.1-8B-Instruct \
--max-num-batched-tokens 8192 \
--max-num-seqs 64go deeper
Know that a serving engine has configurable limits on how many requests run at once and how many tokens go into one step, and that raising them is not free.
Explain what each budget governs: the token budget sets step size and therefore prefill chunk size, the sequence cap sets admission concurrency, and KV-cache memory is what actually bounds the second.
Show a method rather than values — pick the SLO, sweep load with production length distributions, change one budget at a time, and recognize the point where the answer is capacity rather than configuration.
Own the framing that these budgets encode a product decision about whose latency matters, and be ready to argue for separate replica pools with different budgets instead of one compromise setting for mixed workloads.
## Two budgets, two different currencies A continuously batched engine rebuilds its batch every iteration under two independent limits, and they are denominated differently — one in tokens per step, one in concurrent sequences. Confusing them is the most common tuning mistake. The **per-iteration token budget** answers "how many token positions may this forward pass contain?" During pure decode, each running sequence contributes exactly one position, so the budget is barely binding. It becomes the governing number as soon as prefill is in play: with chunked prefill enabled, the budget minus the decode tokens is precisely the size of the prefill slice the scheduler may take this step. vLLM spells it `max_num_batched_tokens`. The **concurrent-sequence cap** answers "how many sequences may be in the running set at once?" vLLM spells it `max_num_seqs`; TGI bounds admitted work with `--max-concurrent-requests`. It is an admission control number, not a memory number — which is the source of most of the confusion around it. ## What the token budget really controls Step duration. That is the whole story, and it is worth saying out loud in an interview because it converts a config value into a user-visible metric. Every running sequence emits one token per iteration, so the wall-clock time of an iteration is the inter-token latency of every stream on the server. A large budget means fatter steps: more prefill work absorbed per step, better prompt ingestion, better time-to-first-token — and a longer gap between output tokens for everyone currently generating. A small budget means finer interleaving, smoother streams, more scheduler overhead per token, and slower prompt ingestion. So you do not tune it toward a number you read in a blog post; you tune it against whichever SLO you wrote down. If the contract is "first token within 800 ms", bias the budget up. If the contract is "a stream that never stalls", bias it down. If your prompts are short (a few hundred tokens), the budget rarely binds and tuning it is mostly wasted effort. ## What the sequence cap really controls Not much, on its own — and this is the part candidates get wrong. Raising `max_num_seqs` from 64 to 256 does not create memory. Each admitted sequence needs KV-cache blocks proportional to its current length, and the cache is a fixed pool carved out of GPU memory at startup. If the pool holds, say, 90k tokens of cache in total, then 256 sequences averaging 4k tokens each simply cannot coexist; the scheduler will admit until the pool is full and then stop, or evict already-running sequences to make room. The configured cap is an upper bound; the memory pool is the real bound. The cap is nevertheless useful in two directions. Set *too low*, it artificially throttles a server that had cache to spare, and you will see the running-sequence count pinned at the cap with cache utilization comfortable — that is free throughput being left on the table. Set *too high*, it invites the scheduler into a regime where it repeatedly admits work it cannot sustain and then evicts it, converting steady throughput into churn. Deliberately capping it *below* what memory allows is also a legitimate latency tactic: fewer concurrent sequences means shorter steps and better per-user token rates, at lower aggregate throughput. ## The interaction you must be able to narrate The two budgets meet inside a single step. Suppose 60 sequences are decoding and the budget is 2048. Sixty tokens go to decode; roughly 1988 remain for a prefill slice. Now raise the sequence cap to 400 and let it fill: 400 decode tokens leave only ~1648 for prefill, so prompt ingestion slows even though you never touched the token budget. Concurrency and prefill throughput compete for the same per-step allowance. Under long-context workloads, where each decode step also drags a much larger KV cache through memory, the same nominal concurrency costs more time per step than it does with short contexts. ## Method, not magic numbers Start from defaults; they are chosen to be reasonable. Then run a load sweep with your *own* prompt-length and output-length distributions — synthetic uniform-length traffic will lie to you, because the entire point of continuous batching is behaviour under heterogeneous lengths. Ramp concurrency, and plot first-token latency and per-token latency against offered load. You are looking for the load at which your latency SLO breaks; that is your capacity per replica, and it is the number that feeds sizing and autoscaling decisions. Change one budget at a time, because their effects overlap and a two-variable change tells you nothing about which one moved the curve. Finally, know when the answer is not a knob. If cache utilization is pinned and prompts are long, no scheduler setting creates capacity — that is a memory, sharding, or replica-count conversation.
- Why does raising the concurrent-sequence cap sometimes reduce throughput instead of increasing it?Because the KV-cache pool, not the cap, is the real limit. Admitting more sequences than the pool can hold pushes the scheduler into repeatedly admitting and then evicting work; the evicted sequences' progress has to be recovered, so useful tokens per second falls. Longer steps from higher concurrency also stretch inter-token latency, so the same throughput number arrives with a worse user experience.
- Your prompts are 200 tokens and your outputs are 800. Which budget matters more?The concurrent-sequence cap. Short prompts mean prefill is cheap and rarely consumes the per-step token allowance, so the token budget seldom binds. Throughput is decided by how many sequences can decode simultaneously, which is bounded by KV cache for 1000-token sequences. Tune concurrency and cache headroom first; the token budget is close to irrelevant in this shape of traffic.
- How would you validate a scheduler change before rolling it out?Replay a realistic load profile — production prompt and output length distributions, not fixed-length synthetic traffic — at several concurrency levels against both the old and new settings, and compare first-token and per-token latency percentiles at equal offered load rather than at equal throughput. Change one budget at a time so the curve you moved is attributable, and hold the model, quantization and hardware constant.
saying these in an interview costs you the question
- Treats the sequence cap as if it allocated memory
- Sets the token budget from a blog post rather than an SLO
- Ignores that step duration is every stream's token latency
- Benchmarks with fixed-length synthetic prompts
- Tunes both budgets at once and cannot attribute the change