What limits which tensor-parallel sizes an LLM checkpoint supports?
answer
- it must divide something in the model
- heads are the sharded unit
- grouped-query attention shrinks the real limit
- packed weights need clean group boundaries
- legal is not the same as sensible
basics
~20 sThe degree must divide the model's structure: the attention head count above all, plus the MLP intermediate size, and with grouped-query attention the much smaller KV-head count becomes the binding limit. Hardware adds its own cap — the GPUs must share one fast interconnect domain.
solid answer
~50 sTensor parallelism cuts each layer along fixed dimensions, so the degree cannot be arbitrary. The hard rule is that the total number of attention heads must be divisible by the tensor-parallel size — you cannot give 64 heads to 6 ranks. The MLP's intermediate size must divide evenly too, which is why degrees in practice are powers of two. With grouped-query attention the KV heads are far fewer than query heads (a 70B-class model may have 64 query heads but only 8 KV heads), and that smaller number is what actually binds: past it, an engine either refuses the configuration or replicates KV heads across ranks, which costs extra cache memory per GPU. Quantized checkpoints add another divisibility constraint, since packed weights and per-group scales must split cleanly along the sharded axis. Finally, hardware: a degree that spans a slow link or a socket boundary is legal but usually slower, so the practical ceiling is the number of GPUs in one high-bandwidth domain.
go deeper
Know that you cannot pick any number of GPUs for one model — the count has to divide the model's internal structure, and the server will refuse to start otherwise.
Name the constraints: total attention heads divisible by the degree, the MLP intermediate size likewise, and the KV-head count as the tighter limit under grouped-query attention.
Read a checkpoint's config and choose a degree that satisfies divisibility, keeps KV heads unreplicated, respects quantized group boundaries, and stays inside one interconnect domain.
Treat supported degrees as part of model-selection criteria: a checkpoint whose head counts force an awkward layout constrains every future fleet decision, including which GPU SKUs you can standardise on.
## Why the degree is not a free parameter Tensor parallelism does not partition work abstractly — it slices specific tensors along specific axes. Every slice has to come out even, so the model's own dimensions decide which degrees are legal, and the hardware decides which of those are sensible. ## The attention constraint Attention is sharded by head: each rank owns a subset of the attention heads and computes them independently before the output projection is all-reduced. That requires the total head count to be divisible by the tensor-parallel size. 64 heads supports 1, 2, 4, 8, 16, 32, 64; it does not support 6 or 12. Serving engines validate this at startup and fail fast with a message about heads not being divisible by the parallel size — a very common first-launch error when someone picks a degree by counting free GPUs rather than reading the config. ## Grouped-query attention moves the real limit Modern checkpoints use grouped-query attention: many query heads share a much smaller number of key/value heads. A model with 64 query heads and 8 KV heads shards cleanly to TP=8, because each rank gets 8 query heads and 1 KV head. Beyond that there is no KV head left to give. Engines resolve this in one of two ways: reject the configuration, or *replicate* the KV heads so that several ranks hold identical copies. Replication works, but it means the KV cache no longer shrinks proportionally with the degree — each rank stores a duplicated slice — so the memory saving you expected from a higher degree does not materialize. Read the model config's head counts before choosing a degree; `num_attention_heads` and `num_key_value_heads` are the two numbers that matter. ## The MLP and vocabulary constraints The feed-forward block is sharded along the intermediate dimension, so that size must divide by the degree too. Embedding and output projections are often sharded along the vocabulary dimension, which usually divides freely but can require padding. In practice these are satisfied automatically when the head constraint is, because model dimensions are chosen as friendly powers of two. ## Quantized checkpoints Quantized weights carry structure of their own: values packed several to a machine word, and scale/zero-point tensors defined per group of channels (group sizes of 128 are common in 4-bit schemes). The sharding boundary has to fall on a group boundary and on a packing boundary, so a degree that is legal for a bf16 checkpoint can be rejected for the quantized version of the same model. Kernel support is a second-order constraint: some optimized kernels assume a minimum per-rank tile size, so a very high degree can leave shards too small for the fast path and quietly fall back to a slower one. ## Mixture-of-experts checkpoints For MoE models there is a further consideration: the expert count interacts with how experts are distributed, and a layout that shards experts is usually preferred to slicing each expert's matrices thinly. That is a distinct strategy rather than a divisibility rule, but it means the best degree for an MoE checkpoint is often not the same as for a dense one of similar size. ## The hardware ceiling Everything above says what is *legal*. What is *good* is bounded by topology: the ranks all-reduce roughly twice per layer per token, so a group that spans a PCIe hop, a CPU socket, or a network link will usually run slower than a smaller group that stays inside one NVLink domain. The practical ceiling is therefore "GPUs in one node connected by a fast fabric", commonly 8. Beyond that, add pipeline stages or replicas rather than raising the tensor-parallel degree. ## Where the degree is set Each engine spells it differently and the value is not portable: vLLM takes `--tensor-parallel-size`, TGI takes `--num-shard`, and TensorRT-LLM fixes the parallel layout when the engine is compiled, so changing it there means rebuilding rather than restarting. The divisibility rules above are properties of the checkpoint and apply to all of them. ## Debugging a rejected degree When a server refuses to start, read the head counts in the model config first, halve the degree, and retry. When it starts but the memory saving is smaller than expected, suspect KV-head replication. When it starts and is slower than a lower degree, suspect topology or a kernel fallback rather than the model.
- A model has 64 query heads and 8 KV heads. What happens at tensor-parallel size 16?Query heads split fine — 4 per rank — but there are only 8 KV heads for 16 ranks. The engine either rejects the configuration or replicates each KV head across two ranks. If it replicates, the model still runs but the KV cache stops shrinking in proportion to the degree, since duplicate copies are stored, so the memory benefit you were counting on largely disappears while the collective overhead still doubles.
- Why are tensor-parallel degrees almost always powers of two?Because model dimensions are chosen as powers of two — head counts, intermediate sizes, hidden sizes — so power-of-two degrees divide everything cleanly. Optimized kernels and collective algorithms are also tuned for those shapes and for balanced ring and tree topologies, and GPU nodes come in counts of 2, 4 and 8. A degree of 3 or 6 is usually either rejected outright or lands on a slower code path.
- Can you change the tensor-parallel degree without reloading the model?No. The degree determines how every weight tensor is sliced and which rank holds which slice, so it is fixed when the weights are loaded and the process group is formed; changing it means restarting the server. For engines that compile per-GPU kernels, such as TensorRT-LLM, it is fixed even earlier — at engine build time — so a different degree requires rebuilding the engine, not just relaunching it.
saying these in an interview costs you the question
- Picking the degree from the number of idle GPUs
- Ignoring KV-head count under grouped-query attention
- Assuming any divisor of the GPU count is valid
- Expecting a quantized checkpoint to shard exactly like its bf16 version
- Thinking the degree can be changed at runtime