skip to content

In multi-GPU LLM inference, how do tensor and pipeline parallelism differ?

level: middleimportance: must knowfreq 72%

answer

  1. two ways to cut a transformer
  2. inside a layer versus across layers
  3. collective per layer versus one handoff
  4. all-reduce cost versus pipeline bubble
  5. fast link inside node, slow link across

basics

~20 s

Tensor parallelism splits each layer's weight matrices across GPUs, so all of them work on every token and must all-reduce partial results at each layer. Pipeline parallelism gives each GPU a different block of layers and passes activations forward once per stage.

solid answer

~50 s

Both split one model across devices, but they cut it along different axes. **Tensor parallelism** (the Megatron-style cut) shards every weight matrix inside a layer — attention heads and the MLP's hidden dimension — so each GPU holds 1/N of the weights and 1/N of the KV cache, and every GPU participates in every token. The price is a collective (an all-reduce) roughly twice per transformer layer, on the critical path of every decode step, so it needs a fast interconnect and is normally kept inside one node. **Pipeline parallelism** cuts the layer stack crosswise: GPU 0 runs layers 0–39, GPU 1 runs 40–79, and only one activation tensor moves point-to-point at each boundary — small and tolerant of slow links, so it is what you use across nodes. But a single request still traverses every stage in sequence, so pipelining does not reduce per-token latency and idles stages ("bubbles") unless enough concurrent requests are in flight to keep every stage busy.

go deeper

for a junior

Know that a model too big for one GPU can be split, and be able to say the two basic cuts: split each layer across GPUs, or give each GPU a different group of layers.

for a middle

Explain the mechanics: tensor parallelism all-reduces roughly twice per layer because each rank holds a partial sum, while pipeline parallelism sends one activation tensor per stage boundary. Say which one lowers per-token latency.

for a senior

Show you have deployed both — argue tensor parallelism inside a node and pipeline parallelism across nodes from the traffic each generates, and explain why a pipeline needs sustained concurrency before it pays off.

for a principal

Own the layout decision end to end: TP degree x PP degree x replica count against a fixed GPU budget, including failure domains, whether the SLO is latency- or throughput-shaped, and what each choice costs per million tokens.

## Why a model gets split at all A served model has to fit in GPU memory alongside its KV cache and activations, and it has to produce tokens fast enough. When one GPU cannot hold the weights, or when reading those weights once per token is too slow, the model itself is split across devices. For transformer inference there are two general ways to cut it, plus a third that only applies to mixture-of-experts models. ## Tensor parallelism: cut each layer sideways Tensor parallelism shards the matrices *inside* every layer. In the attention block the attention heads are divided among ranks — each GPU computes its own subset of heads — and in the MLP the up-projection is split by columns while the down-projection is split by rows. Each GPU therefore produces a *partial* result over the hidden dimension, and the ranks must sum those partials before the next sub-layer can proceed. That summation is an all-reduce, and there is one after the attention block and one after the MLP: roughly two collectives per transformer layer, per forward pass, for every token generated. What you buy: each GPU stores only 1/N of the weights and, because KV heads are split with the attention heads, only 1/N of the KV cache. Aggregate memory bandwidth also scales, and since single-stream decode is bandwidth-bound, tensor parallelism is the split that can genuinely lower per-token latency for one request. What you pay: N GPUs sit on the critical path of every layer, and the whole group only goes as fast as its slowest link. ## Pipeline parallelism: cut the stack crosswise Pipeline parallelism assigns each GPU a contiguous *stage* of layers. The only thing crossing a stage boundary is the hidden-state activation for the tokens being processed — a single tensor of roughly (tokens x hidden_size x dtype bytes), sent point-to-point. That is orders of magnitude less traffic than tensor parallelism's per-layer collectives, and it is why pipeline parallelism is the strategy that survives a slow or high-latency link such as an Ethernet fabric between nodes. The catch is the bubble. For one request the stages run strictly in sequence: stage 2 cannot start until stage 1 finishes, so a lone request sees no latency improvement (and a little extra, from the handoffs), while every other stage sits idle. Serving engines hide the bubble by keeping many requests in flight, so different stages work on different micro-batches at the same time. Pipeline parallelism is therefore a memory-and-throughput tool that assumes concurrency, not a latency tool. ## How they compose The two are orthogonal and are routinely combined: tensor-parallel within a node where the GPUs share a high-bandwidth fabric, pipeline-parallel across nodes. Total GPUs = TP degree x PP degree (x replicas). Both split memory, so both can be used purely to make a model fit; only tensor parallelism reliably makes a single request faster. ## Where the engines expose this The knob names belong to each engine and differ: vLLM takes `--tensor-parallel-size` and `--pipeline-parallel-size` at launch; TGI's `--num-shard` sets the tensor-parallel degree; TensorRT-LLM fixes the parallel layout when the engine is *built*, not when it is served, so changing the degree there means rebuilding the engine. Do not assume one spelling carries across engines. ## Common misreads "Pipeline parallelism makes generation faster" — it does not, for a single request. "Tensor parallelism is always better" — only where the interconnect can carry the collectives; over a weak link it can be slower than not sharding. "Splitting the model gives me more total memory for everything" — tensor parallelism does hand each rank a smaller weight and KV-cache slice, but it also imposes a collective per layer, and the group fails as a single unit: lose one rank and the whole replica is down.

  • Which of the two actually reduces latency for a single request with no other traffic, and why?
    Tensor parallelism. Decode is memory-bandwidth-bound: each GPU reads only its shard of the weights per token, so aggregate bandwidth rises and the step gets shorter, as long as the per-layer all-reduces are cheap. Pipeline parallelism cannot help — the token still passes through every stage in order, and the extra stage handoffs make it marginally slower.
  • What is a pipeline bubble at inference time, and how does a serving engine hide it?
    A bubble is stage idle time: while stage 2 processes a batch, stage 1 has nothing to do unless another batch is ready. Engines hide it by keeping many requests in flight and feeding successive micro-batches into the pipeline, so all stages stay busy. At low concurrency the bubble cannot be hidden, which is why pipeline parallelism buys throughput, not latency.
  • How does tensor parallelism affect the KV cache each GPU stores?
    KV heads are sharded with the attention heads, so each rank holds roughly 1/N of the KV cache for the same sequences. Compared with running N independent replicas — which store N full copies of the weights — a tensor-parallel group stores one copy split N ways, leaving more free memory per GPU for cache and therefore supporting longer contexts or larger batches per replica.

saying these in an interview costs you the question

  • Claiming pipeline parallelism speeds up a single request
  • Thinking tensor parallelism sends data only once per request
  • Assuming pipeline stages run concurrently on the same token
  • Treating the two splits as interchangeable regardless of interconnect
  • Believing the engine flag names are the same across vLLM and TGI

context