skip to content

How do you combine tensor and pipeline parallelism when a checkpoint fits on no single device?

level: principalimportance: nice to knowfreq 30%

answer

  1. ask what binds: memory or links
  2. one collective per block versus per boundary
  3. the fast domain has a fixed size
  4. depth costs bubble, width costs bandwidth
  5. global batch falls out of the layout

basics

~20 s

Match each strategy to the interconnect it needs. Tensor parallelism puts a blocking collective in every block, so keep it inside one high-bandwidth machine. Pipeline parallelism sends one small activation per stage boundary, so it tolerates slower links between machines.

solid answer

~50 s

Decide by what actually binds. Tensor parallelism puts a blocking collective inside every block and its volume scales with tokens times hidden size, so it only pays inside a high-bandwidth domain — the fast links within one machine. Pipeline parallelism sends a single boundary activation point-to-point per micro-batch per stage boundary, which is small and latency-tolerant, so it maps naturally onto the slower links between machines; its price is the fill-and-drain bubble and the need for balanced stages. So the usual layout is tensor-parallel up to the number of devices in one high-bandwidth group, pipeline-parallel across groups with enough micro-batches to keep the bubble small, and replicas on top. Then check the second-order costs: pipeline depth forces a larger micro-batch count, which raises the global batch size and drags the learning-rate schedule with it. And before any of this, confirm that sharding the optimizer state alone would not have been enough.

go deeper

for a junior

Recall that a model can be cut across its matrices or cut across its depth, and that the two send very different amounts of data between devices.

for a middle

Be able to say what each strategy actually transmits — a collective inside every block versus one activation per stage boundary — and why that difference decides where each one can run.

for a senior

Show you would name the binding resource before choosing, map each degree onto the interconnect hierarchy, and profile stage times rather than counting layers.

for a principal

Own the whole layout as one decision with the training recipe: fix the acceptable global batch, then choose the degrees subject to it, and be explicit about which constraint each cut is buying you out of.

## Framing the decision When a model does not fit, there are two fundamentally different ways to cut it, and they fail on different resources. The right question is not *which is better* but *which constraint is currently binding* — device memory or interconnect bandwidth — and where in the cluster each cut can be afforded. ## The two cuts and what they cost on the wire **Tensor parallelism** splits individual weight matrices across devices. Every block ends in a collective that sums or gathers partial results, and that collective is on the critical path: the next layer cannot start until it completes, and there is no independent work inside the block to overlap it with. Its volume scales with the number of tokens times the hidden size — it grows with micro-batch and sequence length, and it recurs for every block, every micro-batch, every step. The upside is that it is fine-grained and bubble-free: all devices in the group work on the same tokens at the same time, so there is no idle structure to amortise away. **Pipeline parallelism** splits the model by depth. A stage's own computation involves no collective whatsoever; the only traffic is the activation tensor handed to the next stage and its gradient handed back — one tensor per boundary per micro-batch, sized by tokens times hidden, but *once per stage boundary* rather than once per block. That is orders of magnitude less traffic than tensor parallelism over the same model, and it is point-to-point rather than a group collective, so it is far more tolerant of latency and of a slower link. The upside is cheap communication; the downside is structural idle time — the fill-and-drain bubble — plus the requirement that stages be balanced. ## The mapping that follows Clusters are not flat. Devices inside one machine talk over links that are dramatically faster than the network between machines. That hierarchy determines the layout almost entirely: - **Tensor-parallel width = the size of the high-bandwidth domain**, and no wider. Stretch the tensor-parallel group across the network and every block's blocking collective now crosses the slow link; multiplied by every block and every micro-batch, that latency dominates the step and can cost far more than the memory it relieved. There is also a modelling ceiling: attention heads are not split, so the head count caps the useful width. - **Pipeline stages span machines.** Each boundary crossing carries one tensor, infrequently, and can be overlapped with compute, which is exactly what a slower link can absorb. - **Replicas sit on top**, each replica being a full tensor-and-pipeline arrangement, with the gradient combination between replicas being the outermost and least frequent traffic. ## Choosing the degrees Start from the memory requirement per device and work backwards. 1. **Check whether you need model surgery at all.** If the shortfall is optimizer bookkeeping rather than the weights themselves, sharding that state across replicas is nearly free and involves no splitting. Only when a full weight copy plus the activations a device must retain still will not fit do you cut the model. 2. **Set the tensor-parallel degree** to the smallest width that brings per-device weight memory into range, capped by the high-bandwidth domain size. 3. **Set the pipeline depth** to the smallest depth that closes the remaining gap. Every extra stage adds a slot to both the fill and the drain, so depth is a cost, not a benefit — buy only what you need. 4. **Set the micro-batch count** several times the stage count to keep the bubble modest, then check what that does to the global batch. ## The second-order cost people miss The parallelism layout is not invisible to the optimizer. Global batch size equals micro-batch size times micro-batch count times replica count. Deepening the pipeline forces a larger micro-batch count to keep the bubble small, which inflates the global batch, which changes the learning-rate schedule, warmup length and the number of steps to convergence. A layout chosen purely on throughput can hand you a training recipe you never intended. Treat the global batch as a first-class constraint and solve the layout subject to it, rather than reading it off afterwards. ## Balance beats layer counting Pipeline stages must be balanced by *measured time*, not by layer count. In a deep encoder-decoder, a stage that happens to hold twice the layers of its neighbours sets the cadence for the whole pipeline, and every other device idles the difference on every micro-batch — a loss that scales with the micro-batch count instead of being amortised by it. Non-uniform pieces make this worse: an input embedding, a tied output projection and the loss computation have very different costs from a plain block, so the first and last stages routinely need fewer layers than the middle ones. Profile the stages, then move layers until the times match. ## What to say when asked A strong answer names the binding resource first, maps each strategy to the level of the interconnect hierarchy it can afford, states the degrees in the order you would fix them, and closes with the global-batch consequence. A weak answer recites the names of the strategies without ever saying what makes one of them unaffordable across a network.

  • What happens if you extend the tensor-parallel group across two machines?
    Every block's collective now crosses the slow network link while sitting on the critical path, and it recurs for every block and every micro-batch. Step time typically degrades far more than the per-device memory relief is worth. If memory is still short, the better moves are deepening the pipeline, sharding optimizer state, or reducing what each stage must retain — anything that does not put a blocking collective on a slow link.
  • Once the tensor-parallel width is fixed, how do you pick the number of pipeline stages?
    Take the smallest depth that makes per-device memory fit. Depth is a pure cost: each extra stage adds a slot to the fill and to the drain, and forces a higher micro-batch count to keep the bubble in check, which in turn inflates the global batch. Then balance the chosen stages by measured stage time, not by layer count, so that no single stage sets the cadence for everyone else.
  • How does the chosen layout constrain the optimization recipe?
    Global batch size is micro-batch size times micro-batch count times replica count. Deepening the pipeline pushes the micro-batch count up to suppress the bubble, so the global batch grows and the learning-rate schedule, warmup and step budget all have to move with it. Decide the acceptable global batch first and solve the layout subject to it; discovering it after the fact usually means retuning the whole recipe.

saying these in an interview costs you the question

  • Treats tensor and pipeline parallelism as interchangeable
  • Spreads a tensor-parallel group across a slow network link
  • Adds pipeline stages as if depth were free
  • Never mentions what the interconnect can actually carry
  • Ignores that the layout fixes the global batch size
  • Balances pipeline stages by layer count rather than measured time

context