Why does tensor parallelism scale poorly over PCIe compared with NVLink?
answer
- the collective is in the inner loop
- two per layer, every token
- link bandwidth times layer count
- small messages mean latency, not bandwidth
- check the topology before blaming the model
basics
~20 sTensor parallelism runs an all-reduce roughly twice per transformer layer, on the critical path of every token. NVLink moves those messages at hundreds of GB/s between GPUs; PCIe offers roughly an order of magnitude less bandwidth and higher latency, so the collectives dominate the step.
solid answer
~50 sTensor parallelism puts a collective inside the inner loop: after the attention block and after the MLP, every rank holds a partial sum that must be all-reduced before the next sub-layer runs. For an 80-layer model that is about 160 collectives per decode step — per token. On an NVLink/NVSwitch node each GPU has hundreds of GB/s of peer bandwidth and sub-10-microsecond small-message latency, so that traffic disappears into the compute. On a PCIe-only server each pair shares roughly 32 GB/s per direction on Gen4, often routed through the host bridge if peer-to-peer is unavailable, and small-message latency is much worse. At batch 1 the messages are tiny, so you are latency-bound, not bandwidth-bound, and 160 round trips per token is simply added wall-clock. That is why doubling GPUs over PCIe typically yields well under 2x, and why the usual advice is to keep the tensor-parallel group inside one NVLink domain and use pipeline parallelism or separate replicas beyond it.
go deeper
Know that GPUs sharing one model must constantly exchange data, and that the link between them (NVLink versus PCIe) can be the limiting factor rather than the GPUs themselves.
Be able to say where the traffic comes from — an all-reduce after attention and after the MLP in every layer — and estimate the message size from batch, hidden size and dtype.
Diagnose it on real hardware: compare TP=1 versus TP=2 throughput per GPU, read nvidia-smi topo -m, check NCCL's chosen transport, and decide whether to shrink the tensor-parallel group or switch to pipeline parallelism.
Turn this into a purchasing and placement rule: which GPU SKUs and node topologies you standardise on, how large a tensor-parallel domain you allow, and when a cheaper PCIe fleet running replicas beats a smaller NVLink fleet running shards.
## Where the traffic comes from Tensor parallelism shards each layer's matrices, so every rank computes a partial result over the hidden dimension. Those partials must be summed — an all-reduce — before the next sub-layer can run. There is one after attention and one after the MLP, which means about 2 collectives per layer, per forward pass, for every token generated. A 70B-class model with 80 layers therefore performs roughly 160 collectives *per decode step*. Nothing else in inference sits on the critical path that many times. ## The arithmetic, to an order of magnitude The message for one all-reduce is the activation slab: batch tokens x hidden_size x bytes per element. With 32 concurrent sequences in decode, hidden size 8192 and bf16, that is about 512 KB. A ring all-reduce moves roughly 2(N-1)/N times that per GPU — call it ~0.5 MB at N=2, more at higher degrees. Multiply by ~160 collectives and you are pushing tens of megabytes per GPU per token step. At NVLink-class bandwidth (hundreds of GB/s) that is a fraction of a millisecond and hides behind the matmuls. At PCIe Gen4 x16, roughly 32 GB/s per direction, the same traffic is milliseconds — comparable to or larger than the compute it is supposed to overlap. At batch 1 the picture inverts: the message is only tens of kilobytes, so bandwidth barely matters and *fixed latency per collective* dominates. Even a modest 30-50 microseconds of overhead per all-reduce, multiplied by 160, is several milliseconds added to every token. This is why low-concurrency, latency-sensitive deployments feel the interconnect most sharply. ## What makes PCIe worse than its headline number Headline PCIe bandwidth is a ceiling, not a promise. Whether two GPUs can DMA directly to each other depends on the topology: GPUs under the same PCIe switch can do peer-to-peer, GPUs on different root complexes (different CPU sockets) often cannot, and traffic then bounces through host memory, halving effective bandwidth and adding latency. Some virtualized or IOMMU/ACS configurations disable peer-to-peer entirely. Inspect this before blaming the model: `nvidia-smi topo -m` prints the pairwise link matrix (NV#, PIX, PHB, SYS), and running the server with `NCCL_DEBUG=INFO` makes NCCL print the transport it actually chose. Discovering that your "8-GPU box" is two 4-GPU islands across the socket boundary explains a lot of disappointing scaling. ## The practical rules that follow - **Keep the tensor-parallel group inside one high-bandwidth domain.** On an NVSwitch node that is all 8 GPUs; on a PCIe box it may realistically be 2, or the 4 under one switch. - **Cross a slow boundary with pipeline parallelism instead**, which sends one activation tensor per stage rather than a collective per layer. - **Or do not shard at all.** If the model fits on one GPU, independent replicas have zero collective traffic and usually beat a PCIe tensor-parallel group on throughput per dollar. - **Measure, don't assume.** Run the same model at TP=1 and TP=2 on the same hardware and compare tokens per second per GPU. Sub-linear is normal; a scaling factor near 1.1x on two GPUs means the interconnect is eating the gain. ## Version and hardware notes The numbers move with generations — PCIe Gen5 x16 roughly doubles Gen4, and each NVLink generation raises per-GPU peer bandwidth again — but the *ratio* has stayed roughly an order of magnitude in NVLink's favour, and it is the ratio that decides whether the collectives hide. Also note that a tensor-parallel group is one failure domain: it launches, stalls and dies together, and a stuck collective typically presents as a hung server rather than an error, so watch for NCCL timeouts in the logs. ## What a weak answer sounds like "PCIe is slower, so everything is slower." The point is not general slowness — it is that tensor parallelism uniquely places a synchronizing collective inside the per-token loop, so an interconnect deficit is multiplied by layer count and by every token you generate.
- How would you confirm that the interconnect, and not the model, is your bottleneck?Benchmark the same model at TP=1 and TP=2 on the same box and compare tokens per second per GPU; a gain far below linear points at the collectives. Then check `nvidia-smi topo -m` for whether the pair is NVLink, under one PCIe switch, or across sockets, and run the server with `NCCL_DEBUG=INFO` to see which transport NCCL selected. Falling back to host-staged transfers is a common culprit.
- Why does the interconnect penalty look worse at batch 1 than at batch 64?The all-reduce message size scales with the number of tokens in flight, so at batch 1 it is only tens of kilobytes — far too small to be bandwidth-limited. What dominates is the fixed per-collective latency, paid roughly twice per layer per token regardless of size. At batch 64 the messages are large enough that raw bandwidth matters and the fixed cost is amortized over far more useful work.
- If PCIe is your only option and the model does not fit on one GPU, what do you do?Prefer pipeline parallelism across the weak link: it sends one activation tensor per stage boundary instead of a collective per layer, so it tolerates low bandwidth, and it splits memory just as effectively. Accept that it needs sustained concurrency to fill the stages. Quantizing the weights so the model fits on a single GPU is often the better answer, since it removes cross-GPU traffic from the decode loop entirely.
saying these in an interview costs you the question
- Saying the collectives happen once per request, not per layer
- Expecting linear speedup from a second GPU
- Ignoring peer-to-peer topology and socket boundaries
- Quoting headline PCIe bandwidth as achievable throughput
- Assuming bigger batches make the interconnect matter less rather than more