skip to content

DDP & Process Groups

0 questions

DistributedDataParallel replicates a model per rank and overlaps bucketed all-reduce with backward, while torchrun wires the process groups. Most multi-GPU interview war stories start here.

questions

no questions here yet

this part of the tree is still being written

>
0 questions in this topic or below it