DDP & Process Groups
0 questions
DistributedDataParallel replicates a model per rank and overlaps bucketed all-reduce with backward, while torchrun wires the process groups. Most multi-GPU interview war stories start here.
questions
no questions here yet
this part of the tree is still being written
>
0 questions in this topic or below it