skip to content

Sharding & Model-Parallel Engines

0 questions

FSDP2's fully_shard and DeepSpeed ZeRO shard weights, gradients and optimizer state; Megatron-Core sets tensor, pipeline, context and expert sizes. LLM-training rounds ask for the memory math.

questions

no questions here yet

this part of the tree is still being written

>
0 questions in this topic or below it