Sharding & Model-Parallel Engines
0 questions
FSDP2's fully_shard and DeepSpeed ZeRO shard weights, gradients and optimizer state; Megatron-Core sets tensor, pipeline, context and expert sizes. LLM-training rounds ask for the memory math.
questions
no questions here yet
this part of the tree is still being written
>
0 questions in this topic or below it