GPU Training Fundamentals
You will learn the hardware realities of training: why batch size is bounded by GPU memory, what mixed precision and loss scaling buy you, and when gradient accumulation or data parallelism is the right fix. Interviewers use these to separate practitioners who have trained real models from paper readers.
on this pageshowhide
explore
- Arithmetic and Precision10 questions
- Compute vs Memory Bound3 questions
- Floating-Point Formats3 questions
- Mixed-Precision Training4 questions
- Memory Budget14 questions
- Model State and Activations3 questions
- Gradient Accumulation4 questions
- Activation Checkpointing4 questions
- Out-of-Memory Triage3 questions
- Scaling Across Devices8 questions
- Data Parallelism4 questions
- Model and Pipeline Parallelism4 questions
questions
page 2 of 2How do you combine tensor and pipeline parallelism when a checkpoint fits on no single device?
basics
~20 sMatch each strategy to the interconnect it needs. Tensor parallelism puts a blocking collective in every block, so keep it inside one high-bandwidth machine. Pipeline parallelism sends one small activation per stage boundary, so it tolerates slower links between machines.
A training job OOMs the day before a deadline — in what order do you apply the memory levers?
basics
~20 sMeasure the true peak first, then go cheapest-to-riskiest: shrink the micro-batch, restore the effective batch with gradient accumulation, add activation recomputation, drop precision, then shard state across devices. Anything that changes what the model sees comes last.
showing 31–32 of 32