skip to content

GPU Training Fundamentals

You will learn the hardware realities of training: why batch size is bounded by GPU memory, what mixed precision and loss scaling buy you, and when gradient accumulation or data parallelism is the right fix. Interviewers use these to separate practitioners who have trained real models from paper readers.

on this pageshow

explore

questions

page 2 of 2

How do you combine tensor and pipeline parallelism when a checkpoint fits on no single device?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Match each strategy to the interconnect it needs. Tensor parallelism puts a blocking collective in every block, so keep it inside one high-bandwidth machine. Pipeline parallelism sends one small activation per stage boundary, so it tolerates slower links between machines.

open as a page

A training job OOMs the day before a deadline — in what order do you apply the memory levers?

level: principalimportance: nice to knowfreq 36%

basics

~20 s

Measure the true peak first, then go cheapest-to-riskiest: shrink the micro-batch, restore the effective batch with gradient accumulation, add activation recomputation, drop precision, then shard state across devices. Anything that changes what the model sees comes last.

open as a page

showing 31–32 of 32