A training job OOMs the day before a deadline — in what order do you apply the memory levers?
answer
- measure before you change anything
- cheapest currency first
- reversible before semantics-changing
- stop as soon as it fits
- fidelity changes go last
basics
~20 sMeasure the true peak first, then go cheapest-to-riskiest: shrink the micro-batch, restore the effective batch with gradient accumulation, add activation recomputation, drop precision, then shard state across devices. Anything that changes what the model sees comes last.
solid answer
~50 sOrder the levers by what they cost you, not by how much memory they free. First measure the real peak step, so you apply one lever instead of four. Then: shrink the micro-batch — instant, no code, but it changes the effective batch the recipe was tuned for. Restore that with gradient accumulation, which costs wall-clock roughly in proportion to the number of micro-batches. Then recompute activations instead of storing them, buying a large saving for extra forward compute. Then reduce numeric precision, which is often faster too but carries numerical risk that needs verifying. Then shard optimizer state and parameters across devices, which only helps if you have more than one and adds communication. Last, and only under real duress, the levers that change what the model sees — shorter sequences, lower resolution, dropped examples — because those change the result, not just its cost.
go deeper
Know the first two rungs and their order: make the micro-batch smaller to see whether the job fits, then use gradient accumulation to get the effective batch back. Say that you would measure before changing anything.
Explain what each lever spends — throughput, extra compute, numerical risk, communication — and why the ones that change what the model sees sit at the bottom of the list rather than the top.
Show the operating judgment: measure the peak, apply one lever, re-measure, and stop when it fits. Be ready to say when a batch budget replaces the whole ladder because the peak came from an outlier, not the model.
Own the policy, not just the run. Decide which levers are standing defaults for the team, which need validation before a deadline can depend on them, and how a run that used a fidelity lever is marked so nobody compares it to results from a different task.
## The principle: order by what it costs, not by what it saves Every memory lever buys room with a different currency. Some cost engineering time, some cost throughput, and a few cost **fidelity** — they change the model you end up with, so the run is no longer comparable to the ones before it. Under deadline pressure the instinct is to reach for whichever lever frees the most bytes. The better instinct is to spend the cheapest currency first and stop as soon as the job fits. ## Step zero: measure the peak Before touching anything, find out how much room you actually need. Instrument the peak step, not the average one, and note the failing request's size against the memory the process already holds. A job that misses by 5% needs one lever; a job that misses by 3x needs a different plan entirely — and knowing which saves you from stacking four changes and never learning which one mattered. This step costs minutes and routinely saves the afternoon. ## The ladder **1. Reduce the micro-batch size.** Instant, no code change, and it attacks the term that dominates most budgets. The cost is that the effective batch has changed, so the learning-rate schedule and any batch-size-dependent behaviour are now off-recipe. Treat this as a diagnostic first: does the job fit at all at a smaller micro-batch? If not, no amount of the later levers is going to rescue it either. **2. Restore the effective batch with gradient accumulation.** This puts the recipe back where it was — same effective batch, same schedule — at the cost of wall-clock time that grows roughly with the number of micro-batches per update, plus a caveat around layers whose statistics are computed per micro-batch. Steps 1 and 2 together are the standard first move, and for a job that misses by a modest margin they are usually the whole answer. **3. Recompute activations instead of storing them.** A large saving on the activation term, paid for with extra forward computation during the backward pass. It is semantics-preserving and needs no change to the optimization recipe, which is why it sits above the precision and sharding levers, but it is a real code change and a real throughput hit, so it comes after the free ones. **4. Reduce numeric precision.** Halving the width of activations and intermediate values is a big saving and is often *faster* rather than slower, which makes it tempting to move earlier. It sits here because it is the first lever with genuine numerical risk: overflow and underflow ranges change, and a run that trains stably at full width can misbehave without the appropriate scaling and accumulation choices. If you have never validated the recipe at reduced precision, doing it for the first time the night before a deadline is how you turn a memory problem into a convergence problem. **5. Shard state across devices.** Splitting optimizer state, gradients and parameters across several devices is mathematically the same computation and can free an enormous amount, but it only exists as an option if you have more than one device, and it adds communication and configuration to a system you are about to run unattended overnight. Powerful, but not a deadline-eve lever unless it is already in place. **Last: the levers that change the result.** Shortening sequences, lowering input resolution, dropping the longest examples, or cutting model width all fit the job by changing the problem. They are legitimate engineering choices — but they belong at the bottom because after using one you can no longer compare the run to yesterday's, and the number you report is a number for a different task. ## Judgment the ladder does not encode The ordering above assumes a single deadline-driven job. Two things change it. If the peak is caused by an outlier batch rather than the model, bounding the batch by a token or pixel budget is cheaper than every lever on the list and should be done instead of them. And if the same team keeps hitting this every month, the right call is not a better ladder but a standing budget: a memory metric logged on every run, a measured peak recorded with each configuration, and a policy that a configuration is only "validated" once it has survived its own worst-case step. ## What a strong answer sounds like The candidate names the ordering, gives the currency each lever spends, and — crucially — says when to stop. Applying every lever because they all help is how a job that needed 10% more room ends up 40% slower, at reduced precision, on a recipe nobody has validated, producing a number that cannot be compared to anything.
- Why not start with reduced precision, given it frees a lot and often runs faster?Because it is the first lever that can change whether the run converges. Reduced-precision training needs the right scaling and accumulation choices, and a recipe validated at full width is not automatically valid at half. When the schedule allows, it is an excellent default; adopted for the first time the night before a deadline, it converts a memory problem into a debugging problem.
- When would you skip the whole ladder?When the peak comes from an outlier batch rather than from the model. If one very long sample sets the peak, budgeting tokens or pixels per batch fits the job immediately, costs no throughput, and changes no recipe. Applying the ladder instead makes every step slower to survive a step that could simply have been bounded.
- How do you keep an emergency fit from becoming permanent technical debt?Record what was changed and why alongside the run, and mark the result as not comparable to earlier runs if a fidelity lever was used. Then schedule the proper fix — a validated precision recipe, a batch budget, or sharding set up calmly — instead of leaving a deadline-eve configuration as the team's new default.
saying these in an interview costs you the question
- Applies every lever at once without measuring the peak
- Reaches for sharding before trying a smaller micro-batch
- Shortens sequences first and reports the result as comparable
- Treats reduced precision as free because it also runs faster
- Cannot say what any given lever costs in throughput or fidelity