skip to content

How much device memory does an Adam-trained 1.5B-parameter model's state need before any activations?

level: middleimportance: must knowfreq 66%

answer

  1. Count bytes per parameter, then multiply
  2. Weight plus gradient is only half
  3. Adam carries two buffers per parameter
  4. Sixteen bytes each, times 1.5 billion
  5. Half precision keeps a 32-bit master copy

basics

~20 s

About 24 GB. Adam training in 32-bit floats costs 16 bytes per parameter - weight, gradient, and two moment estimates at 4 bytes each - so 1.5 billion parameters need roughly 24 GB before any activation exists.

solid answer

~50 s

Count bytes per parameter, then multiply. With 32-bit floats you store the weight (4 bytes), its gradient (4 bytes), and the two per-parameter buffers Adam carries between steps — the first and second moment estimates — at 4 bytes each. That is 16 bytes per parameter, so 1.5 billion parameters is about 24 GB of model state, and that number is fixed before you choose a batch size, an input resolution or anything else. The optimizer choice moves it a lot: plain SGD with no momentum is 8 bytes per parameter (12 GB here), SGD with momentum is 12 (18 GB). Under the common mixed-precision recipe the total stays near 16 bytes per parameter, because half-precision weights and gradients are paired with a 32-bit master weight copy and 32-bit moments — half precision saves mostly on activations, not on optimizer state.

go deeper

for a junior

Memorise the shape of the calculation: bytes per parameter times parameter count. Know that the gradient is as large as the weights and that adaptive optimizers add more per-parameter buffers on top.

for a middle

Derive the 16 bytes from the four tensors rather than quoting it, and be ready to redo the sum for plain SGD, for momentum, and for a frozen backbone. Interviewers here want the reasoning, not the constant.

for a senior

Use the number to make a call: which optimizer a run can afford on the hardware you have, and what remains of the budget for activations once state is subtracted. Explain why the mixed-precision master copy keeps state near 16 bytes.

for a principal

Own the tradeoff between optimizer memory and training stability across a fleet. Be prepared to justify spending state on adaptive moments versus buying batch or resolution with the same gigabytes, and to set a default for teams that cannot tune.

## The per-parameter accounting Model state is the part of the training footprint you can compute exactly on a whiteboard, which is why interviewers ask for it. The recipe is: decide what tensors exist **per parameter**, decide the byte width of each, add them up, and multiply by the parameter count. For Adam or AdamW in full 32-bit precision: | tensor | bytes per parameter | |---|---| | weight | 4 | | gradient | 4 | | first moment estimate | 4 | | second moment estimate | 4 | | **total** | **16** | So `1.5e9 * 16 = 2.4e10` bytes, about 24 GB using decimal gigabytes, or about 22.4 GiB in binary units. Either way, a single 24 GB device cannot host this run at all, and a 40 GB device hosts it with roughly 16 GB left for everything else. Why two moment buffers? Adam's update rule maintains a running average of the gradient and a running average of the squared gradient, and divides one by the square root of the other. Both averages are elementwise, so each needs its own tensor shaped exactly like the weights. That is the price of an adaptive per-parameter step size. ## What changes the number **The optimizer.** This is the biggest single lever on model state: - SGD with no momentum: weight + gradient = **8 bytes per parameter** → 12 GB here. - SGD with momentum (or Nesterov): one velocity buffer added = **12 bytes** → 18 GB. - Adam / AdamW: **16 bytes** → 24 GB. A candidate who can say "switching this run to momentum SGD frees 6 GB, at the cost of tuning the learning rate much more carefully" is doing the job the question is testing. **Precision.** The naive expectation is that half precision halves everything. It does not halve model state under the usual recipe. A typical mixed-precision setup keeps a half-precision copy of the weights for the forward and backward math (2 bytes), half-precision gradients (2 bytes), a 32-bit master copy of the weights that the optimizer actually updates (4 bytes), and 32-bit first and second moments (4 + 4). That adds to 16 bytes per parameter — the same total. The master copy exists because repeatedly adding tiny updates to a half-precision weight loses them to rounding. Where reduced precision genuinely pays is the activation pile, where every retained intermediate halves in size and nothing needs a master copy. **Frozen parameters.** A parameter that receives no update needs no gradient and no moment buffers. Freezing a backbone drops those parameters from 16 bytes to 4 (or 2 if the frozen copy is kept in half precision). Fine-tuning only a small head on a large frozen trunk is therefore dramatically cheaper in state than full fine-tuning, even though the parameter count is identical. **Buffers that are not parameters.** Normalization layers keep running statistics, and embedding-adjacent tables count as parameters like any other. These are usually small, but the accounting should be over the actual tensor inventory, not a guess at the architecture. ## Why this number matters before anything else Model state is the **floor**. It is resident for the entire step and does not shrink with batch size, sequence length, resolution or any input-side choice. Once you have it, the remaining budget for activations is simply `device_capacity - model_state - overheads`, and every input-shape decision is spending from that remainder. That framing catches a common planning error. A team sees a 1.5B-parameter model, notes that the weights are only 6 GB in 32-bit, and concludes it fits comfortably on a 24 GB device. In inference-like conditions it would. In training with Adam the same model needs 24 GB before it has computed anything, and the run cannot start. The reverse error also shows up: assuming the optimizer state is negligible because "it's just some averages". For any model where parameters dominate the footprint, optimizer state is two thirds of the model-state bill (8 of 16 bytes per parameter with Adam), which is the largest single block in the budget. ## How to present it in an interview State the per-parameter figure first, then the multiplication, then the sensitivity. "Sixteen bytes per parameter with Adam in 32-bit, so 24 GB for 1.5 billion; 12 GB if I drop to plain SGD; half precision doesn't move it much because of the master copy." That answer is complete in three sentences and demonstrates that you know which quantity each factor attaches to — which is exactly what the interviewer is probing.

  • What does that number become if you switch to SGD with momentum?
    Twelve bytes per parameter — weight, gradient and one velocity buffer — so about 18 GB instead of 24 GB for 1.5 billion parameters. You free roughly a quarter of the model state, but you give up the per-parameter adaptive step size and usually need more careful learning-rate tuning and warmup to match the same loss.
  • Does freezing most of the network reduce this?
    Substantially. A frozen parameter needs no gradient and no moment estimates, so it drops from 16 bytes to 4 — or 2 if the frozen copy is held in half precision. Training a small head on a large frozen trunk keeps the same parameter count but a small fraction of the state, which is why it fits where full fine-tuning does not.
  • If half precision doesn't shrink optimizer state, where does it actually help?
    On the activation side. Every retained intermediate tensor halves in size, and activations are the pile that scales with batch, resolution and sequence length, so that is where the gigabytes are. Model state stays near 16 bytes per parameter because the optimizer still updates a 32-bit master copy using 32-bit moments.

saying these in an interview costs you the question

  • Counts only the weights and calls that the footprint
  • Forgets that Adam stores two buffers per parameter
  • Claims half precision halves optimizer state
  • Thinks model state depends on batch size
  • Cannot convert parameter count into bytes

context