skip to content

Training Hyperparameters

Learning rate and its schedule, batch size, epochs, gradient accumulation and early stopping — plus the adapter rules that a far higher rate and a small effective batch are what actually work.

on this pageshow

questions

4

How does gradient accumulation reach a target effective batch size on one GPU?

level: middleimportance: must knowfreq 52%

answer

  1. a big batch you cannot hold at once
  2. several small passes, one update
  3. micro-batch times accumulation times devices
  4. memory trick, not a speed trick
  5. adapters prefer it modest, under about 32

basics

~20 s

Gradient accumulation runs several small micro-batches, sums their gradients, and steps the optimizer only once. Effective batch equals micro-batch size times accumulation steps times device count, so a large batch costs time instead of memory.

solid answer

~50 s

The batch size that matters for optimization is the **effective batch** — micro-batch size × accumulation steps × number of devices. Accumulation gets you there when the memory for a large micro-batch does not exist: you run the forward and backward pass on a small micro-batch, keep the gradients, repeat N times, then take one optimizer step and clear them. Activation memory scales with micro-batch × sequence length, which is why micro-batch 2 fits at long sequence lengths where micro-batch 4 hits out-of-memory, and why reaching an effective batch of 16 is done as 2 × 8 rather than 16 × 1. Two caveats matter. First, normalize the loss correctly — dividing by the micro-batch count skews the gradient when micro-batches carry different numbers of loss-bearing tokens. Second, adapter training tolerates very large effective batches worse than full fine-tuning, so most LoRA recipes keep the effective batch under about 32 rather than cranking accumulation up.

code

python · 9 lines
python
def effective_batch(micro_batch, accum_steps, devices=1):
    return micro_batch * accum_steps * devices

def steps_per_epoch(examples, eff_batch):
    return -(-examples // eff_batch)

print(effective_batch(2, 8))            # 16 on one GPU
print(effective_batch(4, 4, devices=2))  # 32 across two GPUs
print(steps_per_epoch(3000, effective_batch(2, 8)))  # 188

go deeper

for a junior

Know that the reported batch size is micro-batch times accumulation steps, and that accumulation exists so a large batch can be trained on a GPU too small to hold it at once.

for a middle

Explain the loop — several forward/backward passes, gradients summed, one optimizer step — and the memory reasoning: activations scale with micro-batch times sequence length while weights and optimizer state do not.

for a senior

Show the tuning judgment: find the largest micro-batch that fits at the required sequence length, set accumulation to hit the target effective batch, watch the loss normalization with variable-length data, and keep adapter runs at a modest effective batch.

for a principal

Own the reproducibility rule that only the effective batch is comparable across hardware, and decide where the budget goes — more devices, shorter context, activation checkpointing or adapter-only training — when a run does not fit.

## Effective batch is the number that matters Optimization does not care how a batch was assembled — it cares how many examples contributed to the gradient before the weights moved. That quantity is the **effective batch size**: effective_batch = micro_batch × accumulation_steps × devices The micro-batch is what one forward/backward pass holds in memory at once. The accumulation steps are how many such passes contribute to a single optimizer update. On multi-GPU runs, data-parallel replicas multiply it again. When a paper or a recipe quotes "batch size 64", it means the effective batch, and reproducing the run means matching that product, not one of its factors. ## The mechanic On each micro-batch you run the forward pass, compute the loss, and run the backward pass, which *adds* into the gradient buffers rather than replacing them. You do not step the optimizer. After N micro-batches you take one optimizer step and zero the gradients. Mathematically this reproduces the gradient of a single batch of N × micro-batch examples — subject to the normalization caveat below — while never holding more than one micro-batch of activations in memory. ## Why memory, not arithmetic, is the constraint During a backward pass the framework must retain intermediate activations from the forward pass. That storage scales roughly with micro-batch size × sequence length (attention adds further terms), and on a long-context fine-tune it dominates everything else. Model weights, gradients and optimizer state, by contrast, are fixed regardless of batch size. This is why the practical tuning loop is: pick the sequence length the task needs, find the largest micro-batch that does not OOM at that length, then set accumulation steps to reach the effective batch you want. A run at 4,096-token sequences might only fit micro-batch 2 on an 80 GB card, so an effective batch of 16 is 2 × 8. Halving the sequence length often lets micro-batch double and accumulation halve, for the same effective batch and better throughput. ## What accumulation does not buy you It is a memory trick, not a speed trick. You still perform the same amount of compute for the same number of examples, and small micro-batches under-utilise the GPU, so accumulation is usually *slower* per example than a genuine large batch would be on hardware that could hold it. It also does not shrink optimizer state or model weights — if the model plus optimizer does not fit, accumulation will not save you; quantization of the frozen base or adapter-only training will. ## The normalization pitfall The backward pass sums gradients across micro-batches, so the loss must be scaled down or the effective step is N times too large. The naive fix is to divide each micro-batch loss by N. That is exact only when every micro-batch contributes the same number of loss-bearing tokens. With variable-length examples, or when loss is computed on completion tokens only, micro-batches carry different token counts, and dividing by N over-weights the short ones. The correct normalization divides by the *total* loss-bearing token count across the accumulation window. This was a real, widely-reproduced bug in popular training stacks, and it is a good senior-level detail: it silently changes what you are optimizing without ever throwing an error. ## How large should the effective batch be Larger batches give lower-variance gradient estimates and more stable training, and they let you raise the learning rate — up to a point, beyond which extra examples per step stop buying anything and simply cost compute. For adapter-based fine-tuning there is a sharper limit: LoRA-style training degrades at large batch sizes noticeably earlier than full fine-tuning does, and increasing adapter capacity does not rescue it. The practical guidance that emerged from systematic studies is to keep the **effective batch under roughly 32** for adapter runs. That is a genuinely useful interview answer because it inverts the usual instinct that a bigger batch is always safer. ## Interactions to state out loud - **Learning rate.** Changing the effective batch changes what peak rate is appropriate; a rate found at effective batch 8 is not automatically right at 64. - **Steps per epoch.** Steps = examples ÷ effective batch. Raising the effective batch shortens the run in steps, which interacts with the schedule length and with how often you evaluate. - **Reproducibility.** Two runs match only if the *product* matches; 4 × 4 on one GPU and 2 × 4 on two GPUs are the same effective batch and should train alike, modulo data ordering.

  • Why is dividing each micro-batch loss by the accumulation count sometimes wrong?
    Because it assumes every micro-batch contributes the same number of loss-bearing tokens. With variable-length examples, or when loss is applied to completion tokens only, some micro-batches carry far fewer tokens, and dividing by the step count over-weights those. The correct normalization divides by the total loss-bearing token count across the accumulation window. The bug throws no error — it just quietly changes the objective.
  • You raise the effective batch from 8 to 64 and quality drops on an adapter run. What is going on?
    Adapter training tolerates large batches worse than full fine-tuning, and that limit is not fixed by giving the adapter more capacity. Studies of the effect put the practical ceiling around an effective batch of 32. Also check the learning rate: a rate tuned at batch 8 is not automatically right at 64, and step count per epoch has dropped eightfold, so the run may simply be shorter than the schedule assumes.
  • Micro-batch 4 fits at 1,024-token sequences but OOMs at 4,096. Why, and what do you change?
    Retained activations scale roughly with micro-batch times sequence length, so quadrupling the context quadruples the dominant memory term while weights and optimizer state stay fixed. Drop the micro-batch to 1 or 2 and raise accumulation steps to keep the same effective batch. If that is still tight, activation checkpointing trades recomputation for memory, and adapter-only training removes most of the optimizer state.

saying these in an interview costs you the question

  • Thinks accumulation makes training faster rather than fitting in memory
  • Quotes micro-batch size as the batch size of the run
  • Believes accumulation reduces optimizer or weight memory
  • Divides the loss by micro-batch count with variable-length examples
  • Assumes a bigger effective batch is always better for adapters

context

open as a page

Why does a LoRA fine-tune need a higher learning rate than full fine-tuning?

level: middleimportance: must knowfreq 58%

basics

~20 s

A LoRA run trains only a small adapter, so each step must move far fewer parameters and needs a bigger step size. The working rule is roughly ten times the full fine-tuning rate — about 1e-4 versus 1e-5.

open as a page

What do warmup and cosine decay do in a fine-tuning learning-rate schedule?

level: juniorimportance: should knowfreq 46%

basics

~20 s

Warmup ramps the learning rate from near zero up to its peak over the first few percent of steps so early updates cannot destabilise the model. Cosine decay then lowers it smoothly back toward zero so late steps refine instead of overwrite.

open as a page

How do you choose the number of epochs for a small fine-tuning dataset?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Start in the 1–3 epoch range: one pass often underfits a new style, three is the usual sweet spot, and more begins memorising a small set. Evaluate on a held-out split at a fixed step interval and keep the best checkpoint rather than the last.

open as a page