skip to content

Training Loop Anatomy

One iteration takes a mini-batch forward, computes the loss, sends gradients back and updates every weight once; an epoch is one pass over the data. Interviewers ask what a step changes.

on this pageshow

questions

4

With 50,000 training examples and batch size 128, how many iterations and updates do 3 epochs take?

level: juniorimportance: must knowfreq 78%

answer

  1. one full pass over the data
  2. divide, then round up
  3. 50,000 over 128 is not whole
  4. one update per iteration
  5. 391 times three

basics

~20 s

An epoch is one full pass over the data. At batch size 128, 50,000 examples give 390 full batches plus a ragged batch of 80, so 391 iterations per epoch, one update each, and 1,173 updates over 3 epochs.

solid answer

~40 s

An **epoch** is one complete pass over the training set; an **iteration** (also called a step) is the processing of one mini-batch, and in the standard loop each iteration ends in exactly one parameter update. 50,000 divided by 128 is 390.625, so an epoch is 390 full batches of 128 covering 49,920 examples plus one ragged final batch of 80 — 391 iterations. Three epochs therefore perform 391 x 3 = 1,173 updates. If the loop is configured to discard the incomplete final batch, it is 390 iterations per epoch and 1,170 updates, with 80 examples unused each pass. The point of the arithmetic is that the optimizer only ever counts steps: epochs are bookkeeping over the dataset, not a unit of learning.

go deeper

for a junior

Be able to define epoch, iteration and batch size in one sentence each and do the division out loud without a calculator. Know that one iteration means one update.

for a middle

Explain the ceiling-versus-floor difference when the batch size does not divide the dataset, and why doubling the batch size halves the update count at fixed epochs.

for a senior

Show that you report runs in steps and samples seen rather than epochs, and that you know an unweighted average of per-batch losses is skewed by a ragged final batch.

for a principal

Own the convention your team reports in. Argue for a step-and-token budget over an epoch count so that runs with different batch sizes and dataset sizes remain comparable.

## The three words Three words describe the same loop at different granularities, and interviewers ask this because candidates routinely blur them. - **Mini-batch**: the group of examples processed together in one forward pass. Its size is the *batch size* — here 128. A batch is a slice of data, not an event in time. - **Iteration** (or **step**): one trip through the inner loop — take the next mini-batch, run the forward pass, compute the loss, run the backward pass, apply one update. "Iteration" and "step" are used interchangeably in the standard loop, because each iteration ends in exactly one update. - **Epoch**: one complete pass over the whole training set, i.e. as many iterations as there are mini-batches in the dataset. ## The arithmetic The number of mini-batches in one epoch is the dataset size divided by the batch size, rounded up when the loop keeps the leftover: ``` iterations_per_epoch = ceil(N / B) ``` With `N = 50,000` and `B = 128`: `50,000 / 128 = 390.625`. So 390 batches are full (390 x 128 = 49,920 examples) and 50,000 - 49,920 = **80 examples remain**, forming a ragged final batch. That is `391` iterations per epoch. Over three epochs: ``` 391 x 3 = 1,173 iterations = 1,173 parameter updates ``` Two variants matter in practice: 1. **Dropping the incomplete batch.** Many loops discard the final partial batch so every batch has an identical shape (which keeps timing uniform and, for anything that computes statistics across a batch, avoids a tiny final batch with unusually unstable statistics). Then it is `floor(N / B) = 390` iterations per epoch and `1,170` updates, and 80 examples go unused in each epoch. If the data is reshuffled every epoch, it is a *different* 80 each time, so nothing is permanently excluded — but it is 80 examples of work quietly thrown away every pass. 2. **Changing the batch size.** At batch size 256 the same 50,000 examples give `floor(50,000 / 256) = 195` full batches (49,920 examples) plus the same ragged 80 — 196 iterations per epoch and `588` updates over three epochs. Doubling the batch size roughly halves the number of updates for the same number of epochs. That is a pure counting fact, and it is why "we trained for 3 epochs" tells you far less about a run than "we took 1,173 updates". ## Why the distinction is load-bearing The optimizer has no idea what an epoch is. It sees a sequence of steps, and every quantity that is scheduled or budgeted is naturally expressed in steps: total training budget, checkpoint intervals, evaluation intervals, logging intervals. Epochs are a convenience for describing how many times the data has been recycled. This has practical consequences: - **Two runs at "the same number of epochs" are not comparable** if their batch sizes differ, because one performed several times as many updates as the other. - **Progress reported in epochs is coarse.** On a large dataset an epoch may be tens of thousands of steps; the interesting behaviour happens between epoch boundaries, so logging and evaluation are usually hung on step counts instead. - **"Steps" and "samples seen" are the two honest units.** Samples seen is `steps x batch_size` (here, after three epochs, 150,000 sample presentations — the 50,000-example set shown three times). ## The common confusions - *"An epoch is one update."* No — an epoch here is 391 updates. This is the single most common junior error. - *"The last batch always has the full batch size."* Only when `N` divides evenly by `B`, or when the loop discards the remainder. Otherwise the last batch is smaller, and if your loss is a mean over the batch, averaging per-batch means over the epoch weights those 80 examples more heavily than the other 49,920 — the unweighted average of 391 batch means is not the mean over 50,000 examples. - *"Batch size is the number of batches."* Batch size is examples per batch; the number of batches is the derived quantity. Be able to do this arithmetic out loud in a few seconds — it is a screening question, and hesitating on it reads as never having watched a training run's own logs.

  • What changes if the loop is configured to discard the incomplete final batch?
    You get 390 iterations per epoch instead of 391, so 1,170 updates over three epochs, and 80 examples are unused each pass. With reshuffling each epoch it is a different 80 every time, so no example is permanently excluded — but every batch now has an identical shape, which keeps step timing uniform and avoids a tiny final batch.
  • Keeping 3 epochs but doubling the batch size to 256 — how many updates now?
    50,000 divided by 256 is 195.3125, so 195 full batches plus the same leftover of 80: 196 iterations per epoch and 588 updates over three epochs. Roughly half the updates of the batch-128 run for exactly the same amount of data seen, which is why epoch counts alone never describe a run.
  • Are 'iteration' and 'step' the same thing?
    In the standard loop, yes. One iteration consumes one mini-batch, computes its loss, computes gradients and applies exactly one parameter update, so the words are used interchangeably and a step counter and an iteration counter are the same counter.

An epoch is a lap around the track; an iteration is one stride. Asking how many strides a lap takes needs the stride length — the batch size.

saying these in an interview costs you the question

  • Thinks one epoch means one parameter update
  • Says the final batch must contain a full 128 examples
  • Confuses batch size with the number of batches
  • Cannot state how many updates a finished run performed
  • Compares two runs by epoch count with different batch sizes

context

open as a page

Narrate one training step for a batch of 64 — which tensors change and which do not?

level: middleimportance: must knowfreq 68%

basics

~20 s

The forward pass turns the 64 inputs into activations and one loss number, the backward pass fills a gradient for every parameter, and the update writes new parameter values. Only parameters, optimizer state and gradient buffers change; the batch, targets and architecture do not.

open as a page

What breaks when a validation pass is run in training mode instead of evaluation mode?

level: seniorimportance: should knowfreq 60%

basics

~20 s

Dropout stays active, so every validation number is a noisy sample of a randomly thinned network, and normalization layers use the validation batch's own statistics and refresh their stored ones. The metric becomes irreproducible, batch-order dependent, and validation data leaks into the model.

open as a page

Your model sees a 2-billion-token stream once — how do you plan and report the run without epochs?

level: principalimportance: nice to knowfreq 38%

basics

~20 s

Count the run in optimizer steps and tokens consumed, not epochs — an epoch that happens once carries no progress information. Fix a step budget up front, hang evaluation and checkpointing on step intervals, and compare runs at equal tokens seen.

open as a page