Your model sees a 2-billion-token stream once — how do you plan and report the run without epochs?
answer
- an epoch that happens once measures nothing
- steps and data consumed
- the budget must be fixed first
- validate on step intervals
- checkpoint the stream position too
basics
~20 sCount the run in optimizer steps and tokens consumed, not epochs — an epoch that happens once carries no progress information. Fix a step budget up front, hang evaluation and checkpointing on step intervals, and compare runs at equal tokens seen.
solid answer
~50 sWith a single pass, "epoch" degenerates: it occurs exactly once, at the end, so it measures nothing. The run's real units are **optimizer steps** and **tokens consumed**, related by `tokens = steps x batch_size_in_tokens`. Fix the total step budget before launch, because everything else is expressed in it — the learning-rate schedule spans the budget, and checkpoints and validation passes fire on step intervals rather than at a boundary that never arrives. Validation uses a fixed held-out shard evaluated every N steps. Per-epoch reshuffling is replaced by shuffling shards and a shuffle buffer inside the stream. Checkpoints must save the stream position alongside the parameters and optimizer state, or a resume silently re-shows the head of the corpus and never reaches the tail. Report tokens seen, steps, and batch size in tokens — comparing two runs by step count alone is meaningless if their batch sizes differ.
go deeper
Know that a training run does not need epochs at all, and that a step is the unit the loop actually counts.
Be able to convert between steps, batch size and data consumed, and explain why a schedule needs a total step budget fixed before launch.
Show that you would checkpoint the data-stream position with the parameters and hang evaluation on step intervals against a fixed held-out shard.
Own the reporting convention. Argue for steps and tokens as the org's units so runs stay comparable across batch sizes, corpora and pass counts.
## Why the epoch stops working An epoch is defined as one complete pass over the training set. That definition is useful only when the number of passes is greater than one and is a design choice. On a 2-billion-token stream consumed exactly once, the epoch count is 1 at the very end and 0 before that — it carries no information about progress, cannot be used to schedule anything, and cannot compare two runs. This is not an exotic regime; it is the default for large-corpus training, and interviewers use it to check whether a candidate treats the epoch as a fundamental unit or as the bookkeeping convenience it is. ## The units that replace it **Optimizer steps.** The loop still does exactly what it always did: take a batch, forward, loss, backward, update. The step counter is monotone, it is what every schedule can be written against, and it is what the run's budget is denominated in. **Tokens (or examples) consumed.** `tokens_seen = steps x tokens_per_batch`. This is the unit that describes how much *data* the model has been exposed to, independent of how it was chopped into batches. You need both, because they answer different questions. Steps tell you how many times the parameters moved; tokens tell you how much evidence produced those movements. Two runs at the same step count but different batch sizes consumed different amounts of data; two runs at the same token count but different batch sizes took different numbers of updates. Reporting only one of the two makes runs incomparable — this is the single most common reporting error in this regime. ## What has to be decided up front - **A total step budget.** Not a nice-to-have: the learning-rate schedule is defined over the whole budget, so the budget must be known before the first step. Changing your mind about the length mid-run leaves you with a schedule that does not land where intended. - **Evaluation cadence.** With no epoch boundary to hang it on, validation runs every N steps against a **fixed held-out shard** that the training stream never yields. Choosing N is a real tradeoff: too frequent and evaluation dominates the run's cost, too rare and you learn about a problem tens of thousands of steps late. - **Checkpoint cadence**, likewise in steps, sized so that the work lost to a crash is acceptable. - **Data ordering.** "Reshuffle before every epoch" has no meaning here. Randomness comes from shuffling the order of shards and from a shuffle buffer that mixes examples locally as the stream is read. The property you want — that consecutive batches are not correlated by how the corpus happened to be written — still has to be engineered, just differently. ## Resuming is a first-class design problem A long single-pass run *will* be interrupted. A checkpoint that stores only parameters is not resumable in any meaningful sense: restarting the stream from the beginning re-shows data the model has already been updated on and guarantees the tail of the corpus is never reached. A correct checkpoint stores parameters, optimizer state, the step counter, and the **position in the data stream**, plus enough of the shuffling state that the resumed order is the intended one. Teams discover this after the first crash, and the tell in an interview is whether you list the stream position without prompting. ## A genuine property of the single-pass regime Within one pass, every batch's loss is computed in the forward pass *before* that batch has ever contributed to an update. So the running training loss is measured on data the parameters have not fitted — it is close to a held-out measurement, and the usual train-versus-held-out gap is largely absent. This is a real and useful property: it means the training curve itself is informative about generalisation while the data lasts. It stops being true the moment any data repeats, which is exactly why the distinction between a single-pass run and a multi-epoch run is worth stating explicitly rather than treating the two as the same loop with a different counter. ## What you actually report A defensible run summary in this regime names: total steps taken, tokens consumed, batch size in tokens, the fraction of the corpus consumed, the held-out metric and the step at which it was measured, and the checkpoint that produced it. An epoch count either does not appear or appears as a fraction below one, which usually confuses more than it informs. ## The principal-level judgment The strategic call is the reporting convention itself. If your team reports epochs, then the moment someone doubles a batch size or swaps a corpus, the historical numbers stop meaning anything and nobody notices for a quarter. Standardising on steps and tokens costs nothing and keeps every run in the org comparable — including runs that pass over data once, runs that pass over it many times, and runs on corpora of different sizes. Own that convention, and require the stream position in every checkpoint contract.
- Two runs used different batch sizes — what do you compare them on?Data consumed, in tokens or examples, not step count. Equal steps at different batch sizes means one run saw several times as much data as the other, so a step-for-step comparison flatters the small-batch run's data efficiency and hides the difference. Report both numbers so either comparison is available.
- How do you resume such a run after a crash?Restore parameters, optimizer state, the step counter and the position in the data stream together, including enough shuffling state to reproduce the intended order. Restoring only the parameters restarts the stream at the beginning: the model re-trains on data it has already absorbed and the corpus tail is never reached, quietly turning a single-pass plan into something else.
- Why is the running training loss a decent proxy for held-out loss in a single-pass run?Each batch's loss is computed in the forward pass before that batch has contributed any update, so it measures performance on data the parameters have not fitted. That makes the training curve nearly a held-out curve. It stops holding the moment data repeats, which is why the property belongs to the single-pass regime specifically.
- Where does validation data come from when the training data is an endless stream?From a fixed held-out shard that the training stream is configured never to yield, evaluated on a step interval. It must be fixed rather than freshly sampled from the stream, otherwise the metric moves for two reasons at once — the model changing and the evaluation set changing — and run-to-run comparison is lost.
Epochs are laps; a single-pass stream is a point-to-point road. You do not report laps on a road trip — you report distance covered and how far is left.
saying these in an interview costs you the question
- Reports epochs when the corpus is seen once
- Compares runs by step count across different batch sizes
- Checkpoints parameters but not the stream position
- Assumes validation must happen at epoch boundaries
- Treats a fraction-of-an-epoch number as progress