skip to content

What cross-entropy loss should a freshly initialized 1000-class classifier report on its first batch?

level: juniorimportance: should knowfreq 48%

answer

  1. one forward pass, no training needed
  2. an untrained softmax is essentially uniform
  3. negative log of one over the class count
  4. ln(1000) is about 6.9
  5. too high means confidently wrong at step zero

basics

~20 s

About 6.9, which is ln(1000). An untrained head spreads probability roughly uniformly, so the correct class receives about 1/1000 and the loss is -ln(1/1000). A much larger reading means the initial outputs are far from zero.

solid answer

~50 s

Cross-entropy on one example is `-ln(p_true)`. At initialization the output scores are near zero, the softmax over them is close to uniform, so `p_true` is about `1/C` and the loss should sit near `ln(C)` — `ln(1000) = 6.9` for a thousand balanced classes, `ln(10) = 2.3` for ten, `ln(2) = 0.69` for a binary head. Reading 11.4 instead means the head is already making confident wrong predictions, so the pre-softmax outputs are not near zero: the last layer's weights or bias are scaled far too large, or a nonlinearity was applied where raw scores were expected. A value well *below* `ln(C)` is equally suspicious — check whether the loss is averaged the way you think and what distribution the targets actually have. It is the first rung of the sanity ladder and costs one forward pass.

go deeper

for a junior

Be able to produce the number on demand: the loss is minus the log of the probability given to the true class, and an untrained model gives about one over the class count. Know ln(2), ln(10) and ln(1000).

for a middle

Explain why an untrained softmax is near-uniform, and translate an observed value back into an implied probability so you can say how far off the initial scores must be. Know what to expect when the bias encodes class priors.

for a senior

Demonstrate that you check this before any long job and use it to eliminate whole bug classes — output width, target format, loss reduction — before spending time on the optimizer. Extend it to structured outputs and to regression losses.

for a principal

Argue for cheap, computable expectations as gates in a team's training workflow: an assertion that the first-step loss matches its analytic value catches contract breaks between data and model code long before anyone reads a curve.

## The arithmetic For a single example with correct class `y`, cross-entropy is `-ln(p_y)`, where `p_y` is the probability the model assigns to `y` after the softmax. Before any training, the layer producing the pre-softmax scores has small random weights and a zero bias, so all `C` scores are close to zero, the softmax of a near-constant vector is close to uniform, and `p_y` is close to `1/C`. The averaged loss over a batch is therefore expected to be about ``` -ln(1/C) = ln(C) ``` A few values worth memorizing: `ln(2) = 0.69`, `ln(10) = 2.30`, `ln(100) = 4.61`, `ln(1000) = 6.91`. A binary head with a sigmoid output has the same story with `C = 2`: expect about 0.69 at step zero. ## Why the check earns its place It costs one forward pass and no training at all, and it is the only rung of the sanity ladder that can be evaluated before a single parameter has moved. It confirms three things at once: the output width matches the number of classes, the loss is reading the targets in the format it expects, and the initial scores are on a sane scale. Because it is so cheap, it belongs above the memorize-one-batch rung — there is no point trying to drive the loss down from a starting point that is already wrong. ## Reading a value that is too high Suppose a 1000-class model reports 11.4 rather than 6.9. Since `-ln(p_y) = 11.4` implies `p_y` is about `1.1e-5`, roughly a hundredth of the uniform `1/1000`, the model is not undecided — it is already confidently wrong. That happens when the pre-softmax scores are far from zero. Common causes: - The final layer's weights are initialized at too large a scale, so the scores spread over many units and the softmax concentrates on an arbitrary class. - A bias was initialized to something large or non-uniform. - Something between the head and the loss transforms the scores — a squashing nonlinearity applied where raw scores were expected, or a scaling factor left in place — so what reaches the loss is not what you think. The number itself carries information: `11.4 - 6.9 = 4.5` nats of excess is a large, structural discrepancy, not the kind of jitter a random draw produces. Random noise around a uniform softmax moves the initial loss by a few hundredths. ## Reading a value that is too low A loss noticeably below `ln(C)` at step zero is not good news. Two honest explanations exist. First, the head's bias may have been deliberately initialized to the log class priors, a standard trick under heavy class imbalance; then the correct expectation is not `ln(C)` but the entropy of the label distribution, which is strictly smaller — for a two-class problem at a 99/1 split that is about 0.056. Second, the target distribution may not be what you assume, or the reduction over positions or examples may be a sum where you expected a mean, or the reverse. Either way the correct move is to compute the expected value for the label distribution you actually have, then compare. ## Structured outputs The same check extends past flat classification. A per-pixel segmentation head predicting `C` classes at every location should also start near `ln(C)` when the loss is averaged over pixels, because each pixel's prediction is independently near-uniform. That case is instructive: a segmentation model whose initial loss lands exactly on the uniform-prior value has already demonstrated that the output width, the pixel-wise target format and the loss reduction line up. If that same model then stalls on a one-batch memorization test, the surviving suspects are the update path and the optimizer settings, not the shape of the head — the init check has ruled out an entire class of wiring bug before any training happened. ## The regression analogue For squared error the same reasoning gives a different constant. A model whose outputs start near zero produces an initial mean squared error of roughly the second moment of the targets, `mean(y^2)`, which equals `mean(y)^2 + var(y)`. On standardized targets that is about 1.0. Wildly larger means the targets were never standardized, or the head has a large bias; much smaller usually means the targets have almost no spread, which is worth knowing before you interpret any later loss number.

  • What if the initial loss comes in noticeably below ln(C)?
    Two honest causes. The output bias may have been initialized to the log class priors, in which case the right expectation is the entropy of the label distribution, not ln(C) — much lower under heavy imbalance. Otherwise suspect the reduction or the targets: a sum where you expected a mean, positions excluded from the average, or a label distribution that is not what you assumed.
  • A segmentation model's initial loss lands exactly on the uniform-prior value, but it then stalls on the one-batch test. What has the init check ruled out?
    It has ruled out the head's shape and the target format: the output has the right number of classes per pixel, targets are read in the expected form, and the loss reduction is sane. What remains is the update path and the optimizer settings — gradients not reaching parameters, a step size orders of magnitude off, or saturated activations. The wiring bug class is off the table.
  • Does an equivalent check exist for a regression model?
    Yes. With outputs starting near zero, squared error should begin near the second moment of the targets, mean(y)^2 + var(y) — about 1.0 on standardized targets. A far larger value usually means the targets were never scaled; a far smaller one means they barely vary, which changes how you should read every later loss number.

saying these in an interview costs you the question

  • Says the initial loss should be near zero
  • Quotes 1/C instead of the negative log of 1/C
  • Confuses the class count with its natural logarithm
  • Treats a below-expected initial loss as a good sign
  • Assumes ln(C) holds even when the bias encodes class priors

context