skip to content

What is a logit in a neural classifier, and why do losses take logits rather than probabilities?

level: juniorimportance: must knowfreq 72%

answer

  1. raw score, not a probability
  2. unbounded, any real number
  3. squashing happens last, or not at all
  4. exponential and logarithm cancel
  5. avoids overflow and log of zero

basics

~20 s

A logit is the raw, unbounded score a network's final layer emits before any squashing function turns it into a probability. Losses take logits so the squashing and the logarithm are computed together, which avoids overflow and log-of-zero.

solid answer

~50 s

A logit is just the raw output of the last linear layer: any real number, positive or negative, with no constraint that logits sum to one. A probability is what you get after squashing them - a sigmoid for one binary score, a softmax across a row of class scores. The reason a classification loss is defined on logits rather than probabilities is numerical, not stylistic: the loss ultimately wants `log(p)`, and going logit -> probability -> log throws away exactly the information the log needs. A tiny probability can round to exactly `0`, and `log(0)` is negative infinity; a probability near `1` loses most of its significant digits. Computed straight from logits, `log(p)` is a well-behaved finite number even at -400 nats. So keep logits as the model's output, and hand them to the loss unsquashed.

go deeper

for a junior

Be ready to say in one breath that a logit is the raw unbounded score before squashing, and that sigmoid or softmax is what turns it into a probability. Know that a binary logit of 0 means probability 0.5.

for a middle

Explain the mechanics: which squashing function belongs to which head, that only logit differences matter in a multiclass row, and why a loss is defined on logits so the exponential and the logarithm can cancel instead of being computed in sequence.

for a senior

Show the production judgment: keep the head unsquashed, squash only at reporting time, and store logits rather than probabilities when scores are persisted, because probabilities silently collapse to zero at the tail and cannot be recovered.

for a principal

Own the interface decision. Whether your serving contract carries logits or probabilities determines what downstream teams can do - recalibration, ensembling and threshold tuning all need the raw scale - and a probability-only contract quietly makes those impossible to do correctly later.

## What a logit is The last layer of a classifier is usually a plain affine map: `z = W h + b`. The numbers in `z` are the **logits**. They are unconstrained real numbers - a logit can be `-17.3`, `0.0`, or `+800`. Nothing about them is normalized, nothing about them is bounded, and in a multiclass model nothing forces them to sum to anything in particular. A **probability** is what you get after a squashing function is applied: - binary case, one logit `x`: `p = sigmoid(x) = 1 / (1 + exp(-x))`, giving a number in `(0, 1)`; - multiclass case, a row of logits `x_1 ... x_n`: `p_i = exp(x_i) / sum_j exp(x_j)`, giving a row that is non-negative and sums to `1`. Both maps are strictly increasing in the score, so ordering is preserved: the largest logit corresponds to the largest probability. The name "logit" comes from the inverse of the sigmoid, the log-odds `log(p / (1 - p))` - a raw binary score really is on a log-odds scale, where `0` means even odds, `+2.2` means about 90 percent, and `-2.2` means about 10 percent. ## Why the loss wants the logit, not the probability Every likelihood-based classification loss is built out of `log(p)`. The negative log-likelihood of the correct class `c` is `-log(p_c)`. So the chain a naive implementation performs is ``` logits -> exponentiate -> normalize -> probability -> logarithm -> loss ``` and the middle of that chain is a floating-point trap. **The exponential can overflow.** In single precision the largest representable number is about `3.4e38`, so `exp(x)` becomes infinity once `x` exceeds roughly `88`; in double precision the wall is around `709`. A confident decoder over a large vocabulary can genuinely emit a logit of `+800`. Exponentiating that first gives `inf`, and `inf / inf` gives `NaN` - the loss is destroyed before the log is ever reached. **The probability can underflow.** Go the other way: a logit far below the row maximum produces a probability that is not merely small but rounds to *exactly* `0`. In single precision `exp(d)` underflows to zero once the gap `d` is below about `-104`. Then `log(0)` is negative infinity (or an outright domain error), the per-example loss is `+inf`, and one such row turns the whole batch's mean loss into infinity. The next backward pass produces `NaN` parameters and the run is finished. **Precision is lost even when nothing blows up.** A probability near `1` - the confidently-correct case - is stored with almost no room to spare: `1 - 1e-9` is representable in single precision only as `1.0`, so the loss reads as exactly `0` when it should be `1e-9`. Going straight from the logit, that same loss is an ordinary small number computed to full relative precision. Working from logits sidesteps all three, because `log(p_i)` can be written as a subtraction of two well-scaled quantities instead of a log of a possibly-zero ratio. The output is a finite number even when the probability itself is far too small to represent. ## Practical consequences 1. **The model's output is logits.** A classifier head should not end in a squashing function when its output feeds a likelihood loss. Squash once, at reporting time, for humans and downstream consumers. 2. **Store logits, not probabilities, when you save scores.** Probabilities are a lossy encoding: everything below the underflow floor collapses to zero and is unrecoverable, while the logit that produced it is an ordinary number. 3. **Reading logits is a real skill.** The magnitude of a single logit means nothing on its own in the multiclass case - only the gaps between logits in the same row matter. In the binary case the single logit is directly interpretable as log-odds: `0` is `p = 0.5`, `+3` is about `0.95`, `-3` is about `0.047`. 4. **Thresholding can happen on logits.** Because the sigmoid is monotone, comparing the probability to `0.5` is identical to comparing the logit to `0`, and a probability threshold `t` maps to the logit threshold `log(t / (1 - t))`. ## The short version Logits are the model's native currency: unbounded, precise, and cheap to work with. Probabilities are a presentation format that happens to destroy precision at both ends of its range. Losses are defined on logits so that the exponential and the logarithm - which are inverses - can cancel against each other analytically instead of being computed one after the other in finite precision.

  • A binary classifier emits a logit of -3.0 for the positive class - what probability is that, and what is the decision at a 0.5 threshold?
    Applying the sigmoid gives `1 / (1 + exp(3)) = 1 / 21.09`, about `0.047`, so the model is roughly 95 percent sure the example is negative and a 0.5 threshold predicts negative. Because the sigmoid is strictly increasing, that comparison is equivalent to asking whether the logit is above zero - no squashing is needed to make the call.
  • Why not just store the probabilities and take the logarithm later when you need it?
    Because the probability is a lossy encoding. Any class whose score sits far enough below the row maximum rounds to exactly zero once normalized, and `log(0)` is negative infinity - the original score is gone and cannot be recovered. The logit that produced it, say `-400` relative to the maximum, is an unremarkable number that survives storage and gives an exact log-probability later.
  • In a multiclass row, does a single logit of +12 mean the model is confident?
    On its own it means nothing. Softmax depends only on the differences between logits in a row, so `+12` alongside `+11.9` is a near coin flip, while `+12` alongside `-5` is overwhelming confidence. Read gaps, not magnitudes. The binary case is different: there a single logit is a log-odds value and is directly interpretable.

Logits are the unrounded running total; probabilities are the total rounded to two decimals for the receipt. Do the arithmetic on the running total, print the receipt at the end.

saying these in an interview costs you the question

  • Calls a logit a probability that merely has not been normalized yet
  • Assumes logits must be non-negative or lie in [0, 1]
  • Squashes to probabilities and then feeds those into a log-likelihood loss
  • Reads one multiclass logit's magnitude as a confidence level
  • Thinks logits and probabilities are interchangeable because the ordering matches

context