Your character language model reports 1.35 bits per character - what does that number mean?
answer
- it is the training objective, re-based
- log base 2 instead of natural log
- exponentiate to get an effective branching factor
- 2^1.35 is about 2.55
- normalise by characters to compare fairly
basics
~20 sBits per character is the training objective itself in base 2: the mean negative log-likelihood the model assigns to each true next character on held-out text. Exponentiating gives a perplexity of about 2.55 characters of effective choice.
solid answer
~50 sBits per character is the average of `-log2 p(next character | prefix)` over the evaluation text - the same negative log-likelihood the model was trained to minimise, just reported in bits instead of nats. Perplexity is its exponential in the matching base: `2^1.35 = 2.55`. In nats the same figure is `1.35 * ln 2 = 0.936`, and `exp(0.936)` gives the same 2.55, which is the point - perplexity is base-independent as long as you match the exponential to the logarithm. The interpretation is effective branching factor: at each character the model's uncertainty equals that of picking uniformly among about 2.55 options. Compare it against a uniform baseline over the alphabet, roughly `log2(96) = 6.6` bits for printable ASCII, to see how much structure the model has captured. Because the normaliser is characters, this number is comparable across models with different tokenisations, which per-token perplexity is not.
go deeper
Be able to say that this is an average negative log-likelihood per character reported in bits, that lower is better, and that exponentiating it gives perplexity.
Expect to do the conversion out loud, in either base, and to explain the effective-branching-factor reading plus the uniform-alphabet baseline it should be compared against.
Show you can defend a comparison: insist on the same held-out text, no train overlap, and normalisation by characters rather than prediction steps before ranking two models.
Own what the team reports. Decide whether a likelihood figure is the right headline at all for the product's goal, and set the convention - unit, baseline and evaluation corpus - so numbers stay comparable across projects.
## The number and the objective are the same thing A next-character model outputs a distribution over the alphabet at every position. Its training objective is the mean negative log-likelihood of the characters that actually occurred: ``` mean NLL = -(1/N) * sum over positions of log p(true character | prefix) ``` Report that with the logarithm in base 2 and you have **bits per character**. Report it in base e and you have nats per character. Nothing else changed - no separate evaluation metric was computed. A model at 1.35 bits per character is a model whose held-out mean NLL is 1.35 bits, which is `1.35 * ln 2 = 0.936` nats. ## Perplexity Perplexity is the exponential of the mean NLL, taken in the same base as the logarithm: ``` perplexity = 2^(bits per character) = 2^1.35 = 2.55 perplexity = exp(nats per character) = exp(0.936) = 2.55 ``` Same number, as it must be. The interpretation is an **effective branching factor**: a model with perplexity 2.55 is, on average, as uncertain about the next character as someone choosing uniformly among 2.55 equally likely characters. A model that is certain everywhere has perplexity 1 and zero bits per character. A model that spreads mass uniformly over an alphabet of size `A` has perplexity `A` and `log2(A)` bits. That gives the baseline you need to judge 1.35. For roughly 96 printable ASCII characters, uniform guessing costs `log2(96) = 6.6` bits per character. English text also has strong unigram structure, so even a frequency table gets well below that. So 1.35 bits is a real reduction from the alphabet baseline, and the honest way to report it is alongside the baselines it beats, not as a bare figure. ## Why this is maximum likelihood The likelihood the model assigns to an entire held-out document is the product of its per-position probabilities. Products of many small numbers underflow and are impossible to compare, so take the logarithm, which turns the product into a sum; negate, which turns maximise into minimise; and divide by the number of characters, which stops the number growing with document length. The result is mean NLL. The parameters that minimise it are exactly the parameters that maximise the likelihood of the observed text under the model's categorical next-character distribution - the objective is maximum likelihood estimation, and bits per character is just its per-character, base-2 report. This is why the number is directly meaningful in a way that, say, character accuracy is not. It is on the scale of the thing being optimised, and differences in it are differences in log-likelihood. ## Comparing across models: the trap Models that segment text differently cannot be compared on **per-token** perplexity. A model with a coarse segmentation makes fewer predictions per document, and each one is harder; a model that works character by character makes many easy predictions. Per-token perplexity divides by a count that the segmentation choice controls, so the metric moves when nothing about modelling quality has changed. The fix is to normalise by a unit that both models agree on. Compute the total log-likelihood of the same held-out text under each model and divide by the number of **characters** (or bytes) in that text, not by the number of prediction steps. Then bits per character is comparable, because the numerator is the same quantity - the log-probability of the same text - and the denominator is a property of the text rather than of the model. If someone reports perplexity 12 and someone else reports 9 on the same corpus with different segmentations, the ranking is unresolved until both are converted. Two further comparability conditions that are easy to forget: the two models must be scored on the identical held-out text, and neither may have seen it in training. Both are more often violated than the arithmetic is. ## What the number does not tell you Bits per character scores the model's probability for the **true** next character given the **true** prefix. It says nothing directly about what the model produces when it consumes its own output, so a model can improve on this metric and still generate poorly - the metric never puts the model in the state its own mistakes create. It is also sensitive to the domain of the evaluation text: the same model scored on code, on prose, and on log files will report three different numbers, and a suspiciously low figure is often a sign that the evaluation text overlaps the training data rather than a sign of a strong model. So the senior read of `1.35 bits per character` is: I know exactly what was averaged, I can convert it to perplexity 2.55 in one step, I know which baseline to compare it against, I know it is the training objective and hence a likelihood, and I know the two ways someone can report a better-looking number without having a better model.
- Why is minimising average negative log-likelihood the same as maximum likelihood estimation?The likelihood of the whole corpus is the product of the per-position probabilities. Taking the logarithm turns that product into a sum, negating turns maximisation into minimisation, and dividing by the character count makes it independent of document length. None of those steps moves the optimum, so the parameters minimising mean NLL are exactly the maximum-likelihood parameters of the categorical next-character model.
- Two teams report perplexity 12 and 9 on the same corpus with different text segmentations - is 9 better?Unknown as reported. Per-token perplexity divides by a number of prediction steps that the segmentation controls, so a coarser segmentation makes fewer, harder predictions and a finer one makes many easy ones. Convert both to bits per character - total log-likelihood of the same text divided by its character count - and then they are comparable.
- Does a lower bits-per-character figure guarantee better generated text?No. The metric scores the probability of the true next character given the true prefix, so it never evaluates the model in the state its own errors put it in. It is also domain-sensitive and drops sharply if the evaluation text overlaps training data. Treat it as a likelihood measurement, and judge generation quality separately.
Perplexity is a branching factor: 2.55 means the model is about as unsure as someone flipping a slightly loaded three-sided die at every character.
saying these in an interview costs you the question
- Treats it as an accuracy rather than an average log-likelihood
- Exponentiates in the wrong base and reports the wrong perplexity
- Compares per-token perplexity across different text segmentations
- Quotes the figure with no baseline for the alphabet
- Cannot connect it back to the training objective