skip to content

What does the Shannon entropy of an event log's status field, measured in bits, actually tell you?

level: juniorimportance: must knowfreq 62%

answer

  1. a property of the distribution
  2. average, not worst case
  3. uncertainty per value, before you look
  4. weighted by how often each value occurs
  5. H equals minus sum p log2 p

basics

~20 s

Shannon entropy is the average surprisal of a field's values in bits: the sum of p times log2(1/p) over every value. It measures how uncertain the next value is, not how many distinct values exist.

solid answer

~40 s

Shannon entropy `H = -sum p(x) log2 p(x)` is the **average** information content of one value of that field, in bits. Each possible value contributes its own surprisal `log2(1/p)` weighted by how often it actually occurs, so a value that is rare is very surprising but rarely counted, and a value that is near-certain is common but tells you almost nothing. Read it as uncertainty per value *before* you look: a field whose values are all equally likely carries the most, a field stuck on one value carries none. It is a property of the value distribution, not of any single record and not of the width you chose to store the field in. Always say what it is *per* — per value, per record, per second — because the number is a rate.

go deeper

for a junior

Recall the shape of the definition: entropy averages over a field's values, each weighted by its probability, and the unit is bits. Being able to say what it measures matters more than reciting the formula.

for a middle

Explain why the weighting matters: a value at probability 0.001 is extremely surprising yet adds about 0.01 bits, because it almost never occurs. Show that the figure depends only on the probabilities, not the labels or the storage width.

for a senior

Use it as a sizing unit in review: state entropy per field value, name the distribution you measured it on, and reject figures quoted with no denominator. Keep measured content separate from the width a field currently occupies.

for a principal

The judgment is what the measurement is for. A low-entropy field is a candidate for a cheaper representation, but that prize competes with query cost, schema churn and operability. Entropy sets the ceiling of the saving, not its worth.

## What the number is a property of **Shannon entropy** is a property of a *probability distribution*, never of a record, a file or a byte layout. Fix one field in an event log — say the `status` attached to every ingested record — and suppose that across the feed each possible value `x` occurs with long-run frequency `p(x)`. The entropy of that field is `H = - sum over x of p(x) * log2 p(x)` and because the logarithm is taken base 2, the unit is **bits**. Writing `log2(1/p(x))` for the *surprisal* of one particular value, the same formula reads as the **probability-weighted average surprisal**: the expected information carried by one value of the field. Three words in that sentence carry the whole meaning: - **average** — not the worst case, and not the surprisal of the rarest value the field can take; - **one value** — it is a *rate*, so the denominator must always be stated: bits per value, per record, per second; - **before** — it measures uncertainty about a value you have not yet read. After you read it, the value is known and there is nothing left to measure. ## Why the weighting is the interesting part The formula mixes two quantities that pull in opposite directions. Surprisal `log2(1/p)` grows without limit as a value becomes rarer: a value at `p = 0.001` carries about 9.97 bits of surprisal. But that value shows up in one record in a thousand, so its contribution to the average is `0.001 * 9.97 = 0.01` bits — negligible. A value at `p = 0.9` carries only 0.152 bits of surprisal, and even at nine records in ten its contribution is `0.9 * 0.152 = 0.137` bits. The consequence catches people out: **neither the rarest nor the commonest value dominates the sum**. The per-value contribution `-p log2 p` is zero at both ends (`p` approaching 0 and `p = 1`) and peaks in between, at `p = 1/e` (about 0.368), where a single value contributes about 0.53 bits. Mid-frequency values are what make a field expensive. ## What entropy is not | It is not | Why the confusion, and what is true | |---|---| | The count of distinct values | The count only fixes the **maximum**: `H <= log2 n` for `n` possible values. A field with a thousand values can average 3 bits if the distribution is skewed. | | The field's stored width | A value stored in one byte may carry 0.56 bits of content. The width is a representation choice; entropy measures the distribution. | | A property of one record | One record's value has a surprisal; the *field* has an entropy. Quoting entropy for a single value is a category error. | | A measure of usefulness | A field of random identifiers has high entropy and may be worthless; a near-constant flag may be the one that matters. Entropy measures unpredictability, not value. | ## Reading it on a real field In an ingestion pipeline sizing a day of structured records, the natural use is per field. You take the observed frequency of each value, compute `H`, and get a number in bits that says how much the field varies. A status field at 0.9 / 0.07 / 0.03 comes out at 0.557 bits; a field that is a distinct identifier on every record comes out near the log of its alphabet size. The two numbers tell you which fields are nearly free content and which ones genuinely differ record to record — before anyone argues about how to represent them. **Whether any particular encoding can actually reach that figure is a separate result about code lengths, and the entropy number by itself does not assert it.** ## Boundaries and conventions 1. **Entropy is never negative.** Every term is `p * log2(1/p)` with `0 <= p <= 1`, so every term is at least zero. A negative result means the probabilities do not sum to one, or a count is wrong. 2. **A value with `p = 0` contributes zero** by the standard convention that `p log2 p` tends to 0 as `p` tends to 0. Values that cannot occur do not cost anything. 3. **Relabelling changes nothing.** Entropy depends only on the multiset of probabilities, so renaming values, reordering them, or re-encoding them leaves `H` identical. Only changing *how often* each value occurs can move it. The practical habit worth forming: whenever you quote an entropy, quote the denominator and the distribution you measured it on. A bare "that field is 4 bits" is not a measurement anyone can check.

  • Two fields each measure 3 bits of Shannon entropy — do they hold the same number of distinct values?
    No. Three bits is `log2 8`, so a uniform eight-value field hits it exactly, but a field with a thousand values and a skewed distribution can average three bits too. Entropy pins the average uncertainty; only the maximum, `log2 n`, is tied to how many distinct values exist.
  • Does a field's Shannon entropy change if you rename its values?
    No. Entropy depends only on the multiset of probabilities, so relabelling three statuses as `1`, `2`, `3` leaves the number identical. Anything that touches only labels — renaming, reordering, re-encoding — cannot move it. Only a change in how often each value occurs can.
  • Can a Shannon entropy figure come out negative?
    No. Each term is a probability times a non-negative surprisal `log2(1/p)`, so every term is at least zero and `H >= 0`, with equality only when one value is certain. A negative figure means the probabilities do not sum to one or the counts feeding them are wrong.

saying these in an interview costs you the question

  • Says entropy counts how many distinct values a field has
  • Treats entropy as a property of one record's value
  • Confuses entropy with the byte width the field is stored in
  • Quotes an entropy figure without saying per value or per record
  • Assumes high entropy means the field is more useful