skip to content

Shannon entropy of a status field is 0.56 bits at 0.9/0.07/0.03 but 1.58 bits when uniform — why?

level: middleimportance: must knowfreq 55%

answer

  1. how spread out, not how many
  2. uniform is the maximum
  3. certainty costs zero bits
  4. nine in ten are guessable
  5. top is log2 n, bottom is 0

basics

~20 s

Entropy peaks at log2 n when the n values are equally likely and falls to zero when one is certain. Skew moves a field toward certainty: at 0.9/0.07/0.03 most records are predictable, so the average drops to 0.56 bits.

solid answer

~40 s

Entropy is bounded: `0 <= H <= log2 n` for an n-value field. Uniform is the top — three equally likely values give `log2 3 = 1.585` bits — and a field stuck on one value is the bottom, exactly 0 bits. The skewed field sits near the bottom because nine records in ten carry a value you could have guessed without looking. Breaking the sum down, the 0.9 value contributes 0.137 bits, the 0.07 value 0.269 and the 0.03 value 0.152, totalling 0.557 bits, about 35% of the uniform maximum. Note which term is largest: the common value contributes *least*, because it is unsurprising, and the rarest contributes little because it is rare. Mid-frequency values dominate.

code

pseudocode · 15 lines
pseudocode
values = [ "ordinary", "retried", "failed" ]
p      = [  0.90,       0.07,      0.03    ]

H = 0
for each i in 0 .. 2:
    term = -p[i] * log2(p[i])
    print values[i], "contributes", term
    H = H + term

print "H =", H

// ordinary contributes 0.137
// retried  contributes 0.269   <- largest, at only 7% of records
// failed   contributes 0.152
// H = 0.557 bits per value

go deeper

for a junior

Remember the two anchors: equally likely values give the most entropy, a single certain value gives zero. Being able to say which of two distributions is higher is enough at this stage.

for a middle

Explain the bound 0 <= H <= log2 n and work one small sum out loud. The point to land is that the dominant value contributes the fewest bits because it is unsurprising, not the most because it is frequent.

for a senior

Push back on cardinality arguments in review: ask for the distribution, normalise by log2 n when comparing fields, and say what the measurement is blind to — order and cross-field structure are not in a single-field figure.

for a principal

Decide what a low-entropy field is worth acting on. Heavy skew is a standing invitation to a cheaper representation, but a distribution that drifts turns yesterday's saving into tomorrow's migration.

## The two ends of the scale For a field with `n` possible values, Shannon entropy is pinned between two limits: - **Maximum, `log2 n`, at the uniform distribution.** When every value is equally likely there is no guess better than any other, and each value carries the full `log2 n` bits. Three values give `log2 3 = 1.585` bits; eight give exactly 3. - **Minimum, 0, under certainty.** When one value has probability 1, its surprisal `log2(1/1)` is 0 and every other term is `0 * log2 0`, taken as 0 by convention. A constant field carries no content at all. Everything else lies in between, and **skew is simply movement from the top toward the bottom**. The more the mass concentrates on one value, the better you can guess without looking, and the less each observation tells you. ## Working the skewed field term by term Take an ingestion pipeline whose `status` field runs 0.9 ordinary, 0.07 retried, 0.03 failed: | Value | Probability `p` | Surprisal `log2(1/p)` | Contribution `-p log2 p` | |---|---|---|---| | ordinary | 0.90 | 0.152 bits | **0.137 bits** | | retried | 0.07 | 3.837 bits | **0.269 bits** | | failed | 0.03 | 5.059 bits | **0.152 bits** | | total | 1.00 | — | **0.557 bits** | Against the uniform maximum of 1.585 bits, the skewed field carries about 35% as much. Against a naive fixed-width code — three values need 2 bits — it is about 28%. The ordering in the last column is the part worth arguing in an interview. The value that appears in nine records out of ten contributes the **least** of the three, and the mid-frequency value contributes the **most**. That is not an accident of these numbers: the function `-p log2 p` is zero at `p = 0`, zero at `p = 1`, and peaks at `p = 1/e` (about 0.368) where a single value contributes roughly 0.53 bits. ## Why skew, not cardinality, is what moves the number Two fields with the same number of possible values can sit anywhere on the scale, and two fields with wildly different cardinality can sit at the same place: 1. **Adding a rare value barely moves it.** Give the status field a fourth value at `p = 0.001`, taken from the dominant value. Its own term is `0.001 * log2 1000 = 0.010` bits, and the dominant value's term shifts by about 0.001 bits. Total: roughly 0.568 bits, up about 0.01. A long tail of rare values is not what makes a field expensive. 2. **Flattening the existing values moves it a lot.** Shift the same three values to 0.5 / 0.3 / 0.2 and entropy rises to about 1.49 bits, nearly the uniform maximum, without adding a single new value. 3. **Removing a value's variability collapses it.** If retries and failures stop, the field is constant and entropy is exactly 0 — the field can be lifted out of the record entirely without losing information. ## Reading it in a pipeline review The practical translation is: **entropy measures how surprised you are on average, and heavy skew means you are rarely surprised.** When someone reports that a field "has lots of different values", that is a statement about cardinality and it does not settle the question. Ask for the distribution. A field with a hundred values where one holds 99% of the mass is closer to a constant than to a hundred-way choice. The same reading works in the other direction, and it is the honest caveat on the whole exercise: a field that measures near its `log2 n` maximum has no cheap structure to exploit *in its value distribution*. There may still be structure elsewhere — order, correlation with another field, a pattern inside the value itself — but single-field entropy is blind to all of it by construction. ## Common traps - **Assuming the dominant value carries the bits.** It carries the records; the mid-frequency value carries the bits. - **Assuming a rare value is expensive.** It is surprising per occurrence and almost free on average. - **Comparing entropies across fields with different cardinality without normalising.** 2 bits on a 4-value field is the maximum; 2 bits on a 1000-value field is deep skew. Divide by `log2 n` when you want them on the same scale.

  • What is the Shannon entropy of a field that only ever holds one value, and what does that mean for the record?
    Exactly 0 bits: the sole value has probability 1 and `log2 1 = 0`, while absent values contribute 0 by convention. The field is fully determined before you read it, so it carries no content at all and could be lifted out of every record into the schema without losing information.
  • A fourth status value appears at probability 0.001 — roughly how much does the entropy rise?
    About 0.01 bits, from 0.557 to roughly 0.568. Its own term is `0.001 * log2 1000 = 0.010` bits and the mass it takes from the dominant value shifts that term by about 0.001. Rare values move the average almost not at all.
  • Which single value contributes the most bits at 0.9 / 0.07 / 0.03, and why is it not the most common one?
    The 0.07 value, at 0.269 bits against the dominant value's 0.137. The per-value contribution `-p log2 p` is zero at both `p = 0` and `p = 1` and peaks at `p = 1/e` (about 0.368, worth roughly 0.53 bits), so mid-frequency values dominate the sum.

Forecasting weather where it is dry nine days in ten: you are right almost every day by saying "dry" without thinking, so each forecast tells you very little on average. Entropy is the average effort of guessing, not the effort on the rare wet day.

saying these in an interview costs you the question

  • Says entropy depends only on the number of distinct values
  • Claims one rare value raises entropy sharply because it is surprising
  • Thinks a constant field still costs log2 of its declared width
  • Assumes the most frequent value contributes the most bits
  • Believes skew raises entropy because the data looks lopsided