skip to content

How do you turn a field's per-symbol Shannon entropy into bits per record and a daily volume figure?

level: middleimportance: should knowfreq 40%

answer

  1. it is a rate, not a total
  2. say bits per what
  3. multiply by values per record
  4. independence licenses the multiplying
  5. bits to bytes: divide by eight

basics

~20 s

Shannon entropy is a rate: bits per value. Multiply by the number of values a record carries for bits per record, then by records per day and divide by eight for bytes — valid while the values are independent draws.

solid answer

~40 s

Entropy comes out of the formula as **bits per value**, so the whole exercise is keeping track of the denominator. If each record carries one status value at 0.557 bits, then a feed of 200 million records a day carries about 111 million bits, or roughly 14 MB of status content. If each record carries eight such values, independently drawn from the same distribution, that is `8 * 0.557 = 4.46` bits per record and about 111 MB a day. Compare that with a fixed one-byte code per value — 1.6 GB a day — and the gap is the headroom. The multiplication is only licensed by **independence**: if the values within a record are correlated, the product overstates the true content and you need the two-variable machinery to say by how much.

go deeper

for a junior

Hold on to the unit: entropy comes out as bits per value. Multiplying to a per-record or per-day figure means multiplying by counts and dividing by eight to reach bytes.

for a middle

Carry the arithmetic end to end and name the assumption. Independent, identically distributed values let you multiply the rate by the count; correlated values do not, and the product then runs high rather than low.

for a senior

In a capacity review, separate measured content from what the current representation spends, quote the ratio rather than the absolute, and say which way the estimate errs so nobody plans against a bound they think is exact.

for a principal

Decide whether the headroom is worth spending engineering on at all. A fourteen-to-one gap on one field of one feed may still lose to the operational cost of a cleverer representation that everything downstream must understand.

## Entropy is a rate, so name the denominator The formula `H = -sum p log2 p` produces a number in **bits per value of one field**. Almost every mistake in this area is a denominator mistake: quoting "the field is 0.56 bits" and then treating that as a per-record or per-day quantity. Three different numbers are in play in a sizing conversation, and they differ by factors that matter: - **bits per value** — what the formula gives you directly; - **bits per record** — the per-value rate times the number of values of that field a record carries; - **bytes per day** — bits per record times records per day, divided by 8. ## Multiplying out one field for one day Take an ingestion pipeline at 200 million records a day, with a status field measured at 0.557 bits per value. 1. **One value per record.** `200,000,000 * 0.557 = 111.4` million bits. Divided by 8, about **13.9 MB** of status content per day. 2. **Eight values per record** — say a per-stage outcome for each stage of the pipeline, each drawn independently from the same distribution. Per record: `8 * 0.557 = 4.46` bits. Per day: `8 * 111.4 = 891` million bits, about **111 MB**. 3. **What a naive representation costs.** One byte per value is `8 * 200,000,000 = 1.6` billion bytes, about **1.6 GB** a day. | Quantity | Figure | What it tells you | |---|---|---| | Entropy per value | 0.557 bits | Average content of one value | | Bits per record, 8 values | 4.46 bits | Content of that field group in one record | | Per day at 200M records | about 111 MB | The content-side budget | | Fixed one byte per value | about 1.6 GB | What an indifferent representation costs | The ratio, roughly 14 to 1, is the number the conversation is actually about. **Whether any particular encoding can reach the content figure is a separate result about code lengths, and the entropy number by itself does not assert it** — but it does tell you the size of the prize before anyone builds anything. ## Where the multiplication stops being valid Multiplying a per-value rate by a count assumes each value is an **independent draw from the same distribution**. Two ways that breaks in real feeds: - **Correlation within a record.** If a failure at stage three makes a failure at stage four near-certain, the eight values are not eight independent draws, and their combined content is *below* `8 * 0.557` bits. The sum is then an upper bound, and measuring the overlap needs the two-variable forms rather than the single-field figure you started with. - **A distribution that is not the same everywhere.** If the status mix differs by tenant, by hour, or by client version, one pooled `H` describes a blend that matches no individual stream. Pooling across heterogeneous sources generally inflates the figure relative to any one source. Both failures point the same way: **the product is an upper bound on the real content, never an underestimate.** That is the useful direction for capacity work, as long as you say so out loud. ## What the daily figure is and is not It is a statement about the **value distribution**, not a prediction of bytes on disk. A feed measuring 14 MB a day of status content may well write 200 MB, because a fixed-width representation spends the same room on every record regardless of how predictable that record's value was. The gap is not a bug; it is the headroom, and whether it is worth taking depends on cost, on query patterns and on how stable the distribution is. It also says nothing about the *rest* of the record. Sizing one field is exactly that: a per-field statement. Adding up all fields is another multiplication with the same independence caveat, and in real schemas fields correlate heavily. ## Doing it in a review - Quote three numbers, always together: the rate, the multiplier, and the period. - Convert to bytes explicitly; the factor of 8 is the most common slip. - State the independence assumption when you multiply, and say which direction the error runs. - Compare against what the field costs today, not against zero — the interesting quantity is the ratio, not the absolute figure.

  • Two fields in the same record are strongly correlated — what does summing their Shannon entropies give you?
    An upper bound on their combined content, not the true figure. Summing is exact only when the fields are independent; correlation means part of the second field is already implied by the first, so the truth sits below the sum. Quantifying that overlap requires the two-variable forms, a different measure from single-field entropy.
  • Your daily figure says 14 MB but the feed actually writes 200 MB for that field — is the measurement wrong?
    Not necessarily. Entropy measures the content of the value distribution; the 200 MB is what the chosen representation costs, and a fixed-width code spends the same room on every record however predictable it is. The gap is headroom. Whether taking it is worthwhile is a separate engineering question.
  • The status mix differs sharply between tenants — what does one pooled entropy figure describe?
    A blend that matches no individual tenant, and generally a higher figure than any single stream, because pooling mixes distributions and flattens the result. If you size per tenant, measure per tenant; the pooled number is only right for a consumer that genuinely sees the mixed feed.

saying these in an interview costs you the question

  • Reports entropy in bits per record without saying how many values
  • Multiplies per-value entropy by record count and forgets the eight
  • Multiplies across correlated fields as if they were independent
  • Treats the daily figure as the size a file will actually be
  • Confuses the per-value rate with the field's declared width