How do you turn a field's per-symbol Shannon entropy into bits per record and a daily volume figure?
answer
- it is a rate, not a total
- say bits per what
- multiply by values per record
- independence licenses the multiplying
- bits to bytes: divide by eight
basics
~20 sShannon entropy is a rate: bits per value. Multiply by the number of values a record carries for bits per record, then by records per day and divide by eight for bytes — valid while the values are independent draws.
solid answer
~40 sEntropy comes out of the formula as **bits per value**, so the whole exercise is keeping track of the denominator. If each record carries one status value at 0.557 bits, then a feed of 200 million records a day carries about 111 million bits, or roughly 14 MB of status content. If each record carries eight such values, independently drawn from the same distribution, that is `8 * 0.557 = 4.46` bits per record and about 111 MB a day. Compare that with a fixed one-byte code per value — 1.6 GB a day — and the gap is the headroom. The multiplication is only licensed by **independence**: if the values within a record are correlated, the product overstates the true content and you need the two-variable machinery to say by how much.
go deeper
Hold on to the unit: entropy comes out as bits per value. Multiplying to a per-record or per-day figure means multiplying by counts and dividing by eight to reach bytes.
Carry the arithmetic end to end and name the assumption. Independent, identically distributed values let you multiply the rate by the count; correlated values do not, and the product then runs high rather than low.
In a capacity review, separate measured content from what the current representation spends, quote the ratio rather than the absolute, and say which way the estimate errs so nobody plans against a bound they think is exact.
Decide whether the headroom is worth spending engineering on at all. A fourteen-to-one gap on one field of one feed may still lose to the operational cost of a cleverer representation that everything downstream must understand.
## Entropy is a rate, so name the denominator The formula `H = -sum p log2 p` produces a number in **bits per value of one field**. Almost every mistake in this area is a denominator mistake: quoting "the field is 0.56 bits" and then treating that as a per-record or per-day quantity. Three different numbers are in play in a sizing conversation, and they differ by factors that matter: - **bits per value** — what the formula gives you directly; - **bits per record** — the per-value rate times the number of values of that field a record carries; - **bytes per day** — bits per record times records per day, divided by 8. ## Multiplying out one field for one day Take an ingestion pipeline at 200 million records a day, with a status field measured at 0.557 bits per value. 1. **One value per record.** `200,000,000 * 0.557 = 111.4` million bits. Divided by 8, about **13.9 MB** of status content per day. 2. **Eight values per record** — say a per-stage outcome for each stage of the pipeline, each drawn independently from the same distribution. Per record: `8 * 0.557 = 4.46` bits. Per day: `8 * 111.4 = 891` million bits, about **111 MB**. 3. **What a naive representation costs.** One byte per value is `8 * 200,000,000 = 1.6` billion bytes, about **1.6 GB** a day. | Quantity | Figure | What it tells you | |---|---|---| | Entropy per value | 0.557 bits | Average content of one value | | Bits per record, 8 values | 4.46 bits | Content of that field group in one record | | Per day at 200M records | about 111 MB | The content-side budget | | Fixed one byte per value | about 1.6 GB | What an indifferent representation costs | The ratio, roughly 14 to 1, is the number the conversation is actually about. **Whether any particular encoding can reach the content figure is a separate result about code lengths, and the entropy number by itself does not assert it** — but it does tell you the size of the prize before anyone builds anything. ## Where the multiplication stops being valid Multiplying a per-value rate by a count assumes each value is an **independent draw from the same distribution**. Two ways that breaks in real feeds: - **Correlation within a record.** If a failure at stage three makes a failure at stage four near-certain, the eight values are not eight independent draws, and their combined content is *below* `8 * 0.557` bits. The sum is then an upper bound, and measuring the overlap needs the two-variable forms rather than the single-field figure you started with. - **A distribution that is not the same everywhere.** If the status mix differs by tenant, by hour, or by client version, one pooled `H` describes a blend that matches no individual stream. Pooling across heterogeneous sources generally inflates the figure relative to any one source. Both failures point the same way: **the product is an upper bound on the real content, never an underestimate.** That is the useful direction for capacity work, as long as you say so out loud. ## What the daily figure is and is not It is a statement about the **value distribution**, not a prediction of bytes on disk. A feed measuring 14 MB a day of status content may well write 200 MB, because a fixed-width representation spends the same room on every record regardless of how predictable that record's value was. The gap is not a bug; it is the headroom, and whether it is worth taking depends on cost, on query patterns and on how stable the distribution is. It also says nothing about the *rest* of the record. Sizing one field is exactly that: a per-field statement. Adding up all fields is another multiplication with the same independence caveat, and in real schemas fields correlate heavily. ## Doing it in a review - Quote three numbers, always together: the rate, the multiplier, and the period. - Convert to bytes explicitly; the factor of 8 is the most common slip. - State the independence assumption when you multiply, and say which direction the error runs. - Compare against what the field costs today, not against zero — the interesting quantity is the ratio, not the absolute figure.
- Two fields in the same record are strongly correlated — what does summing their Shannon entropies give you?An upper bound on their combined content, not the true figure. Summing is exact only when the fields are independent; correlation means part of the second field is already implied by the first, so the truth sits below the sum. Quantifying that overlap requires the two-variable forms, a different measure from single-field entropy.
- Your daily figure says 14 MB but the feed actually writes 200 MB for that field — is the measurement wrong?Not necessarily. Entropy measures the content of the value distribution; the 200 MB is what the chosen representation costs, and a fixed-width code spends the same room on every record however predictable it is. The gap is headroom. Whether taking it is worthwhile is a separate engineering question.
- The status mix differs sharply between tenants — what does one pooled entropy figure describe?A blend that matches no individual tenant, and generally a higher figure than any single stream, because pooling mixes distributions and flattens the result. If you size per tenant, measure per tenant; the pooled number is only right for a consumer that genuinely sees the mixed feed.
saying these in an interview costs you the question
- Reports entropy in bits per record without saying how many values
- Multiplies per-value entropy by record count and forgets the eight
- Multiplies across correlated fields as if they were independent
- Treats the daily figure as the size a file will actually be
- Confuses the per-value rate with the field's declared width