A team wants to drop the data-centre column because H(data centre | region) is only 0.54 bits: how do you decide?
answer
- small is not zero
- convert bits into rows first
- residual concentrates in the exception rows
- derivation rules drift over the retention window
- price the pair jointly instead of deleting
basics
~20 sOnly a conditional entropy of zero licenses dropping a column as derivable. At 0.54 bits, roughly one row in eight disagrees with its region's home data centre, and those are exactly the rows an incident review needs.
solid answer
~50 sSmall is not zero, and the gap is the whole decision. `H(data centre | region) = 0.544 bits` in a log where each region routes 7 rows in 8 to its home data centre means 12.5% of rows depart from the mapping, and the residual bits are concentrated there: a departing row carries 3 bits of surprisal against 0.19 for a conforming one, so about 69% of the residual comes from that eighth of the log. Those departures are the failovers, which is precisely the data a postmortem wants. Two further cautions: the number is an average over one measured window, and the region-to-data-centre mapping is an operational policy that changes, so reconstructing old rows with today's rule is not reconstruction. The middle option usually wins: keep the column but price it by the chain rule at about 2.54 bits a row instead of 4.
go deeper
Take away the rule of thumb: a field is only recomputable from another when it is fixed by it in every row. A usually-predictable field is still recording something.
Turn the bits into rows before arguing. Say what share of rows breaks the mapping and what those rows are, rather than quoting an average as if it were a guarantee.
Bring the operational angle: the mapping is a policy that changed inside the retention window, so reconstruction dates wrongly, and the exceptional rows are the ones incident review depends on.
Own the trade: storage saving against the queries the organisation can still answer a year out, with a joint encoding as the option that captures most of the saving without deleting anything.
## What the number licenses, and what it does not `H(Y | X) = 0` is the only value that says **derivable**: every region maps to exactly one data centre, so the second column can be recomputed from the first without loss. Any positive value says the opposite, and 0.544 bits is emphatically positive. The temptation is to read a small average as "nearly derivable" and treat nearly as good enough. It is not, because the decision is not about the average; it is about which rows the residual covers. So the first move is to convert the bits into rows. In this log each region sends 7 rows in 8 to its home data centre and 1 in 8 to a failover, which is what makes `H(7/8, 1/8) = 0.544 bits`. Dropping the column would silently rewrite 12.5% of the log. ## The residual is concentrated, not spread A per-row average hides a very uneven distribution of surprise. | row class | share of rows | surprisal a row | share of the residual | |---|---|---|---| | matches its region's home data centre | 87.5% | 0.19 bits | about 31% | | routed to the failover data centre | 12.5% | 3 bits | about 69% | The second line is the operationally interesting one: failovers, drains, capacity spillover. An average of 0.544 bits a row makes the column look like rounding error; the table shows that most of what the column tells you is carried by the minority of rows you would actually go looking for. **Low residual entropy and high residual value routinely coexist**, because the two are measuring different things. ## The derivation rule is a policy, not a law A derived column is only as good as the rule that derives it, and a region-to-data-centre mapping is an operational decision that changes: a region is re-homed, a data centre is retired, a new one absorbs half of a region. Reconstruction then has a dating problem. - Recomputing old rows with today's mapping produces confidently wrong values, which is worse than a missing column because nothing signals the error. - Keeping a versioned history of the mapping means carrying a second dataset that must itself be correct, retained as long as the log, and joined at read time. - A conditional entropy measured over last month's rows says nothing about next month's, so the justification decays even when nobody edits anything. ## The third option is usually the right one The choice is rarely keep-at-full-width versus delete. The chain rule offers the middle path: describe the region in full and the data centre only as its departure from what the region implies. That prices the pair at `H(region) + H(data centre | region) = 2 + 0.544 = 2.544 bits` a row against 4 bits for two fixed-width fields, a saving of about 36% that loses no row's actual value. A cruder version of the same idea keeps a one-bit flag for "matched the home mapping" plus the actual value only when it did not. Either way, no row is reconstructed from a rule. ## How to decide 1. **Convert bits to rows.** What share of rows departs from the derivation rule? Call the column derivable only when that share is zero. 2. **Ask who reads the departing rows.** Incident review, capacity planning, billing attribution. A consumer of the exceptional rows is a veto on dropping them. 3. **Check the slices, not only the average.** One region in a long drain can hold most of the residual while the aggregate stays low, and a slice with few rows will look deterministic when it is not. 4. **Test the rule's stability over the retention window.** If the mapping changed once in the period you must keep, reconstruction is already wrong. 5. **Price the middle option.** If joint description already gets most of the saving, the residual argument for deleting data is thin. 6. **Decide what you can defend after an incident**, which is the actual question: not the bytes saved, but whether the record can still answer where a request went. ## The framing to reject "The entropy is small, so the column is redundant" mixes an average with a guarantee, and it treats bytes as the only cost in play. The entropy figure is a strong input into how to **store** the pair and a weak input into whether to **keep** it. Use it for the first; decide the second on who reads the exceptional rows.
- What measurement would change your mind?A per-region breakdown showing the departures concentrated in one separately recorded maintenance window, a mapping unchanged across the full retention period, and no consumer that queries the column for the exceptional rows. With all three, dropping it becomes defensible; with any one missing, it is not.
- If you keep both columns, what does the chain rule say they should cost?H(region) + H(data centre | region) = 2 + 0.544, about 2.54 bits a row against 4 bits for two fixed-width fields. You describe the region in full and the data centre only as its departure from the region's implied value, which keeps every row's actual value while capturing most of the saving.
- Why is a confidently reconstructed wrong value worse than a missing column?A missing column announces itself and forces the reader to handle the gap. A value recomputed from a mapping that changed after the rows were written looks exactly like a recorded fact, so it propagates into dashboards, attributions and postmortems with no signal that it was inferred at all.
saying these in an interview costs you the question
- Drops the column because the residual entropy is small on average
- Treats a low conditional entropy as a deterministic derivation rule
- Assumes today's region-to-data-centre mapping reconstructs older rows
- Overlooks that the residual bits carry the failover rows
- Argues purely from bytes saved with no consumer analysis