skip to content

Conditioning never raises entropy on average, yet one region's data-centre uncertainty exceeds the whole column's: how is that consistent?

level: seniorimportance: should knowfreq 34%

answer

  1. the inequality binds an average
  2. slices are individually unconstrained
  3. a bad slice must be a rare slice
  4. weighting by slice share saves it
  5. equality means independence, not weak dependence

basics

~20 s

H(Y|X) <= H(Y) constrains the frequency-weighted average across all conditioning values, not each one. A single region can leave more data-centre uncertainty than the pooled column; if it is rare enough, the average still cannot rise.

solid answer

~40 s

The inequality `H(Y | X) <= H(Y)` is about the average, because `H(Y | X)` *is* an average: the per-region entropies weighted by how often each region occurs. Individual slices are unconstrained. Take a log where 99% of rows come from one region that always routes to the same data centre, and 1% come from a region under maintenance whose traffic is spread evenly over four data centres. The pooled data-centre column is almost constant, about 0.076 bits. The maintenance region's slice is 2 bits, far above it. Yet H(data centre | region) = 0.99 * 0 + 0.01 * 2 = 0.02 bits, comfortably under 0.076. The weighting is what saves the inequality, and it is also why a per-slice breakdown can look alarming while the aggregate looks calm.

go deeper

for a junior

Hold on to the headline and its qualifier together: knowing a related field does not increase uncertainty on average. The words 'on average' are part of the statement, not a hedge.

for a middle

Explain why the average is safe: the conditional entropy is a share-weighted mean of per-slice entropies, so a slice that is worse than the pooled column has to be rare.

for a senior

Show it in a breakdown: quote an aggregate and a per-slice figure that disagree, say which decision each supports, and flag slices too small to estimate honestly.

for a principal

Decide what gets reported. An aggregate that hides a pathological slice is a monitoring choice, and it is your call which cut the organisation is held to.

## What the inequality actually quantifies over The standard statement is that **conditioning reduces entropy**: `H(Y | X) <= H(Y)`, with equality exactly when the two columns are independent. The word doing the work is hidden in the notation. `H(Y | X)` is not one number about one situation; it is `sum over x of p(x) * H(Y | X = x)`, a **frequency-weighted average** of per-slice entropies. The inequality binds that average. It says nothing whatever about an individual slice `H(Y | X = x)`, which may sit anywhere from 0 up to the logarithm of the number of distinct values. Informally: learning the region cannot hurt you *on expectation*, but a particular region can absolutely be bad news. ## A log where one value makes things worse Rows record the **region** and the **data centre** the request routed to. - 99% of rows come from a steady region, which routes every request to one data centre. - 1% of rows come from a region under maintenance, whose requests are spread evenly across four data centres. Work the numbers: | quantity | value | |---|---| | data centre given the steady region | 0 bits | | data centre given the maintenance region | 2 bits | | data-centre column pooled over all rows | about 0.076 bits | | data centre given region, weighted | 0.99 * 0 + 0.01 * 2 = 0.02 bits | The pooled column is nearly constant because the steady region dominates: the busiest data centre holds 99.25% of rows and the other three 0.25% each, which is 0.076 bits. One slice is 2 bits, more than twenty-five times the pooled figure, and the average is still 0.02 bits. No rule was broken. ## Why the average survives The weighting forces a bad slice to be rare. Since every term in the sum is non-negative, `p(x) * H(Y | X = x) <= H(Y | X) <= H(Y)`, so any slice whose entropy is far above `H(Y)` must occupy a correspondingly small share of rows. In the example, a 2-bit slice cannot hold more than about 3.8% of the log while `H(Y)` stays at 0.076 bits. The mechanism behind the inequality itself is concavity: the entropy of an averaged distribution is at least the average of the entropies, and the pooled column is precisely the average of the slices. Two useful corollaries, both frequently mis-stated: 1. Conditioning on **more** columns cannot raise the average either: `H(Y | X, Z) <= H(Y | X)`. Again, on average only. 2. Equality throughout means independence, not "weak dependence". Any cell of the joint table that departs from the product of its margins makes the inequality strict. ## Reading a per-slice breakdown at work The distinction is not pedantry; it is how you read a breakdown. An aggregate conditional entropy of 0.02 bits says the region column predicts the data centre extremely well overall, and it would justify modelling the pair as nearly deterministic. The per-slice view says one region is behaving completely unpredictably. Both are true. The practical rules that follow: - **Report the slice, not only the average**, whenever slices differ wildly in size; a rare slice is invisible in the aggregate by construction. - **Beware small slices.** A slice with three rows estimates its entropy badly and almost always downward, so a slice can also look deterministic when it is not. - **Do not average slices unweighted.** An unweighted mean of per-region entropies is not `H(Y | X)` and obeys no such inequality; it can easily exceed `H(Y)` and mislead. ## The statements to avoid - "Conditioning always reduces uncertainty." It never *increases* it on average, and it may leave it unchanged; for a specific value it can increase it. - "If a slice is worse than the pooled column, the table is inconsistent." It is not; it only means that slice is rare. - "A low conditional entropy means every slice is predictable." It means the unpredictable slices are small. The honest one-sentence form is: conditioning never increases entropy **on average**, and any single conditioning value may still leave you worse off than knowing nothing.

  • How rare must a slice be if its entropy exceeds the pooled column's?
    Bounded directly by the average. Every term is non-negative, so p(x) times that slice's entropy cannot exceed H(Y|X), which cannot exceed H(Y). With a pooled entropy of 0.076 bits and a slice at 2 bits, the slice holds at most about 3.8% of the rows.
  • Does conditioning on a second column also never increase the average?
    Correct: H(Y | X, Z) <= H(Y | X), by the same argument applied within each slice. The cost is practical rather than theoretical, since conditioning on two columns needs one distribution per observed combination, and those combinations get sparse fast, which biases the estimate downward.
  • What does equality in the inequality tell you?
    That the two columns are independent: every slice's distribution matches the pooled one, so knowing the first column changes nothing about the second. Equality is exact and fragile. A single joint cell that departs from the product of its margins already makes the inequality strict.

A shelf label usually narrows your search, but one badly organised shelf can leave more places to look than the room's average. You land on it rarely enough that labels still pay off overall.

saying these in an interview costs you the question

  • Claims conditioning always lowers uncertainty for every individual value
  • Calls a joint table inconsistent because one slice beats the pooled entropy
  • Averages per-slice entropies without weighting by slice share
  • Reads a low aggregate as meaning every slice is predictable
  • Treats equality in the inequality as merely weak dependence