skip to content

Why can KL divergence between last month's and this month's traffic mix come out infinite?

level: seniorimportance: should knowfreq 40%

answer

  1. a brand-new category breaks it
  2. zero in the denominator of the log ratio
  3. empty bins are an artefact of the window
  4. add a pseudo-count to every bin
  5. the symmetric mixture version stays bounded

basics

~20 s

KL is infinite whenever the current month has mass in a category the reference month gives zero probability. A brand-new channel or an empty bin makes the log ratio diverge, so the score explodes for a reason unrelated to drift size.

solid answer

~50 s

`KL(P||Q) = sum_x P(x) log(P(x)/Q(x))` is infinite as soon as some category has `P(x) > 0` and `Q(x) = 0`. In drift monitoring that is routine: a channel that launched this month has zero mass in last month's reference, or a rare bin picked up no traffic in one of the two windows. The score then reports infinity even though the practical shift may be tiny. Three fixes, in order of preference. Smooth both distributions — add a small pseudo-count to every category, so no bin is exactly zero — and hold the category list and its ordering fixed across periods. Or switch to Jensen-Shannon, `0.5*KL(P||M) + 0.5*KL(Q||M)` with `M` the mixture, which is symmetric and bounded by 1 bit, so it degrades gracefully instead of exploding. Either way, fix the direction and the binning before anyone sets an alert threshold.

go deeper

for a junior

Be ready to say KL goes infinite when one distribution assigns zero probability to something the other actually produces, such as a channel that only exists this month.

for a middle

Explain the mechanics of the fixes: a fixed shared category list, pseudo-count smoothing before normalising, and the mixture construction that keeps Jensen-Shannon finite and bounded.

for a senior

Show you have operated this: frozen bin edges, an explicit other bucket, a threshold calibrated from historical period-over-period scores, and per-category contributions attached to every alert.

for a principal

Own the policy question: decide what drift is worth paging for, who owns the response, and resist a monitoring metric becoming a target that teams optimise instead of the downstream outcome it proxies.

## Where the infinity comes from Suppose you summarise traffic mix as a distribution over acquisition channels, compute last month's shares as `Q` (the reference) and this month's as `P` (the current window), and score drift with ``` KL(P||Q) = sum_x P(x) * log( P(x) / Q(x) ) ``` The convention is that a term with `P(x) = 0` contributes 0, but a term with `P(x) > 0` and `Q(x) = 0` contributes `+infinity`. In words: this month produced traffic from a channel last month says is impossible. The formula does not know the difference between "impossible" and "we did not happen to see it", so a single new channel with 0.3% of traffic sends the whole score to infinity, next to which a genuine 20-point shift in the main channels registers as nothing. ## The three ways this bites in practice **New or retired categories.** A campaign launches, a partner is switched off, a new device type appears. Any category present in one window and absent in the other creates a zero in one of the two vectors. **Sparse bins from small windows.** Even with a stable category list, a low-share category can draw zero events in a short window purely by sampling noise. The zero is an estimation artefact, not a property of the underlying distribution, and it is more likely the shorter the window and the longer the tail. **Binning a continuous variable.** If you discretise session duration or order value into bins, the edge bins routinely go empty in one period. Worse, if bin edges are recomputed per period from that period's quantiles, the two distributions are no longer defined on comparable outcomes at all, and the number is meaningless whether or not it is finite. ## Fixes **Fix the outcome space first.** Both distributions must be over the same fixed, ordered list of categories or bins, with edges derived once from the reference window and frozen. Everything unseen or new goes into an explicit "other" bucket rather than silently changing the alphabet. This is a prerequisite, not an optimisation: a KL between distributions on different supports is not a comparison. **Smooth.** Add a small pseudo-count to every category in both vectors before normalising — for example, add `a` to each of `k` categories with `n` observations, giving `(count + a) / (n + k*a)`. No bin is then exactly zero, and KL is finite. The cost is that the score now depends on `a`: a large pseudo-count pulls both distributions toward uniform and damps real drift, a tiny one leaves the score dominated by whichever rare bin happened to be empty. Pick `a` once, document it, and keep it fixed, because changing it silently rescales every historical value. **Use a bounded symmetric divergence.** The Jensen-Shannon divergence ``` JS(P,Q) = 0.5*KL(P||M) + 0.5*KL(Q||M), M = (P + Q)/2 ``` is finite by construction: the mixture `M` has mass wherever either distribution does, so no log ratio blows up. It is symmetric, so you no longer have to defend a direction, and in base 2 it lies in `[0, 1]` bits, with 1 attained only when the two distributions have disjoint support. A bounded score is what makes a fixed alert threshold meaningful — you can say "page at 0.15 bits" and have that mean the same thing every month. For the fair-versus-biased pair `P = (0.5, 0.5)`, `Q = (0.9, 0.1)`, JS is about 0.147 bits while the two KL directions are 0.737 and 0.531 bits — the same ordering of severity, on a scale with a ceiling. ## Judgment beyond the infinity **Direction is a decision.** `KL(current||reference)` penalises the current window producing things the reference thought rare; `KL(reference||current)` penalises the reverse. They are different numbers with different alarm behaviour. Choose one, write down why, and never quietly swap it. **A divergence score is a detector, not a diagnosis.** The aggregate says something moved; the per-category contributions `P(x) log(P(x)/Q(x))` say what. Always surface the top contributing categories with the score, or an on-call engineer has an alert with nowhere to go. **Volume changes the noise floor.** Two windows with different traffic volumes give estimates of different precision, so a raw threshold will fire more often on the smaller window. Either fix the window size, or calibrate the threshold from the historical distribution of period-over-period scores rather than picking a round number. **Not all drift matters.** A mix shift that leaves downstream behaviour unchanged is a fact, not an incident. The divergence number should route attention, and the decision to act belongs to whoever owns the downstream metric.

  • Does smoothing the two distributions fully solve the problem?
    It removes the infinity but introduces a tuning knob. The pseudo-count you add sets how much a rare or newly-seen category can move the score: too large and both distributions are pulled toward uniform, damping real drift; too small and the score is dominated by whichever tail bin was empty. Choose it once, document it, and keep it fixed, because changing it retroactively rescales every historical value and invalidates the threshold.
  • Why is Jensen-Shannon easier to alert on than KL?
    It is bounded. In base 2, `JS` lies between 0 and 1 bit, hitting 1 only for distributions with disjoint support, so a fixed threshold means the same thing every period. It is also symmetric and always finite, so you never have to defend a direction or handle an infinite value. KL is unbounded, so any threshold you pick is scale-free only until an unusual month arrives.
  • The score is high. How do you turn that into an actionable diagnosis?
    Decompose it. The total is a sum of per-category terms `P(x) log(P(x)/Q(x))`, so rank the categories by contribution and report the top few alongside the score. That usually points straight at a launched campaign, a retired source, or a tracking change. Then check whether the shifted categories actually differ in downstream behaviour, because a mix shift with no downstream effect is an observation, not an incident.

saying these in an interview costs you the question

  • Recomputes bin edges separately for each period
  • Treats a zero count as a genuine zero probability
  • Alerts on a raw KL threshold with no bound or calibration
  • Silently swaps the direction between reference and current
  • Reports an aggregate score with no per-category breakdown

context