In surprisal terms, one signal fires in 999 minutes out of 1000 and another in one minute out of 1024 - which should page a human, and why?
answer
- certainty carries nothing
- the expected arrival is not news
- the missing signal is the news
- bits live on the low-probability side
- informative is not the same as severe
basics
~20 sThe rare signal, at about 10 bits per firing, carries real information; the near-certain one carries about 0.0014 bits and confirms what you already knew. For the near-certain signal the informative event is its absence, worth about 9.97 bits.
solid answer
~50 sCompute both. The near-certain signal fires with `p = 0.999`, so its surprisal is `-log2(0.999)`, about **0.0014 bits** - its arrival resolves essentially no uncertainty. The rare one fires with `p = 1/1024`, so it carries **10 bits**. On information grounds the rare firing is the one that changed your picture of the system. For the near-certain signal the useful trigger is the complement: a missing minute has `p = 0.001` and carries `log2(1000)`, about **9.97 bits**. So page on the rare firing and on the *absence* of the routine signal, not on its presence. One caveat worth saying out loud: surprisal measures information, not cost. A rare but harmless event is highly informative and still not worth waking someone for, so rarity is a necessary input to the routing decision, not the whole of it.
go deeper
Recall the direction of the scale: the more likely an outcome was, the fewer bits its occurrence carries, and a near-certain arrival is worth almost nothing.
Do the arithmetic both ways: compute -log2 p for the arrival and for the absence, and show that for a near-certain signal all the bits sit on the absence side.
Show operating judgment: trigger on the informative outcome, say where the probability came from, and recompute when rates drift so a once-useful trigger is retired rather than left firing.
Own the tradeoff: information content ranks triggers, expected impact ranks consequences, and the routing policy has to combine both explicitly rather than letting severity labels stand in for measurement.
## Put numbers on both signals first Surprisal is `-log2 p` for a single outcome, so the whole comparison is four small calculations - one for each signal and one for each complement. | observation | probability in a given minute | surprisal in bits | |---|---|---| | routine signal present | 0.999 | 0.0014 | | routine signal missing | 0.001 | 9.97 | | rare signal fires | 1/1024 | 10 | | rare signal silent | 1023/1024 | 0.0014 | Two things jump out. First, the routine signal's arrival is worth roughly a **seven-hundredth of a bit** - it resolves almost nothing, because you would have bet on it at 999 to 1 anyway. Second, the informative half of a near-certain signal is its **complement**: the missing minute carries about ten bits, essentially the same figure as the rare signal's firing. ## The information lives in the rare outcome This is the general shape of the result, and it is worth stating as a rule rather than an anecdote: - A signal that almost always fires tells you something only when it **fails** to fire. - A signal that almost never fires tells you something when it **does** fire. - In both cases the bits sit on the low-probability side, because `-log2 p` is large only where `p` is small. That is why a health check that emits a steady stream of confirmations is not, by itself, a source of information; the pipeline that watches for a *gap* in that stream is. The confirmations are the carrier; the gap is the news. ## What this actually decides Applied to routing, the measure gives you a defensible ordering rather than a gut feeling: 1. **Estimate the probability per observation window** for every trigger you are considering, from published or measured rates. 2. **Convert to bits.** Anything scoring a small fraction of a bit is confirming a prior, not reporting news. 3. **Invert the near-certain triggers.** Where the fire-side score is near zero, the absence-side score is large; trigger on the absence. 4. **Re-rank when the rates move.** A trigger that was informative at one rate becomes noise once the underlying event becomes routine. Step 4 is the one teams skip. A trigger's information content is a function of the current rate, so a signal that fired once a week and carried real bits becomes near-worthless once a change makes it fire every minute - and nothing in the alerting configuration notices. The same arithmetic that justified the alert originally condemns it afterwards. ## Information is not severity, and not volume Three separations keep this honest: - **Information is not cost.** A one-in-a-million event may be informative and harmless; a routine one may be catastrophic. Surprisal ranks how much an observation told you, and the decision to wake a human must also weigh what it costs to ignore. - **Information is not volume.** A signal firing sixty times an hour does not carry more information per firing than one firing once; if anything, higher frequency means a higher probability and therefore *fewer* bits per occurrence. - **Information is not loudness.** Severity labels are a human annotation. They are worth having, but they are not a measurement of anything, and they drift. The interviewer is usually probing exactly this boundary: a candidate who says "page on the rare one because rare means important" has skipped a step. Rare means *informative*. Importance is a second axis you must supply. ## When the model is wrong the conclusion is wrong Every number above rests on an assumed probability. If the routine signal's true availability is 0.9 rather than 0.999, its presence carries `-log2(0.9)` = about 0.15 bits and its absence about 3.32 bits - still asymmetric, but a tenth of the previously claimed gap, and a trigger on absence would now fire far more often than the earlier arithmetic suggested. Two practical habits follow: measure the rate rather than inheriting it from a document, and recompute when behaviour changes. Publishing a bit figure without the rate and the window it came from is publishing a number nobody can check. ## The answer an interviewer wants to hear A complete answer does four things in about a minute: puts `-log2 p` on both signals, notices that the near-certain one scores almost zero, flips it to its complement and notices the complement scores about ten bits, and then explicitly separates informativeness from severity before making the routing call. Averaging these per-outcome values across a whole distribution is a different measure with a different purpose and is not what this question is asking for.
- If the near-certain signal is missing for one minute, how many bits does that carry?About 9.97 bits. Its absence has probability 0.001, and `-log2(0.001) = log2(1000)`, which is 9.9658. That is why watching for the gap in a steady stream is worth far more than watching the stream itself.
- Does a rare event always deserve a page?No. Surprisal ranks how much an observation told you, not what it costs you. A rare but harmless event scores high on information and still should not wake anyone. Rarity is one input to the routing decision; expected impact is the other.
- The rare signal starts firing every minute after a deployment. What happens to its information content?It collapses. At a probability near 1 its firing carries a small fraction of a bit, so the trigger that once carried ten bits now confirms a prior. Nothing in the alert definition changes - only the rate does - which is why bit figures must be recomputed when behaviour drifts.
saying these in an interview costs you the question
- Pages on the constant signal because it fires most often.
- Says a 99.9 percent likely arrival still carries about one bit.
- Treats the absence of a routine signal as carrying no information.
- Confuses how severe an event is with how informative it is.
- Assumes more firings per hour means more information per firing.
- Keeps using a bit figure after the underlying rate has changed.