In information terms, an alert that fires in one minute out of 1024 carries ten bits of surprisal - what does that number mean?
answer
- rarity, not how often it fires
- one fair coin toss is one bit
- the log base fixes the unit
- each halving of probability adds one
- 1/1024 is ten halvings from certainty
basics
~20 sTen bits is the surprisal of an outcome with probability 1/1024, since -log2(1/1024) = 10. One bit is the uncertainty resolved by one fair yes/no answer, so rarer outcomes carry more bits and a certain one carries zero.
solid answer
~40 sSurprisal, or self-information, of a single outcome is `-log2 p`: the number of bits of uncertainty that learning the outcome destroys. An alert with `p = 1/1024` per minute scores `-log2(2^-10) = 10` bits, because 1/1024 is ten successive halvings away from certainty and **each halving of probability adds exactly one bit**. One bit is the information in one fair coin toss, or in the answer to one well-chosen yes/no question: with 1024 equally likely candidates, ten yes/no questions pin down one. The scale runs opposite to how often something happens - a routine event carries very few bits, an event that always happens carries exactly zero, and as `p` falls the bits grow without bound. It is a property of one outcome under an assumed probability model, not of the alert's message size.
go deeper
Recall the formula and the unit: surprisal is -log2 p, one bit is one fair yes/no answer, and a certain outcome is worth zero. Being able to say why 1/1024 gives ten is enough at this stage.
Explain the mechanics: halvings of probability, why the value grows without bound as the probability falls, and why fractional bits are normal. Separate the measure from the encoded size of a message.
Show the judgment: state where the probability came from, say what happens to every derived figure when that rate drifts, and refuse to equate a high bit count with a high-severity event.
Frame the tradeoff: a bit figure is only as good as the model behind it, so decide how rates get published, reviewed and re-measured before letting anyone route work on the strength of those numbers.
## Surprisal: the information in one outcome **Surprisal** (also called **self-information**) attaches a number to a *single* outcome under an assumed probability model. If an outcome has probability `p`, its surprisal is `I = -log2 p`, measured in **bits**. It answers exactly one question: how much uncertainty does learning that this outcome occurred destroy? Three properties fall straight out of the formula: - It is **decreasing in p**. The rarer the outcome, the larger the number. - It is **zero at p = 1**. An outcome that was guaranteed tells you nothing when it arrives. - It **grows without bound as p falls toward zero**. There is no ceiling on how many bits one outcome can carry. ## Why one in 1024 is exactly ten bits Because 1/1024 = 2^-10, and `-log2(2^-10) = 10`. The arithmetic-free way to read it is to count halvings from certainty: start at `p = 1` (0 bits) and halve ten times - 1/2, 1/4, 1/8, ... 1/1024. Each halving is worth one bit. | probability of the outcome | surprisal in bits | read it as | |---|---|---| | 1 | 0 | certain; nothing is learned | | 1/2 | 1 | one fair yes/no answer | | 1/4 | 2 | two fair yes/no answers | | 1/8 | 3 | three halvings from certainty | | 1/32 | 5 | five halvings from certainty | | 1/100 | 6.64 | bits need not be whole numbers | | 1/1024 | 10 | the alert in this question | | 1/1,000,000 | 19.93 | a genuinely rare event | ## One bit is one fair yes/no answer The unit is *defined* by the fair coin: an outcome with probability 1/2 carries one bit. The twenty-questions framing makes the same point operationally. With 1024 equally likely candidates, a well-chosen yes/no question halves the candidate set, so ten questions cut 1024 down to one. Saying an outcome carries ten bits and saying it takes ten ideal yes/no answers to single it out are the same statement. Non-integer values are normal and not a defect: 6.64 bits simply means no whole number of fair yes/no answers matches that outcome exactly. Bits are a continuous scale, not a count of questions you must actually ask. ## The probability model is an input, not a fact Surprisal is always relative to an assumed `p`. If the published rate says an alert type fires once in 1024 minutes but the real rate is once in eight, the ten-bit figure is fiction - the honest figure is three bits. Two habits follow: 1. State where `p` came from (a published rate, a measured rate, an assumption). 2. Recheck it when behaviour drifts; a stale `p` silently corrupts every bit figure derived from it. ## What the number is not - **Not the size of the alert payload.** Surprisal is a property of the outcome's probability. How many bits an encoder actually writes on the wire is a separate question about code design. - **Not severity.** A ten-bit event may be harmless and a one-bit event may be an outage. Information is not cost. - **Not a count of firings.** Ten bits describes one occurrence's rarity, not ten occurrences. - **Not an average.** Averaging surprisal across every outcome of a distribution yields a different quantity with a different use; this measure is about one outcome. - **Not linear in the probability.** A measure like `1 - p` saturates near one and cannot distinguish one-in-a-thousand from one-in-a-million. ## Reading it on a stream of signals On an on-call stream where each signal type has a published rate, the procedure is short: 1. Fix an observation window (say one minute) and write down each signal's probability of firing in that window. 2. Convert to bits by counting halvings, or by `-log2 p` when the number is not a power of two. 3. Compare. A signal whose firing scores a hundredth of a bit is telling you something you already knew; a signal whose firing scores ten bits genuinely changed your picture of the system. That comparison is the whole payoff of the measure at this level. It gives you a single number, on a common scale, for how much any one observation actually told you - independent of how loud it was, how big its message was, or how strongly someone felt about it.
- What is the surprisal of an outcome that is certain to occur, and why?Zero bits. `-log2(1) = 0`, and that matches the meaning: if you already knew the outcome, learning it resolves no uncertainty. It is the anchor of the whole scale - every other outcome is some number of halvings away from this point.
- Does a ten-bit alert mean the alert needs ten bits on the wire?No. Surprisal measures how much uncertainty the occurrence resolves under an assumed probability model. How many bits a particular encoder writes for that alert is a separate question about code design, and a real encoder may use far more or, over many messages, approach fewer.
- Can one outcome carry a fractional number of bits?Yes, and most do. An outcome with `p = 1/100` carries `log2(100) = 6.64` bits. Only probabilities that are exact powers of one half land on whole numbers. Fractional bits are meaningful as a measure even though you cannot ask 0.64 of a yes/no question.
It is the twenty-questions score. With 1024 equally likely candidates, ten well-chosen yes/no answers identify one, and that is precisely what ten bits means.
saying these in an interview costs you the question
- Says a frequently firing signal carries more information than a rare one.
- Treats the bits as the storage size of the alert message.
- Claims a guaranteed outcome still carries some small amount of information.
- Reads ten bits as ten firings rather than one outcome's rarity.
- Uses 1 - p as the measure, which caps information at one.
- Assumes surprisal must always be a whole number of bits.