skip to content

questions

4

In information terms, an alert that fires in one minute out of 1024 carries ten bits of surprisal - what does that number mean?

level: juniorimportance: must knowfreq 62%

answer

  1. rarity, not how often it fires
  2. one fair coin toss is one bit
  3. the log base fixes the unit
  4. each halving of probability adds one
  5. 1/1024 is ten halvings from certainty

basics

~20 s

Ten bits is the surprisal of an outcome with probability 1/1024, since -log2(1/1024) = 10. One bit is the uncertainty resolved by one fair yes/no answer, so rarer outcomes carry more bits and a certain one carries zero.

solid answer

~40 s

Surprisal, or self-information, of a single outcome is `-log2 p`: the number of bits of uncertainty that learning the outcome destroys. An alert with `p = 1/1024` per minute scores `-log2(2^-10) = 10` bits, because 1/1024 is ten successive halvings away from certainty and **each halving of probability adds exactly one bit**. One bit is the information in one fair coin toss, or in the answer to one well-chosen yes/no question: with 1024 equally likely candidates, ten yes/no questions pin down one. The scale runs opposite to how often something happens - a routine event carries very few bits, an event that always happens carries exactly zero, and as `p` falls the bits grow without bound. It is a property of one outcome under an assumed probability model, not of the alert's message size.

go deeper

for a junior

Recall the formula and the unit: surprisal is -log2 p, one bit is one fair yes/no answer, and a certain outcome is worth zero. Being able to say why 1/1024 gives ten is enough at this stage.

for a middle

Explain the mechanics: halvings of probability, why the value grows without bound as the probability falls, and why fractional bits are normal. Separate the measure from the encoded size of a message.

for a senior

Show the judgment: state where the probability came from, say what happens to every derived figure when that rate drifts, and refuse to equate a high bit count with a high-severity event.

for a principal

Frame the tradeoff: a bit figure is only as good as the model behind it, so decide how rates get published, reviewed and re-measured before letting anyone route work on the strength of those numbers.

## Surprisal: the information in one outcome **Surprisal** (also called **self-information**) attaches a number to a *single* outcome under an assumed probability model. If an outcome has probability `p`, its surprisal is `I = -log2 p`, measured in **bits**. It answers exactly one question: how much uncertainty does learning that this outcome occurred destroy? Three properties fall straight out of the formula: - It is **decreasing in p**. The rarer the outcome, the larger the number. - It is **zero at p = 1**. An outcome that was guaranteed tells you nothing when it arrives. - It **grows without bound as p falls toward zero**. There is no ceiling on how many bits one outcome can carry. ## Why one in 1024 is exactly ten bits Because 1/1024 = 2^-10, and `-log2(2^-10) = 10`. The arithmetic-free way to read it is to count halvings from certainty: start at `p = 1` (0 bits) and halve ten times - 1/2, 1/4, 1/8, ... 1/1024. Each halving is worth one bit. | probability of the outcome | surprisal in bits | read it as | |---|---|---| | 1 | 0 | certain; nothing is learned | | 1/2 | 1 | one fair yes/no answer | | 1/4 | 2 | two fair yes/no answers | | 1/8 | 3 | three halvings from certainty | | 1/32 | 5 | five halvings from certainty | | 1/100 | 6.64 | bits need not be whole numbers | | 1/1024 | 10 | the alert in this question | | 1/1,000,000 | 19.93 | a genuinely rare event | ## One bit is one fair yes/no answer The unit is *defined* by the fair coin: an outcome with probability 1/2 carries one bit. The twenty-questions framing makes the same point operationally. With 1024 equally likely candidates, a well-chosen yes/no question halves the candidate set, so ten questions cut 1024 down to one. Saying an outcome carries ten bits and saying it takes ten ideal yes/no answers to single it out are the same statement. Non-integer values are normal and not a defect: 6.64 bits simply means no whole number of fair yes/no answers matches that outcome exactly. Bits are a continuous scale, not a count of questions you must actually ask. ## The probability model is an input, not a fact Surprisal is always relative to an assumed `p`. If the published rate says an alert type fires once in 1024 minutes but the real rate is once in eight, the ten-bit figure is fiction - the honest figure is three bits. Two habits follow: 1. State where `p` came from (a published rate, a measured rate, an assumption). 2. Recheck it when behaviour drifts; a stale `p` silently corrupts every bit figure derived from it. ## What the number is not - **Not the size of the alert payload.** Surprisal is a property of the outcome's probability. How many bits an encoder actually writes on the wire is a separate question about code design. - **Not severity.** A ten-bit event may be harmless and a one-bit event may be an outage. Information is not cost. - **Not a count of firings.** Ten bits describes one occurrence's rarity, not ten occurrences. - **Not an average.** Averaging surprisal across every outcome of a distribution yields a different quantity with a different use; this measure is about one outcome. - **Not linear in the probability.** A measure like `1 - p` saturates near one and cannot distinguish one-in-a-thousand from one-in-a-million. ## Reading it on a stream of signals On an on-call stream where each signal type has a published rate, the procedure is short: 1. Fix an observation window (say one minute) and write down each signal's probability of firing in that window. 2. Convert to bits by counting halvings, or by `-log2 p` when the number is not a power of two. 3. Compare. A signal whose firing scores a hundredth of a bit is telling you something you already knew; a signal whose firing scores ten bits genuinely changed your picture of the system. That comparison is the whole payoff of the measure at this level. It gives you a single number, on a common scale, for how much any one observation actually told you - independent of how loud it was, how big its message was, or how strongly someone felt about it.

  • What is the surprisal of an outcome that is certain to occur, and why?
    Zero bits. `-log2(1) = 0`, and that matches the meaning: if you already knew the outcome, learning it resolves no uncertainty. It is the anchor of the whole scale - every other outcome is some number of halvings away from this point.
  • Does a ten-bit alert mean the alert needs ten bits on the wire?
    No. Surprisal measures how much uncertainty the occurrence resolves under an assumed probability model. How many bits a particular encoder writes for that alert is a separate question about code design, and a real encoder may use far more or, over many messages, approach fewer.
  • Can one outcome carry a fractional number of bits?
    Yes, and most do. An outcome with `p = 1/100` carries `log2(100) = 6.64` bits. Only probabilities that are exact powers of one half land on whole numbers. Fractional bits are meaningful as a measure even though you cannot ask 0.64 of a yes/no question.

It is the twenty-questions score. With 1024 equally likely candidates, ten well-chosen yes/no answers identify one, and that is precisely what ten bits means.

saying these in an interview costs you the question

  • Says a frequently firing signal carries more information than a rare one.
  • Treats the bits as the storage size of the alert message.
  • Claims a guaranteed outcome still carries some small amount of information.
  • Reads ten bits as ten firings rather than one outcome's rarity.
  • Uses 1 - p as the measure, which caps information at one.
  • Assumes surprisal must always be a whole number of bits.
open as a page

Why is a single outcome's information content defined as -log2 p rather than as 1/p or 1 - p?

level: middleimportance: must knowfreq 48%

basics

~20 s

Because information from independent observations should add, while their probabilities multiply. Only a logarithm turns a product into a sum, and it also scores certainty as zero. Measures like 1/p multiply instead of adding, and 1 - p saturates.

open as a page

In surprisal terms, one signal fires in 999 minutes out of 1000 and another in one minute out of 1024 - which should page a human, and why?

level: seniorimportance: should knowfreq 38%

basics

~20 s

The rare signal, at about 10 bits per firing, carries real information; the near-certain one carries about 0.0014 bits and confirms what you already knew. For the near-certain signal the informative event is its absence, worth about 9.97 bits.

open as a page

What changes about a measured information content when it is expressed in nats instead of bits?

level: middleimportance: nice to knowfreq 22%

basics

~20 s

Only the unit changes. The base of the logarithm sets the unit - base 2 gives bits, the natural logarithm gives nats - and the two differ by the constant factor ln 2, so every ordering, ratio and comparison is unchanged.

open as a page