Which Kolmogorov axioms must a probability assignment satisfy to be valid?
answer
- three rules, nothing more
- never negative, whole space is one
- disjoint events add up
- the cap at 1 is derived
- sum-to-one only for disjoint and exhaustive
basics
~20 sThree: every event gets a probability of at least zero; the whole sample space gets probability 1; and the probabilities of disjoint events add. Everything else, including the complement rule and the cap at 1, is derived from these.
solid answer
~40 sKolmogorov's axioms are non-negativity, `P(A) ≥ 0` for every event; normalisation, `P(S) = 1` for the sample space `S`; and countable additivity, `P(A1 ∪ A2 ∪ ...) = P(A1) + P(A2) + ...` for pairwise disjoint events. Everything else follows: `P(empty) = 0`, the complement rule `P(A^c) = 1 - P(A)`, monotonicity, the cap `P(A) ≤ 1`, and the addition rule. So if someone proposes three mutually exclusive and exhaustive churn reasons with probabilities 0.5, 0.4 and 0.3, reject it: additivity makes those a total of 1.2 for the whole sample space, contradicting normalisation. A negative value is rejected on sight. The axioms only enforce internal consistency; they never tell you which numbers are the right ones.
go deeper
Be able to state the three axioms in your own words and spot an obviously illegal assignment, such as a negative value or exhaustive categories that total more than one.
Derive the complement rule and the upper bound of 1 from the axioms rather than listing them alongside. Explain that additivity applies only to disjoint events, which is the hinge of most trick versions.
Show judgment on real model output: distinguish a genuine axiom violation from overlapping labels that legitimately exceed a total of one, and separate internal consistency from whether the numbers are well calibrated.
Own where probability numbers come from in the organisation: symmetry, observed frequency or expert judgment, and make sure validity checks run automatically on model output rather than being spotted in a review.
## The three axioms A probability model is a sample space `S` together with a function `P` assigning a number to each event. Kolmogorov's axioms are the minimal conditions that make `P` a probability: 1. **Non-negativity** — `P(A) ≥ 0` for every event `A`. 2. **Normalisation** — `P(S) = 1`, where `S` is the entire sample space. 3. **Countable additivity** — if `A1, A2, A3, ...` are pairwise disjoint (no two share an outcome), then `P(A1 ∪ A2 ∪ ...) = P(A1) + P(A2) + ...` That is the whole foundation. Notice what is *not* an axiom: there is no axiom saying probabilities are at most 1, no axiom about complements, and nothing at all about independence or conditional probability. Those are either consequences or separate definitions. ## What follows from them - **The empty event has probability 0.** The empty set is disjoint from itself, and additivity forces `P(empty) = 0`. - **Complement rule**: `A` and `A^c` are disjoint and their union is `S`, so `P(A) + P(A^c) = 1`, giving `P(A^c) = 1 - P(A)`. - **Upper bound**: since `P(A^c) ≥ 0`, the complement rule gives `P(A) ≤ 1`. The familiar "probabilities live between 0 and 1" is half axiom and half theorem. - **Monotonicity**: if every outcome of `A` is in `B`, write `B` as the disjoint union of `A` and `B` minus `A`; both pieces are non-negative, so `P(A) ≤ P(B)`. - **Addition rule**: `P(A ∪ B) = P(A) + P(B) - P(A ∩ B)`, derived by cutting the union into disjoint pieces. ## Checking a proposed assignment The practical version of this question is: someone hands you numbers, and you say whether they can be probabilities. **A negative value.** Suppose a model outputs `-0.05` for one category. That is not a rounding curiosity; it violates non-negativity outright, and the model is not a probability model. This happens in practice when numbers come from a fitting procedure that was never constrained to the simplex. **Values summing above 1.** Suppose three churn reasons are described as mutually exclusive and exhaustive, with probabilities 0.5, 0.4 and 0.3. Because they are disjoint, additivity says their union has probability `0.5 + 0.4 + 0.3 = 1.2`. Because they are exhaustive, that union is `S`, so normalisation demands 1. The assignment is contradictory and must be rejected. **Values summing below 1.** If those same three exhaustive reasons totalled 0.85, the model implicitly leaves 0.15 of the probability unassigned, which means the categories are not really exhaustive: some outcome is missing, often an "other" or "unknown" case. **Values summing above 1 that are perfectly fine.** This is the nuance that separates a rehearsed answer from an understood one. If three *overlapping* labels have probabilities 0.6, 0.5 and 0.3, the total 1.4 breaks nothing. A user can carry several labels at once, so additivity does not apply and there is no contradiction. The sum-to-one check is only meaningful for a set of events that are pairwise disjoint and cover the sample space. Before declaring numbers invalid, verify that claim about the events. ## What the axioms deliberately do not do The axioms are a consistency contract, not a source of numbers. They cannot tell you whether a coin's probability of heads is 0.5 or 0.6; both assignments satisfy all three axioms. Where the numbers come from is a separate question, answered by symmetry arguments, by observed long-run frequencies, or by a considered subjective judgment. A candidate who says "the axioms tell you the probability" has confused the rules of arithmetic with the measurement. The axioms also say nothing about independence. `P(A ∩ B) = P(A) P(B)` is the *definition* of independence, an extra property a particular model may or may not have, and it is a common error to recite it as a fourth axiom. ## Interview framing Expect this to arrive disguised: "here are the outputs of a scoring system, is anything wrong with them?" The strong answer names the axiom that is violated, states whether the events involved are actually disjoint and exhaustive, and distinguishes a genuine contradiction from a legitimate overlapping-tag total. If nothing is violated, say the assignment is internally consistent and then ask the separate question of whether it is *well calibrated*, which the axioms cannot judge. ## Traps to avoid - Listing "probabilities are between 0 and 1" as an axiom. The upper bound is derived. - Adding independence as a fourth axiom. - Declaring any set of numbers that sums above 1 invalid without checking that the events are disjoint. - Treating a sum below 1 over supposedly exhaustive categories as a rounding issue rather than a missing outcome.
- Is "probabilities are at most 1" one of the axioms?No, it is a consequence. Non-negativity gives `P(A^c) ≥ 0`, and the complement rule `P(A) + P(A^c) = 1` then forces `P(A) ≤ 1`. Only non-negativity, normalisation and countable additivity are assumed; the upper bound, `P(empty) = 0` and monotonicity are all theorems derived from those three.
- Three user tags have probabilities 0.6, 0.5 and 0.3, totalling 1.4. Is the model broken?Not necessarily. Additivity applies only to disjoint events, so a total above 1 is a contradiction only if the tags are mutually exclusive. If a user can carry several tags at once, 1.4 is perfectly legal and simply reflects overlap. Establish whether the events partition the sample space before judging the arithmetic.
- Do the axioms tell you what the probability of an event actually is?No. They constrain assignments to be mutually consistent but never select one. Assigning 0.5 or 0.6 to heads both satisfy all three axioms. The numbers come from elsewhere: physical symmetry, observed long-run frequencies, or an explicit subjective judgment. Consistency and correctness are separate questions, and only the first is settled by the axioms.
saying these in an interview costs you the question
- Reciting independence as a fourth axiom
- Calling the upper bound of 1 an axiom rather than a derived result
- Rejecting any probabilities that sum above 1 without checking disjointness
- Claiming the axioms determine the actual probability values
- Ignoring that additivity requires the events to be pairwise disjoint