skip to content

In the two-children problem, why does 'at least one is a girl' give 1/3, not 1/2?

level: middleimportance: should knowfreq 34%

answer

  1. write down all four ordered outcomes
  2. which outcomes does the statement rule out
  3. only BB is eliminated in one version
  4. there is no well-defined other child
  5. how you learned it changes the answer

basics

~20 s

Because 'at least one is a girl' leaves three equally likely families, girl-boy, boy-girl and girl-girl, of which one has two girls. Naming a specific child, such as the elder, leaves two cases and gives 1/2.

solid answer

~40 s

Order the children as (elder, younger) and assume each is independently a girl or boy with probability 1/2. The four equally likely outcomes are `GG`, `GB`, `BG`, `BB`. The event 'at least one is a girl' rules out only `BB`, leaving `{GG, GB, BG}`, so `P(both girls | at least one girl) = (1/4) / (3/4) = 1/3`. The event 'the elder is a girl' leaves `{GG, GB}`, so that probability is `(1/4) / (1/2) = 1/2`. The difference is purely the size of the conditioning set: the first phrase eliminates one of four cases, the second eliminates two. The children's sexes stay independent in both computations. Note that how you learned the fact matters too: if you met one child at random and she was a girl, the answer is 1/2.

go deeper

for a junior

Be able to list the four ordered outcomes GG, GB, BG, BB and apply P(A given B) = P(A and B) / P(B) to get 1/3. Enumeration first, arithmetic second.

for a middle

Explain why the mixed-sex cases count twice and why there is no well-defined other child, then contrast the result with conditioning on the elder child being a girl.

for a senior

Raise the protocol question unprompted: the answer depends on whether a property of the pair was reported or a particular child was observed, and say which conditioning event each produces.

for a principal

Be ready to generalise the lesson — that a probability claim is meaningless without the data-generating process behind the observation — to how your organisation words and audits evidence in analyses.

## Setting up the sample space Assume a family has exactly two children, that each child is a girl or a boy with probability 1/2, and that the two children's sexes are independent. Writing outcomes as ordered pairs (elder, younger), the sample space is ``` GG, GB, BG, BB each with probability 1/4 ``` Independence is what makes those four probabilities equal: `P(GB) = P(G)P(B) = 1/2 * 1/2 = 1/4`, and likewise for the rest. Everything that follows is the definition of conditional probability applied to this space. ## The two conditionings **'At least one is a girl'.** This event is `{GG, GB, BG}` — it excludes only `BB` — and has probability 3/4. 'Both are girls' is `{GG}` with probability 1/4, and it sits inside the conditioning event. So ``` P(GG | at least one girl) = P(GG) / P(at least one girl) = (1/4) / (3/4) = 1/3 ``` **'The elder is a girl'.** This event is `{GG, GB}` with probability 1/2, and again `{GG}` sits inside it: ``` P(GG | elder is a girl) = (1/4) / (1/2) = 1/2 ``` Same family, same underlying independence, different answers — because the two statements carve out conditioning sets of different sizes. The first shrinks the space from four outcomes to three, the second from four to two, and the target outcome `GG` is one of them either way. ## Why the intuition misfires The common wrong answer to the first version is 1/2, reasoning that 'the other child is equally likely a girl or a boy, so the answer is 1/2'. The flaw is the phrase *the other child*. 'At least one is a girl' does not identify a specific child, so there is no well-defined 'other' one. In the mixed families `GB` and `BG` the girl is a different child, and both count. The 1/2 answer silently merges `GB` and `BG` into a single outcome 'one of each', which is not equally likely with `GG` — mixed-sex families are twice as likely as two-girl families, exactly because there are two orderings. That is the whole lesson: **conditioning on a property of the pair is not the same as conditioning on a property of an identified member.** The moment you name a member — the elder, the one who answered the door, the one born on a specified day — you split the mixed cases and the arithmetic changes. ## The protocol caveat, which is the real depth of the problem The 1/3 answer is correct for one precise information-generating process: someone examined the whole family and truthfully reported whether the statement 'at least one child is a girl' is true. That is not the only way you could come to know a girl exists. Suppose instead you meet one of the two children uniformly at random and observe that she is a girl. Now condition on that: the other child is independently a girl or a boy, so the probability both are girls is 1/2. Nothing about the family changed; the *sampling protocol* that produced your information changed, and with it the conditioning event. This is why careful statements of the puzzle specify how the information arrived, and why a strong answer volunteers the distinction rather than waiting to be caught by it. An interviewer asking this question is usually probing exactly whether you notice that 'I know a girl exists' is ambiguous between 'the pair contains a girl' and 'this particular child is a girl'. ## What the problem is really testing Three things, in order of how often candidates trip on them: 1. **Do you write the sample space down?** Four ordered outcomes, equally likely by independence. Nearly every error is an error of enumeration, not of arithmetic. 2. **Do you apply the definition rather than intuition?** `P(A|B) = P(A and B) / P(B)` mechanically gives both answers in seconds. 3. **Do you notice that independence is untouched?** The two children's sexes remain independent throughout. Conditioning on a joint property of the pair changes what *you* know without making the underlying events dependent — the probabilities you compute are about your information state, not about a causal link between siblings. ## Assumptions to state out loud The idealised model — exactly two children, each independently a girl with probability 1/2 — is what produces the clean 1/3. Real birth ratios are not exactly even and sexes within a family are not perfectly independent, so the puzzle is a statement about the model, not a demographic claim. Naming the assumptions before computing is the mark of someone who has thought about the problem rather than recited the answer.

  • How does the answer change if you are told the elder child is a girl?
    It becomes 1/2. Naming a specific child cuts the sample space to `{GG, GB}` in (elder, younger) order, so `P(GG) / P(elder is a girl) = (1/4)/(1/2) = 1/2`. Identifying a member splits the two mixed-sex orderings instead of counting both, which is exactly what moves the answer from 1/3.
  • Why does meeting one child at random and seeing a girl give a different answer?
    Because the conditioning event is different. Random-child sampling identifies a particular child as a girl, and the other child is independently a girl with probability 1/2. The reported-property version only tells you the pair is not two boys, which leaves three equally likely families rather than two.
  • Does the 1/3 answer mean the two children's sexes are dependent?
    No. Independence of the two children is what makes the four ordered outcomes equally likely in the first place. Conditioning on a joint property of the pair changes your information state without creating any link between the siblings; the numbers describe what you know, not a causal relationship.
  • What modelling assumptions produce the clean 1/3?
    Exactly two children, each independently a girl with probability 1/2, and information consisting of precisely the statement that the family is not two boys. Real birth ratios are not exactly even and within-family sexes are not perfectly independent, so the result is a claim about the idealised model.

saying these in an interview costs you the question

  • Answers 1/2 by reasoning about the other child
  • Merges the two mixed-sex orderings into one outcome
  • Claims the two children's sexes must be dependent
  • Gives the same answer regardless of how the fact was learned
  • Skips enumerating the sample space before conditioning

context