skip to content

In the geometric distribution, why is the mean number of trials until the first success 1/p?

level: middleimportance: should knowfreq 44%

answer

  1. one over the per-trial success chance
  2. one trial spent, then restart
  3. E equals 1 plus (1-p)E
  4. mean sits well above the median

basics

~20 s

If independent trials each succeed with probability p, the number of trials up to and including the first success is geometric with mean 1/p. With a 2 percent click rate you expect 50 impressions per first click.

solid answer

~40 s

Counting trials up to and including the first success, the geometric mass function is `P(X = k) = (1-p)^(k-1) · p`. The mean follows from a one-line restart argument: you always spend one trial; with probability p you are done, and with probability (1-p) you are back where you started, so `E = 1 + (1-p)E`, giving `E = 1/p`. The variance is `(1-p)/p²`. Watch the convention — some texts count *failures before* the first success, which shifts the mean to (1-p)/p; say which you are using. With p = 0.02 the mean wait is 50 impressions, but the distribution is strongly right-skewed: the median is 35 and about 36 percent of the time you exceed 50. The negative binomial extends this to the r-th success, with mean r/p.

go deeper

for a junior

Recall that with a constant per-trial success chance p, the expected number of trials until the first success is 1/p, and be able to apply it to a rate such as 2 percent giving 50 trials.

for a middle

Derive it with the restart argument E = 1 + (1-p)E, state the variance (1-p)/p squared, and name which of the two counting conventions you are using before quoting a mean.

for a senior

Show you use the whole distribution, not the mean: quote a median and a tail probability, and explain why a heavily skewed wait makes mean-based planning misleading in funnel or campaign work.

for a principal

Own the modelling assumption itself. Decide whether a single constant p is defensible across a heterogeneous audience or a decaying response, and articulate what pooling does to the tail of the observed waits.

## Setup Run independent trials that each succeed with the same probability p. The **geometric** distribution describes how long you wait for the first success. There are two conventions, and confusing them is the single most common error: - **Trials convention** (used here): X = number of trials up to *and including* the first success. Support {1, 2, 3, ...}, `P(X = k) = (1-p)^(k-1) p`, mean `1/p`. - **Failures convention**: Y = number of failures *before* the first success. Support {0, 1, 2, ...}, `P(Y = k) = (1-p)^k p`, mean `(1-p)/p`. The two differ by exactly 1, since X = Y + 1. Both have variance `(1-p)/p²`, because shifting a variable by a constant leaves its spread unchanged. State your convention before you state your answer. ## Why the mean is 1/p — the restart argument Let E be the expected number of trials until the first success. Run one trial. That trial always costs you 1. With probability p it succeeds and you stop; with probability (1-p) it fails and — crucially, because the trials are independent and identically distributed — you face exactly the same problem you started with. So `E = 1 + (1-p) · E`. Solving: `E - (1-p)E = 1`, so `pE = 1` and `E = 1/p`. This argument is worth memorising because it generalises to many waiting-time problems and needs no sum manipulation. ## Why it is 1/p — the direct sum The brute-force route: `E[X] = Σ_{k≥1} k (1-p)^(k-1) p`. Writing q = 1-p, this is `p · Σ_{k≥1} k q^(k-1) = p · 1/(1-q)² = p/p² = 1/p`, using the standard derivative-of-a-geometric-series identity. ## Variance and skew `Var(X) = (1-p)/p²`, so the standard deviation is `sqrt(1-p)/p`, which for small p is very close to the mean 1/p itself. That near-equality is a warning: the geometric is heavily right-skewed, with a long tail of unlucky waits. Take a click-through probability p = 0.02: - Mean: 1/0.02 = **50** impressions. - Standard deviation: sqrt(0.98)/0.02 ≈ **49.5**. - Median: the smallest k with P(X ≤ k) ≥ 0.5, which works out to **35**. - P(X > 50) = 0.98^50 ≈ **0.364**. So the "expected" wait of 50 is exceeded more than a third of the time, and the typical wait is nearer 35. Treating the mean as a typical outcome, or as a planning guarantee, is a real mistake in campaign or funnel work: quoting "about 50 impressions per click" hides that a quarter of the time you will need more than 68 impressions. (P(X > 68) = 0.98^68 ≈ 0.25.) ## Memorylessness The geometric is the **only** discrete distribution that is memoryless: `P(X > m + n | X > m) = P(X > n)`. After 30 straight failures, the expected number of *further* trials is still 1/p. Nothing is "due". This falls straight out of independence — the coin does not know its history — but it contradicts strong intuition, so interviewers like to test it. ## Extending to the r-th success The **negative binomial** counts trials needed to reach the r-th success. Under the trials convention its mass function is `P(X = k) = C(k-1, r-1) p^r (1-p)^(k-r)` for k ≥ r, with - mean `r/p`, - variance `r(1-p)/p²`. These follow immediately from the fact that the total wait is a sum of r independent geometric waits: means add, and variances add because the segments are independent. At r = 1 it reduces to the geometric. So with p = 0.02, waiting for 3 clicks takes 150 impressions on average. ## The assumptions All of this needs independent trials with a **constant** p. If the click probability decays as a user sees the same ad repeatedly, or if impressions are served to a heterogeneous audience with different propensities, the geometric will misdescribe the tail — typically the pooled wait is more variable than the model says, because the low-p users dominate the long waits. ## Common errors Quoting the mean as p rather than 1/p; using the two conventions interchangeably within one answer; assuming the median equals the mean when the distribution is strongly skewed; and expecting a run of failures to make success more likely.

  • What distribution and mean describe the number of trials needed for the r-th success?
    The negative binomial, with mean r/p and variance r(1-p)/p² under the trials convention. It is a sum of r independent geometric waits, so the means and variances simply add. At r = 1 it collapses back to the geometric. With p = 0.02, three successes take 150 trials on average.
  • With a 2 percent success rate, is 50 trials a typical wait?
    No. The mean is 50 but the distribution is strongly right-skewed: the median is 35, and P(more than 50 trials) = 0.98^50 ≈ 0.36. Treating the mean as a typical or worst-case wait understates how often long runs happen, which matters for planning.
  • Why is the geometric distribution called memoryless?
    Because P(X > m + n | X > m) = P(X > n): after m failures, the remaining wait has exactly the same distribution as a fresh start, and its expectation is still 1/p. It is the only discrete distribution with this property, and it follows directly from the trials being independent.

Every trial is an identical lottery ticket. Since the tickets have no memory of the losing ones before them, the average wait is simply one divided by the chance a single ticket wins.

saying these in an interview costs you the question

  • Says the mean number of trials is p
  • Switches conventions without stating which is used
  • Assumes the median matches the mean of 1/p
  • Thinks a long failure run makes success more likely
  • Applies the geometric when p drifts between trials

context