skip to content

How do you derive the maximum likelihood estimate of a coin's heads probability from 7 heads in 10 flips?

level: middleimportance: must knowfreq 68%

answer

  1. turn the product into a sum
  2. differentiate the log-likelihood, set to zero
  3. 7 over p minus 3 over one minus p
  4. second derivative is negative throughout
  5. answer is the sample proportion

basics

~10 s

Write the likelihood p^7 * (1-p)^3, take logs to get 7log(p) + 3log(1-p), differentiate and set the result to zero. That gives 7/p = 3/(1-p), so p-hat = 0.7, the sample proportion.

solid answer

~40 s

Each flip is Bernoulli with heads probability `p`, so with 7 heads and 3 tails the likelihood is `L(p) = p^7 * (1-p)^3` up to a constant that does not involve `p`. Taking logs turns the product into a sum: `l(p) = 7*log(p) + 3*log(1-p)`. Differentiating gives `l'(p) = 7/p - 3/(1-p)`, and setting that to zero yields `7*(1-p) = 3*p`, so `p-hat = 7/10 = 0.7`. In general with `x` successes in `n` trials the same algebra gives `p-hat = x/n`, the sample proportion. I would confirm it is a maximum rather than a minimum by noting the second derivative `-7/p^2 - 3/(1-p)^2` is negative everywhere in `(0,1)`, so the log-likelihood is strictly concave and the stationary point is the unique maximum.

go deeper

for a junior

Know the punchline and be able to state it: the maximum likelihood estimate of a coin's heads probability is the observed proportion of heads, 7 out of 10 here. Recognise the words likelihood and log-likelihood when they appear.

for a middle

Be able to run all four steps at the whiteboard without hesitation, including the log step and the algebra from 7/p = 3/(1-p) to 0.7. Expect a second model, usually exponential or Normal, immediately after.

for a senior

Show the habits that keep the derivation honest in real work: verifying concavity, checking the parameter-space boundary, and noticing when a fitted probability of exactly zero or one is a data artefact rather than a finding.

for a principal

Frame when a hand derivation is worth the effort at all versus numerical optimisation, and what a team should standardise on when closed forms stop existing -- reproducibility of the fitting procedure matters more than elegance at that point.

## Step 1: write down the likelihood Model each flip as an independent Bernoulli trial with unknown heads probability `p`. A single observation contributes `p` if it is a head and `1-p` if it is a tail, which can be written compactly as `p^x * (1-p)^(1-x)` for `x` in {0,1}. Independence multiplies the contributions, so for 7 heads and 3 tails: `L(p) = p^7 * (1-p)^3` If you frame the data as a binomial count rather than an ordered sequence, you pick up a factor `C(10,7) = 120`. It contains no `p`, so it rescales the curve without moving its peak and can be dropped. ## Step 2: take logs Products are awkward to differentiate and underflow numerically once `n` is large. The logarithm is strictly increasing, so it preserves the location of the maximum while turning the product into a sum: `l(p) = log L(p) = 7*log(p) + 3*log(1-p)` ## Step 3: differentiate and solve `l'(p) = 7/p - 3/(1-p)` Setting this to zero -- the score equation -- gives `7/p = 3/(1-p)`, hence `7*(1-p) = 3*p`, hence `7 = 10*p`, hence: `p-hat = 0.7` ## Step 4: check it is a maximum A stationary point is not automatically a maximum, and a good answer says so. Differentiating again: `l''(p) = -7/p^2 - 3/(1-p)^2` Both terms are negative for every `p` strictly between 0 and 1, so the log-likelihood is strictly concave on the open interval, and the single stationary point is its unique global maximum. You can also see it directly: `l(p)` goes to minus infinity as `p` approaches 0 or 1, so the maximum must be interior. ## The general result Repeat with `x` successes in `n` trials: `l(p) = x*log(p) + (n-x)*log(1-p)`, `l'(p) = x/p - (n-x)/(1-p) = 0`, giving `p-hat = x/n` The maximum likelihood estimate of a Bernoulli probability is just the sample proportion. That is worth stating out loud, because it shows the machinery reproduces the obvious answer in the simplest case -- which is exactly why interviewers use it as the warm-up before a harder model. ## The same recipe on a continuous model The steps do not change when the data are continuous. For `n` waiting times `x1, ..., xn` modelled as exponential with rate `lambda`, the density is `lambda * exp(-lambda*x)`, so `l(lambda) = n*log(lambda) - lambda * sum(xi)` `l'(lambda) = n/lambda - sum(xi) = 0` `lambda-hat = n / sum(xi) = 1 / xbar` The estimated rate is the reciprocal of the average waiting time -- if calls arrive on average every 4 minutes, the fitted rate is 0.25 calls per minute. Note the shape of the answer: the likelihood equation collapses to a statement about a simple summary of the data (here the sum, and hence the mean), which is typical of these tidy exponential-family models. ## Boundary cases The derivative-equals-zero step assumes the maximum sits strictly inside the parameter range. Observe 0 heads in 10 flips and `l(p) = 10*log(1-p)`, which is strictly decreasing on `(0,1)` -- there is no interior stationary point at all. The likelihood is maximised at the edge of the range, `p-hat = 0`. That is a legitimate maximiser of the likelihood and also an uncomfortable estimate, since it asserts heads are impossible on the strength of ten flips. The lesson to state in an interview is procedural: always check the endpoints before declaring that a solved score equation is the answer. ## What interviewers are grading They want four things, in order: the likelihood written down correctly from the model, the move to logs justified rather than assumed, clean algebra to `x/n`, and some acknowledgement that a stationary point needs a second-order or boundary argument before you call it the maximum. Candidates who jump straight to "the answer is 0.7 because that is the observed proportion" have given the right number without demonstrating the method, and the follow-up will simply be a model where intuition does not supply the answer.

  • Why maximise the log-likelihood rather than the likelihood itself?
    Because the logarithm is strictly increasing, both are maximised at the same parameter value, so nothing is lost. What is gained is practical: a product of `n` terms becomes a sum, which differentiates term by term and does not underflow to zero in floating point once `n` is large. Constants free of the parameter also become additive and drop out on differentiation.
  • Apply the same steps to n exponential waiting times with rate lambda. What is the MLE?
    The log-likelihood is `n*log(lambda) - lambda*sum(xi)`. Differentiating gives `n/lambda - sum(xi)`, and setting it to zero gives `lambda-hat = n/sum(xi) = 1/xbar`. The estimated rate is the reciprocal of the mean waiting time, so an average gap of 4 minutes gives a fitted rate of 0.25 per minute.
  • What happens to this derivation if you observe 0 heads in 10 flips?
    The log-likelihood becomes `10*log(1-p)`, which decreases strictly on the open interval, so the score equation has no interior solution. The likelihood is maximised at the boundary, `p-hat = 0`. It is a genuine maximiser but a poor estimate, which is why you check endpoints instead of assuming the derivative always locates the answer.
  • How would you convince an interviewer the stationary point is a maximum and not a minimum?
    Compute the second derivative: `-7/p^2 - 3/(1-p)^2`, negative for every `p` in `(0,1)`, so the log-likelihood is strictly concave and the stationary point is the unique global maximum. A quicker argument is that the log-likelihood tends to minus infinity at both ends of the range, so the maximum must be interior.

saying these in an interview costs you the question

  • Differentiates the product form instead of taking logs first
  • Reports a stationary point without checking it is a maximum
  • Gives 7/3, the heads-to-tails ratio, as the estimate
  • Assumes the maximum is always interior to the parameter range
  • Cannot generalise the 0.7 answer to x over n

context