skip to content

Maximum Likelihood Estimation

Writing the likelihood of observed data, taking logs, and solving for the parameter that maximises it, for Bernoulli, Normal and exponential samples. Interviewers want a derivation, not a formula.

on this pageshow

questions

5

How does a likelihood differ from a probability in maximum likelihood estimation?

level: juniorimportance: must knowfreq 80%

answer

  1. same formula, different variable held fixed
  2. data fixed, parameter free to vary
  3. no normalisation over the parameter
  4. not the probability that p is 0.7

basics

~20 s

A probability treats the parameter as fixed and asks how likely the data are; a likelihood fixes the observed data and reads the same formula as a function of the parameter. Likelihood values do not sum to one over parameters.

solid answer

~40 s

They are the same formula read in opposite directions. For a coin flipped 10 times, a probability question fixes the heads rate `p` and asks how likely 7 heads are. A likelihood fixes the observed 7 heads out of 10 and treats `p` as the free variable: `L(p) = C(10,7) * p^7 * (1-p)^3`. So `L(0.7)` is about 0.27, and that number is the chance of seeing exactly 7 heads *if* `p` were 0.7 -- it is not the probability that `p` equals 0.7. Maximum likelihood estimation simply picks the `p` that makes `L(p)` largest. Because `L` is not a distribution over `p`, it does not integrate to one across `p`, and its absolute height carries no meaning; only comparisons between parameter values do.

go deeper

for a junior

Be ready to state the swap in one line: probability fixes the parameter and varies the data, likelihood fixes the data and varies the parameter. Then say what maximum likelihood does with it -- pick the parameter making the observed data most probable.

for a middle

Explain the mechanics: independent observations multiply, so the likelihood is a product, and you work with its logarithm to turn that into a sum. Show that constants free of the parameter can be dropped without moving the maximum.

for a senior

Demonstrate that you police the conditioning in real write-ups. Interviewers listen for whether you can catch a colleague's claim that a likelihood value is the probability a parameter is correct, and whether you report likelihood ratios rather than raw heights.

for a principal

Own the communication risk. Decide how model-fit evidence is phrased to non-statistical stakeholders, since likelihood ratios read as probabilities to most audiences, and set a house convention for reporting comparative evidence rather than bare likelihood numbers.

## The same formula, two different free variables Start from one object: a model that assigns a number to data given a parameter. Write it `f(data | theta)`. Nothing about that expression tells you which symbol you are allowed to vary. That choice is what separates a probability from a likelihood. - **Probability**: hold `theta` fixed at some assumed value, let `data` vary. The result is a genuine distribution over outcomes. It sums (discrete case) or integrates (continuous case) to one across all possible datasets. - **Likelihood**: hold `data` fixed at what you actually observed, let `theta` vary. The result is a function of the parameter, written `L(theta) = f(observed data | theta)`. It is *not* a distribution over `theta`. ## Worked case: 7 heads in 10 flips Suppose a coin is flipped 10 times and lands heads 7 times. The model is binomial with unknown heads probability `p`: `L(p) = C(10,7) * p^7 * (1-p)^3` Evaluate it at a few parameter values: `L(0.5) = 120 * 0.5^10 ~ 0.117`, `L(0.7) = 120 * 0.7^7 * 0.3^3 ~ 0.267`, `L(0.9) = 120 * 0.9^7 * 0.1^3 ~ 0.057`. Reading these correctly: 0.267 is the probability of getting exactly 7 heads in 10 flips **in a world where `p` really is 0.7**. It is not "the probability that `p` is 0.7", and it is not a statement about the coin's bias at all until you compare it with other values. The comparison is the whole point. `L(0.7)` beats `L(0.5)` by a factor of roughly 2.3, so the data are about 2.3 times more probable under `p = 0.7` than under a fair coin. Maximum likelihood estimation formalises exactly this: pick the parameter value under which the observed data would have been most probable. Here that is `p-hat = 0.7`. ## Why the likelihood is not a probability distribution over the parameter Two concrete consequences follow, and interviewers probe both. First, **the likelihood need not integrate to one over the parameter**. For this example, integrating `L(p)` over `p` from 0 to 1 gives exactly `1/11`, not 1. Nothing normalises it, because nothing in the model claims `p` is random. Second, **the absolute value of a likelihood is meaningless on its own**. In the continuous case the model returns a *density*, not a probability, so `L(theta)` can exceed 1 -- a Normal density with a very small standard deviation is tall and narrow, and its value at the peak is large. That is not a broken probability; it is a density evaluated at a point. Only ratios and the location of the maximum are interpretable. This is also why any multiplicative constant that does not involve the parameter can be dropped. The binomial coefficient `C(10,7) = 120` counts the orderings of heads and tails; it does not contain `p`, so `p^7 * (1-p)^3` has its maximum at exactly the same place. Most derivations drop such constants immediately. ## Independence and the product form With independent observations `x1, ..., xn` the joint model factorises, so the likelihood is a product of per-observation terms: `L(theta) = f(x1 | theta) * f(x2 | theta) * ... * f(xn | theta)` Products of many small numbers underflow quickly and are awkward to differentiate, so in practice the **log-likelihood** is used: `l(theta) = sum of log f(xi | theta)`. Because the logarithm is strictly increasing, it has its maximum at exactly the same parameter value, so nothing is lost. ## Saying it cleanly in an interview A crisp phrasing: "Probability runs forward from a parameter to data; likelihood runs backward from data to a parameter, using the same formula. The likelihood is a function of the parameter given fixed data, and it is not a probability distribution over that parameter -- it does not normalise, and its height only matters relative to other parameter values." The classic trap is to answer "`L(0.7) = 0.267` means there is a 27% chance the coin's bias is 0.7". That statement flips the conditioning, and interviewers hear it as a sign that the candidate has not internalised what is being held fixed.

  • Can a likelihood value be greater than one?
    Yes, whenever the model is continuous. There the likelihood is a probability density evaluated at the observed points, and densities are not capped at one -- a Normal density with a small standard deviation is tall and narrow, so its peak exceeds one. Only in discrete models is each factor a genuine probability bounded by one. Either way, only relative values and the location of the maximum matter.
  • Why can you drop the binomial coefficient when maximising a binomial likelihood?
    Because it is a multiplicative constant that does not contain the parameter. `C(10,7) = 120` counts orderings of heads and tails, so scaling `p^7 * (1-p)^3` by it stretches the curve vertically without moving its peak. The maximiser is unchanged, and in log form it becomes an additive constant that vanishes on differentiation.
  • Two parameter values give likelihoods of 0.267 and 0.117. What does that comparison license you to say?
    Only that the observed data are about 2.3 times more probable under the first value than under the second. It is evidence favouring the first value on this data, not a probability that the first value is correct, and not a decision on its own -- the ratio has no fixed threshold attached to it.

One recipe read two ways: forward it tells you what a given oven temperature will produce; backward it tells you which oven temperature best explains the cake you actually pulled out.

saying these in an interview costs you the question

  • Says L(0.7) is the probability that the parameter equals 0.7
  • Claims the likelihood must integrate to one over the parameter
  • Treats a likelihood above one as proof of an error
  • Reads a single likelihood value as meaningful without comparison
  • Uses the words probability and likelihood interchangeably throughout

context

open as a page

How do you derive the maximum likelihood estimate of a coin's heads probability from 7 heads in 10 flips?

level: middleimportance: must knowfreq 68%

basics

~10 s

Write the likelihood p^7 * (1-p)^3, take logs to get 7log(p) + 3log(1-p), differentiate and set the result to zero. That gives 7/p = 3/(1-p), so p-hat = 0.7, the sample proportion.

open as a page

What are the maximum likelihood estimates of the mean and variance of a Normal sample?

level: middleimportance: should knowfreq 52%

basics

~20 s

The mean estimate is the sample average xbar. The variance estimate is the average squared deviation from xbar, that is sum of (xi - xbar)^2 divided by n. Maximising the likelihood returns the divisor n, not n-1.

open as a page

Why is minimising squared error the same as maximum likelihood under Gaussian noise?

level: seniorimportance: should knowfreq 57%

basics

~20 s

Assume the errors are independent Gaussian with constant variance. The log-likelihood then equals a constant minus the sum of squared residuals divided by twice the variance, so maximising it and minimising squared error give the same fitted parameters.

open as a page

If the MLE of a coin's heads probability is 0.7, what is the MLE of the odds p/(1-p)?

level: middleimportance: nice to knowfreq 32%

basics

~10 s

It is 0.7/0.3, about 2.33. The invariance property of maximum likelihood says the estimate of any function of a parameter is that function applied to the parameter's estimate, so no new maximisation is needed.

open as a page