If the MLE of a coin's heads probability is 0.7, what is the MLE of the odds p/(1-p)?
answer
- no re-derivation is needed
- transform the estimate, not the data
- relabelling the axis cannot move the peak
- g of theta-hat
- 0.7 divided by 0.3
basics
~10 sIt is 0.7/0.3, about 2.33. The invariance property of maximum likelihood says the estimate of any function of a parameter is that function applied to the parameter's estimate, so no new maximisation is needed.
solid answer
~40 sApply the function to the estimate: `0.7 / (1 - 0.7) = 7/3 ~ 2.33`. This is the invariance (or equivariance) property of maximum likelihood: if `theta-hat` maximises the likelihood for `theta`, then `g(theta-hat)` is the maximum likelihood estimate of `g(theta)`. For a one-to-one `g` the reason is immediate -- reparameterising relabels the horizontal axis of the likelihood curve without changing which point is highest, so the peak moves to `g(theta-hat)` and nowhere else. The same rule gives the MLE of `sigma` as the square root of the MLE of `sigma^2`, and the MLE of a Normal mean's square as `xbar^2`. It is a genuinely useful shortcut, but it is a statement about maximisers only: other properties of an estimate do not survive a nonlinear transformation just because the estimate itself does.
go deeper
Know the mechanical rule and be able to use it: to estimate a function of a parameter, apply that function to the parameter's estimate. Odds from a fitted probability of 0.7 are 0.7 divided by 0.3.
Explain why it works -- reparameterising relabels the parameter axis, so the location of the likelihood's peak carries over unchanged, and no Jacobian appears because the likelihood is not a density in the parameter.
Show where the shortcut stops. Interviewers want to hear that invariance moves the maximiser only, and that quantities defined by averaging do not pass through a nonlinear transformation the same way.
Own the parameterisation decision itself: choose the scale a team fits and reports on, knowing derived quantities come free, and set expectations about which downstream summaries can be transformed and which must be recomputed.
## The property Maximum likelihood estimation has an unusual and very convenient property: it commutes with reparameterisation. Formally, if `theta-hat` is the maximum likelihood estimate of `theta`, then for any function `g` the maximum likelihood estimate of `g(theta)` is `g(theta-hat)`. This is called **invariance**, or sometimes equivariance. So with `p-hat = 0.7` from 7 heads in 10 flips: - MLE of the odds `p/(1-p)` is `0.7/0.3 = 7/3 ~ 2.33` - MLE of the log-odds `log(p/(1-p))` is `log(7/3) ~ 0.847` - MLE of the probability of two heads in a row, `p^2`, is `0.49` - MLE of the failure probability `1-p` is `0.3` None of these requires writing a new likelihood or differentiating anything. ## Why it holds for a one-to-one transformation Picture the likelihood as a curve over the parameter axis. Applying a one-to-one `g` is a relabelling of that axis -- every parameter value gets a new name, and each new name corresponds to exactly one old one. Nothing about the heights of the curve changes, only the tick marks beneath it. The highest point of a curve does not move when you rename the axis, so the maximiser under the new labelling is precisely the new label of the old maximiser: `g(theta-hat)`. Written out: if `h = g(theta)` and `g` is invertible, the likelihood in the new parameter is `L*(h) = L(g-inverse(h))`. Maximising `L*` over `h` is the same as maximising `L` over `theta` and then applying `g`. There is no Jacobian term, because a likelihood is not a density in the parameter -- it is just a function being maximised, and function maximisation is unaffected by relabelling. That last sentence is the crux, and it is where the property is genuinely surprising. Transformations of *random variables* do pick up Jacobian factors when you change densities. The likelihood does not, precisely because it is not a distribution over the parameter. ## What about transformations that are not one-to-one? Squaring is not one-to-one on the whole real line: `mu` and `-mu` both map to `mu^2`. The property is extended by defining an **induced** (or profile) likelihood for the new parameter: the value assigned to `h` is the largest likelihood attainable over all `theta` with `g(theta) = h`. Under that definition invariance still holds, and `g(theta-hat)` is still the maximum likelihood estimate. So the shortcut is safe in either case; you just need the extended definition to make it precise. ## Where it earns its keep The practical value is that you fit once, in whichever parameterisation is convenient, and read off every derived quantity for free. Fit a rate and want a mean waiting time? Take the reciprocal. Fit a variance and want a standard deviation? Take the square root. Fit a probability and want odds, log-odds, or the chance of a run of successes? Transform the estimate. This is one of the reasons maximum likelihood is such a workhorse: the parameterisation is a modelling convenience, not a commitment. ## The limit of the shortcut Invariance is a statement about **where the maximum sits**, and about nothing else. A nonlinear `g` does not carry other properties across. Concretely, the average value of `g` of a random quantity is generally not `g` of its average -- squaring, taking reciprocals and taking logs all bend the scale, so quantities computed by averaging do not transform by simply applying `g`. Nor does an interval endpoint-transform for free in every case; you have to be explicit about what you are transforming. A weak answer treats invariance as a licence to transform any statistic at will. The correct claim is narrow: the maximiser transforms, because relabelling an axis cannot move a peak. ## Answering it crisply "By invariance, the MLE of any function of a parameter is that function of the MLE, so the odds estimate is `0.7/0.3`, about 2.33. It works because reparameterising relabels the parameter axis without changing the shape of the likelihood, so the location of the maximum is carried along unchanged."
- Why is there no Jacobian correction when you reparameterise a likelihood?Because the likelihood is not a density in the parameter -- it is a function being maximised, so there is nothing to preserve total mass for. Changing variables in a density requires a Jacobian to keep the integral at one; changing variables in a function you are only maximising requires nothing, since the maximum's location survives any relabelling of the axis.
- Does invariance hold when the transformation is not one-to-one, such as squaring?Yes, once the new parameter is given an induced likelihood: the value assigned to `h` is the largest likelihood over all parameter values mapping to `h`. Under that definition `g(theta-hat)` remains the maximiser. The shortcut is safe in practice; the extended definition is only needed to state it precisely.
- If the MLE of a variance is 16, what is the MLE of the standard deviation?It is 4, the square root of 16, by invariance -- the square root is one-to-one on non-negative values, so the fitted variance's peak carries straight over. There is no need to rewrite the log-likelihood in terms of the standard deviation and differentiate again.
Relabelling the parameter axis is like switching a thermometer from Celsius to Fahrenheit: the day's peak temperature is still the same moment, only its number is renamed.
saying these in an interview costs you the question
- Re-derives a new likelihood for every transformed parameter
- Adds a Jacobian factor when reparameterising a likelihood
- Says invariance fails for any transformation that is not linear
- Assumes the average of a transform equals the transform of the average
- Reports the odds as 0.3 over 0.7