skip to content

In a logistic model, what does a coefficient of 18 with a standard error of 4000 tell you?

level: seniorimportance: nice to knowfreq 30%

answer

  1. the standard error is the giveaway
  2. cross-tabulate predictor against outcome
  3. some cell has zero cases
  4. the likelihood never stops improving
  5. no finite maximum exists

basics

~20 s

Separation is the usual cause: a predictor splits the outcome perfectly, so the likelihood has no finite maximum and the estimate runs away while its standard error explodes. Look for a zero cell, then use a penalized likelihood such as Firth's.

solid answer

~40 s

That is the signature of separation, not of a huge real effect. If every account with a corporate email domain converted and no other account did, the likelihood keeps improving as that coefficient grows — the maximum likelihood estimate is infinite, and the optimizer just stops at an iteration limit, leaving an arbitrary large number with an enormous standard error. Complete separation means a predictor splits the outcomes perfectly; quasi-complete means perfect apart from ties, usually a sparse category with a zero cell. Diagnose it by cross-tabulating the suspect predictor against the outcome and checking for fitted probabilities pinned at 0 or 1. Fix it by checking leakage first, then collapsing sparse categories, fitting a penalized likelihood such as Firth's, or reporting the perfect split as a finding.

go deeper

for a junior

Recognise the pattern rather than diagnose it: an implausibly large coefficient with a standard error in the thousands means the fit failed, not that a huge effect was found.

for a middle

Explain the mechanism — the likelihood keeps rising as the coefficient grows, so no finite maximum exists and the reported value is wherever the optimizer stopped. Name the zero-cell cross-tab check.

for a senior

Show the full diagnosis and response: check leakage first, inspect fitted probabilities and warnings, collapse sparse levels, and switch to a penalized likelihood with profile intervals rather than Wald ones.

for a principal

Own the policy angle. Decide when a perfectly predictive feature is reported as a data-quality finding versus modelled around, and make sure convergence warnings cannot be silenced without review.

## The symptom A logistic fit comes back with one coefficient at 18 (an odds ratio near 65 million) and a standard error in the thousands. The Wald p-value is a comfortable 0.99. Nothing about that combination describes a real effect: a genuinely enormous effect estimated from real data does not carry a standard error two hundred times its own size. ## The cause: separation Logistic regression is fitted by maximum likelihood — the software searches for the coefficients that make the observed 0/1 pattern most probable. **Complete separation** occurs when some linear combination of the predictors splits the outcomes perfectly: every row above a cut converted, every row below did not. **Quasi-complete separation** occurs when the split is perfect apart from ties sitting exactly on the boundary, which in practice usually means a sparse category with a zero cell in the predictor-by-outcome table. Under either condition the likelihood is monotone in that coefficient. Pushing the coefficient larger makes the fitted probabilities for the separated group approach 1 and the rest approach 0, which makes the data ever more probable — without limit. There is no finite maximiser. The reported estimate is therefore an artefact of where the optimizer gave up: change the convergence tolerance or the iteration cap and the number changes, which is a good diagnostic in itself. A concrete case: an offer-conversion model where every account with a corporate email domain converted. That one flag perfectly predicts the outcome, so its coefficient races upward and takes its standard error with it. ## Why the standard error explodes and the p-value lies The standard error comes from the curvature of the log-likelihood at the estimate. When the likelihood flattens out along that coefficient's direction — which is exactly what "increasing forever with diminishing gain" looks like — the curvature approaches zero and the estimated variance blows up. The Wald statistic divides the estimate by that standard error, and because the standard error grows faster than the estimate, the ratio collapses toward zero. This is the Hauck-Donner effect: the perfectly predictive variable is reported as insignificant. Trusting that p-value, or the Wald confidence interval built from it, is the deepest version of the mistake. A likelihood-ratio test or a profile-likelihood interval behaves far better here, because neither relies on a quadratic approximation to a likelihood that is not quadratic. ## Detection - **Cross-tabulate** each categorical predictor against the outcome and look for a cell with zero cases. That is the single highest-yield check. - **Inspect fitted probabilities.** Values numerically indistinguishable from 0 or 1 are the fingerprint. - **Read the warnings.** "Did not converge" or "fitted probabilities of 0 or 1 occurred" is not noise to be silenced by raising the iteration limit. - **Refit with a tighter tolerance.** If the coefficient grows again, it is diverging rather than converging. - Watch for it especially with rare outcomes, many-level categoricals, small samples, and models with many predictors relative to events. ## What to do about it **1. Check for leakage first.** A variable that perfectly predicts the outcome is far more often a recording artefact than a discovery: a field populated only after the event, a status flag derived from the outcome, a support ticket that exists only for churned accounts. Fix the data pipeline rather than the estimator. This check comes before any statistical remedy. **2. Collapse or drop sparse categories.** A 40-level categorical with a handful of rows in some levels will separate almost mechanically. Grouping small levels into a sensible "other" bucket often removes the zero cell and restores a finite fit. **3. Use a penalized likelihood.** Firth's method adds a bias-reduction penalty to the likelihood that guarantees finite estimates even under complete separation, and it also reduces the small-sample bias of ordinary maximum likelihood. It is the standard remedy when the separating predictor is genuinely meaningful and must stay in the model. Report profile-likelihood intervals alongside it rather than Wald ones. **4. Consider exact methods for very small samples**, where the asymptotic approximations underlying the usual inference are shaky anyway. **5. Or simply report the separation.** If corporate-domain accounts convert 100% of the time in your data, the honest summary may be that finding plus its sample size, not a coefficient. "12 of 12 converted" tells a stakeholder more than an odds ratio of 65 million, and it makes the fragility of the evidence visible. ## What not to do Do not raise the iteration cap and call it converged — that changes the number without changing the problem. Do not report the coefficient or its odds ratio at face value. Do not conclude the predictor is unimportant because its Wald p-value is large; the p-value is the artefact, not the verdict. And do not assume more data will rescue the fit: if the separation reflects a structural relationship or leakage, extra rows reproduce it rather than break it.

  • Why is the Wald p-value for that coefficient usually large and non-significant?
    Because the standard error grows even faster than the estimate as the likelihood flattens, so the Wald ratio collapses toward zero. This is the Hauck-Donner effect, and it reports a perfectly predictive variable as insignificant. Use a likelihood-ratio test or a profile-likelihood interval instead, since neither depends on a quadratic approximation that no longer holds.
  • How does quasi-complete separation differ from complete separation?
    Complete separation means some combination of predictors splits every outcome perfectly. Quasi-complete separation means the split is perfect except for ties sitting exactly on the boundary, and it typically arises from a sparse category with a zero cell. Both make the maximum likelihood estimate non-finite, so both produce the runaway coefficient and inflated standard error.
  • What non-statistical explanation should you rule out before applying any fix?
    Leakage. A field populated only after the outcome occurs, or derived from it, will separate perfectly by construction — a churn-only support flag, a status column updated at cancellation. Check when each variable is recorded relative to the outcome before reaching for a penalty, because a penalized fit on leaked data just hides the real defect.
  • Would collecting more data resolve separation?
    Not reliably. If the separation comes from a structural relationship or leakage, more rows reproduce the perfect split rather than break it. Extra data helps only in the specific case where separation was a small-sample accident in a sparse category, and even then a penalized fit is the more dependable route.

saying these in an interview costs you the question

  • Blames the optimizer and raises the iteration limit
  • Reads the huge coefficient as an enormous real effect
  • Trusts the Wald p-value from a separated fit
  • Ignores a zero cell in the predictor-by-outcome table
  • Assumes more data always resolves separation
  • Never checks whether the predictor leaks the outcome

context