Why does fitting a Cox proportional hazards model never require specifying the baseline hazard?
answer
- no shape assumed for time itself
- only the ordering of events matters
- compare the failure against the risk set
- a common factor cancels in the ratio
- product of risk-set ratios, one per event
basics
~20 sCoefficients come from a partial likelihood built only from who fails first: at each event time it compares the failing subject's risk score against everyone still at risk. The baseline hazard is a common factor in that ratio and cancels.
solid answer
~40 sThe Cox model is semi-parametric: the covariate part `exp(b'x)` is parametric, while the baseline hazard `h0(t)` is left as an arbitrary function of time. Fitting uses the partial likelihood, which contributes one factor per observed event time: the failing subject's `exp(b'x_i)` divided by the sum of `exp(b'x_j)` over every subject still at risk at that instant. Both numerator and denominator carry `h0(t)` as a common multiplier, so it cancels and never has to be estimated. The information used is purely the ordering of events within risk sets — who failed first among those still eligible — which is why the effective sample size is the number of events, not the number of subjects. If you later want absolute survival curves, you estimate the baseline cumulative hazard separately, typically with the Breslow estimator.
go deeper
Be able to say the model is semi-parametric: covariate effects are modelled, the shape of the hazard over time is not, and that is why no distribution for survival times is assumed.
Expect to write the partial-likelihood factor for one event time and point at exactly where the baseline hazard cancels. Interviewers probe whether you can do the algebra, not just recite the label.
Demonstrate the practical consequences: sizing analyses by event counts, choosing a tie-handling method for coarse time grids, and knowing that absolute risk needs a separately estimated baseline.
Own the choice between a Cox fit and a fully parametric survival model for the team: robustness against baseline misspecification versus the extrapolation and simulation you need when forecasting beyond the observation window.
## Semi-parametric by construction The Cox model specifies `h(t | x) = h0(t) * exp(b1*x1 + ... + bp*xp)` and deliberately says nothing about `h0(t)`. It may rise, fall, spike early, or be flat. That is the semi-parametric bargain: a parametric, log-linear form for how covariates shift the hazard, paired with a completely unspecified shape for how the hazard moves through time. Contrast a fully parametric survival model, where you commit to a family such as exponential (constant hazard) or Weibull (monotone hazard) and estimate its shape parameters alongside the covariate effects. ## The partial likelihood Order the distinct observed event times `t_1 < t_2 < ... < t_k`. At each `t_m`, define the risk set `R(t_m)` as the subjects who are still under observation and have not yet had the event just before `t_m`. If subject `i` is the one who has the event at `t_m`, that time contributes the factor `L_m(b) = exp(b'x_i) / sum over j in R(t_m) of exp(b'x_j)` and the partial likelihood is the product of these factors across all event times. The factor answers a conditional question: given that exactly one event happened at `t_m`, and given who was still at risk, what is the probability that it was subject `i` rather than one of the others? Under the model that probability is each subject's hazard divided by the total hazard in the risk set — and every subject's hazard contains the same `h0(t_m)`, so it cancels top and bottom. Time itself appears only through membership in the risk sets. Maximising the log of this product gives the coefficient estimates. ## What this buys and what it costs **Buys:** robustness. You cannot get the coefficients wrong by misspecifying the shape of the baseline hazard, because you never assumed one. That is the main reason the Cox model became the default in survival analysis. **Costs:** - *No free absolute predictions.* The fitted coefficients give ratios only. To draw a predicted survival curve you need the baseline cumulative hazard, usually estimated after the fact by the Breslow estimator, which sums `1 / (sum of exp(b'x_j) over the risk set)` across event times. - *A small efficiency loss* relative to a correctly specified parametric model. If you genuinely know the hazard is Weibull, the parametric fit uses more information and gives slightly tighter intervals — but you pay for that only if the assumption is true. - *No extrapolation.* The baseline is only estimated where events were observed, so the model cannot say anything about times beyond the last observed event. ## Events are the currency Because the partial likelihood contributes one factor per event time, subjects who never have an observed event supply no factor of their own; they only enlarge the denominators of risk sets while they remain under observation. This is why practitioners size studies by the number of events rather than the number of subjects, and why the common rule of thumb asks for roughly ten events per covariate before the coefficient estimates are stable. A dataset of a million customers with forty churn events supports about four covariates, not four hundred. ## Ties The derivation above assumes one event at a time. Real data record time on a coarse grid, so several subjects share an event time. The exact treatment sums over the possible orderings of the tied events, which is expensive when ties are heavy. Two approximations are standard: Breslow's, which is simpler and biases coefficients toward zero when ties are numerous, and Efron's, which is more accurate and is the sensible default. If time is recorded so coarsely that most events are tied — monthly buckets, for instance — a discrete-time model is a better match than a Cox fit patched up with a tie correction. ## The one-line answer to give Say it as a cancellation argument: the partial likelihood is built from ratios within risk sets, the baseline hazard is a common factor in every term of the ratio, so it divides out and the coefficients can be estimated without ever writing down a functional form for `h0(t)`.
- What does the Cox model give up by leaving the baseline hazard unspecified?Absolute predictions. The coefficients yield ratios only, so producing a predicted survival curve requires estimating the baseline cumulative hazard separately, typically with the Breslow estimator. There is also a modest efficiency loss relative to a correctly specified parametric model such as a Weibull fit, and no ability to extrapolate beyond the last observed event time.
- Why do people say a Cox model's effective sample size is the number of events, not the number of subjects?Each observed event time contributes exactly one factor to the partial likelihood. Subjects who never have an event contribute no factor of their own; they only sit in the denominators of risk sets. So precision is driven by event counts, which is why the usual rule of thumb asks for around ten events per covariate rather than ten subjects.
- How do tied event times affect the partial likelihood?The derivation assumes one event per time point, so ties need handling. The exact method sums over all orderings of the tied events and is costly. Breslow's approximation is simplest but shrinks coefficients toward zero under heavy ties; Efron's is more accurate and the better default. If time is bucketed so coarsely that most events tie, use a discrete-time model instead.
saying these in an interview costs you the question
- Thinks the Cox model assumes exponential or Weibull survival times
- Says the baseline hazard is estimated first, then the coefficients
- Calls the partial likelihood a full likelihood over all subjects
- Believes the fitted coefficients alone yield survival probabilities
- Sizes the study by subjects rather than by events
- Ignores tie handling when event times are coarsely recorded