How would you check the proportional hazards assumption for a treatment indicator in a Cox model?
answer
- look for a trend against time
- who failed versus the risk-set average
- two transformed curves that stay parallel
- fit an explicit interaction with log time
- scaled Schoenfeld slope test
basics
~20 sPlot the scaled Schoenfeld residuals for treatment against time and test their slope; a trend means the hazard ratio is not constant. Confirm with log-minus-log survival curves, which should stay parallel, and by fitting a treatment-by-time interaction.
solid answer
~50 sThree checks that agree are stronger than any one. First, the Schoenfeld residual for treatment at each event time is the treated status of the subject who failed minus the risk-set weighted average; scaled and plotted against time, they should scatter flat around zero, and a test of nonzero slope (the Grambsch and Therneau test) formalises that. Second, plot `log(-log S(t))` for each treatment arm: under proportional hazards those curves are parallel with a vertical gap equal to `log(HR)`, so converging or crossing curves are a red flag. Third, refit with an explicit interaction between treatment and a function of time such as `log t` and see whether the interaction term is meaningful. Judge magnitude, not only the p-value: with tens of thousands of events a trivial departure will test significant, and with fifty events a real one will not.
go deeper
Know that a Cox model assumes the hazard ratio stays the same over time and that this is something you are expected to check rather than take on faith.
Be able to describe at least two concrete checks and what each looks like when the assumption holds: flat scaled residuals against time, and transformed survival curves that stay parallel.
Show diagnostic judgment on real data: separate a statistically significant but trivial drift from one that changes the conclusion, and rule out functional-form misspecification before blaming proportionality.
Own the analysis standard. Decide which covariates get checked, what magnitude of drift triggers a remedy, and how effects that vary over follow-up are communicated so nobody quotes a window-specific average as a permanent number.
## What is being assumed Proportional hazards says the hazard ratio between two covariate profiles is the same at every time point. For a binary treatment indicator, `h(t | treated) / h(t | control) = exp(b)` with no `t` on the right-hand side. The assumption fails when the effect changes over the follow-up — a retention intervention that strongly suppresses churn during the first month and does nothing afterwards is the classic case. ## Check 1: Schoenfeld residuals At each observed event time, the Schoenfeld residual for a covariate is the covariate value of the subject who actually had the event, minus the weighted average of that covariate across the risk set, weighted by each member's fitted risk score. Intuitively, it asks: at this moment, was the subject who failed more or less treated than the model expected, given who was still at risk? If the effect is constant, these residuals contain no time trend and scatter around zero. If treatment protects strongly early and not later, early residuals sit below zero (failures come disproportionately from controls) and later residuals drift up toward zero. Scaling the residuals by an estimate of the variance turns them into a direct estimate of the time-varying coefficient `b(t)`, so the plot can be read as the coefficient path. Fit a smoother through it: flat means proportional, a rising or falling line means not. The associated test — commonly attributed to Grambsch and Therneau — tests whether the slope of the scaled residuals against a time transform is zero, per covariate and globally. ## Check 2: log-minus-log curves Under proportional hazards, `S_treated(t) = S_control(t)^HR`. Take `-log` of both sides to get cumulative hazards related by a factor of `HR`, then take `log` again: `log(-log S_treated(t)) = log(HR) + log(-log S_control(t))` So plotting `log(-log S(t))` against time (or log time) for each arm should give two curves separated by a constant vertical distance of `log(HR)` — parallel, whatever shape they have. Curves that converge, diverge, or cross are direct visual evidence of a violation. Crossing is the most severe case: treatment helps early and hurts later, or vice versa, and no single hazard ratio can represent it. This check works cleanly for categorical covariates with a decent number of events per level; for continuous covariates you would have to bin, which weakens it. ## Check 3: fit the alternative The most direct test is to fit what the alternative claims. Add a term for treatment interacted with a function of time — `log t` is a common choice, or a piecewise indicator for period — so the coefficient becomes `b(t) = b + g * log t`. If `g` is indistinguishable from zero, the constant-effect model is not contradicted. This has the advantage of producing an interpretable description of the violation rather than just a warning: it tells you how the effect moves. ## Reading the result like a practitioner - **Power scales with events.** A test on 50,000 churn events will flag departures far too small to change any decision. A test on 40 events will pass almost anything. Always look at the residual plot alongside the p-value and ask how large the drift in the coefficient is over the period that matters. - **A violated assumption is not automatically fatal.** If proportionality fails, the single fitted coefficient still estimates a kind of average of the time-varying log hazard ratio across the observed follow-up, weighted by where the events fell. That average is a real quantity, but it is specific to your study's follow-up window, so it should not be quoted as if it were a stable property of the treatment. - **Check the covariates you care about first.** A mild violation on a nuisance adjuster matters much less than one on the treatment indicator whose effect is the headline result. - **Beware misspecification masquerading as non-proportionality.** A continuous covariate entered linearly when its true effect is curved, or an omitted strong predictor, can produce residual trends that look like a PH violation. Check the functional form before rewriting the model around time-varying effects. ## What a good answer sounds like Name at least two checks, say what each would look like under the null, and add the judgment layer: quantify the size of the departure, decide whether it changes the conclusion over the horizon the business cares about, and only then reach for a remedy.
- The scaled Schoenfeld residuals for treatment start well below zero and drift up to zero after the first month. What does that tell you?The treatment effect is front-loaded and wears off. Early on, failures come disproportionately from the control arm, so the residuals sit below zero; later the arms behave alike and the residuals centre on zero. The constant-effect assumption fails, and a single reported hazard ratio averages a strong first-month effect with a null one afterwards.
- Why might you accept a proportional hazards test that comes back significant?Because significance is not magnitude. With tens of thousands of events, a coefficient that drifts from 0.70 to 0.75 over the follow-up will test significant while changing no decision. Read the residual plot, quantify how much the hazard ratio moves across the horizon you care about, and check whether the conclusion flips. If it does not, report the average effect and state its window.
- Can something other than a genuine time-varying effect produce a residual trend?Yes. A misspecified functional form — a continuous covariate entered linearly when its effect is curved — or an important omitted predictor can both create apparent non-proportionality. Before restructuring the model around time-varying effects, check the functional form with splines and confirm the trend survives.
saying these in an interview costs you the question
- Treats a significant test as proof the model is unusable
- Checks only a global test and never plots residuals
- Confuses non-proportionality with a non-linear covariate effect
- Thinks crossing transformed survival curves still support one hazard ratio
- Reports a hazard ratio without stating the follow-up window it averages over
- Assumes a non-significant test proves proportionality despite few events