skip to content

In a Cox model where region violates proportional hazards, when do you stratify rather than model a time-varying effect?

level: principalimportance: nice to knowfreq 32%

answer

  1. ask whether you need a number for it
  2. a separate baseline per group
  3. shared coefficients, within-group comparisons
  4. no hazard ratio for the stratifying variable
  5. pre-specify the remedy before looking

basics

~20 s

Stratify when region is a nuisance adjuster you never need an estimate for: each stratum gets its own baseline hazard while the other coefficients stay shared. Model a time-varying effect instead when region's changing effect is itself the finding.

solid answer

~50 s

The decision turns on whether you need a number for that variable. A stratified Cox fit gives each region its own unspecified baseline hazard and one shared set of coefficients for everything else, with comparisons made only within strata. Region then cannot violate anything — but you get no hazard ratio for it, no test of its effect, and less efficiency when strata are thin. So stratify a nuisance adjuster with few levels. If the region effect is the finding, model it: interact region with a function of time, or estimate piecewise hazard ratios by period so the report can say the effect was strong in month one and null afterwards. Keeping the averaged coefficient is defensible when the drift is small, provided you name the follow-up window it averages over. Pre-specify the choice — picking the remedy after seeing residuals is a forking path.

go deeper

for a junior

Know that a stratified Cox model gives each stratum its own baseline hazard and that no hazard ratio is produced for the variable you stratify on.

for a middle

Be able to contrast the two remedies mechanically: stratification frees the baseline per group, while a time interaction keeps one baseline and lets the coefficient move with follow-up time.

for a senior

Demonstrate the tradeoff on real data — thin strata widening every other interval, and the difference between a nuisance adjuster and the variable your stakeholders actually asked about.

for a principal

Own the analysis standard: which covariates are checked, what drift triggers which remedy, and why the rule must be written down before results are seen so post hoc window choices cannot manufacture a finding.

## The situation A Cox model of subscriber churn adjusts for region, and the diagnostic says region's hazard ratio is not constant over follow-up. Nothing is broken about the other coefficients yet, but the model as written asserts something false about region. There are three defensible responses and they answer different questions. ## Option 1: stratify on region A stratified Cox model splits subjects into strata and lets each stratum have its own arbitrary baseline hazard `h0s(t)`, while sharing one coefficient vector across strata: `h(t | x, stratum s) = h0s(t) * exp(b'x)` The partial likelihood becomes a product over strata, with risk sets formed inside each stratum. Because each region's baseline is unrestricted, region can affect the hazard in any time-varying way at all and the assumption cannot be violated by it. What you pay: - **No estimate for region.** There is no coefficient, no hazard ratio, no test. The variable has been adjusted away, not measured. If a stakeholder asks how much region matters, the model cannot answer. - **Efficiency loss with thin strata.** Comparisons happen only within strata. A stratum containing three events contributes almost nothing, and dozens of small strata can meaningfully widen every other interval. - **Proportionality is still assumed for everything else.** Stratifying on region does nothing for a treatment indicator whose own effect fades. - **Interactions become implicit.** If a covariate's effect genuinely differs by region, a shared coefficient still forces one number; stratification frees the baseline, not the slopes. Stratification is the right call when region is a nuisance adjuster with a handful of levels and adequate events in each, and nobody needs a regional effect size. ## Option 2: model the time-varying effect Make the coefficient a function of time. Two practical forms: - **Interaction with a time transform.** Add `region x log(t)` so the log hazard ratio is `b + g*log(t)`. Compact and testable, but it forces a specific functional shape on the drift. - **Piecewise-constant effects.** Split follow-up into pre-specified windows — days 0-30 and 31 onward — and estimate a separate hazard ratio in each. Less elegant, far easier to explain, and it maps directly onto a business statement: the effect was concentrated in the first month. This is the choice when the variable is the finding rather than the background. It converts a violated assumption into a result: not a nuisance to be absorbed, but a description of how the effect evolves. The cost is more parameters, decisions about window boundaries that must be pre-specified rather than chosen after seeing the data, and a report nobody can summarise in a single number. ## Option 3: keep the single coefficient, and caveat it When proportionality fails mildly, the fitted coefficient still estimates an average of the time-varying log hazard ratio over the observed follow-up, weighted by when events occurred. That is a genuine quantity and it may be exactly what a decision needs — but it is a property of your study window, not of the world. If you take this path, say so explicitly, give the follow-up length the average covers, and never let it be quoted as a stable long-run effect or extrapolated to a longer horizon. ## A fourth framing worth naming If several covariates violate proportionality at once, the model form itself may be wrong for the data. An accelerated failure time model makes covariates stretch or compress the time axis rather than scale the hazard — `log T = b'x + error` — and sometimes fits such data more naturally, at the price of committing to a distribution for the error term. Raising this shows you know proportional hazards is a modelling choice, not a law. ## The organisational part The deciding judgment at lead level is procedural as much as statistical. Write into the analysis plan, before looking at results, which covariates get proportionality checks, what magnitude of drift triggers a remedy, and which remedy applies to which kind of variable. Otherwise the team runs the check, sees a borderline result, and picks whichever fix produces the more agreeable headline — and every stratification or window boundary chosen after the fact quietly inflates the false-positive rate of the conclusion built on it. ## The answer to give Stratify nuisance variables you do not need to quantify. Model the time-varying effect for variables whose behaviour is the result. Keep the average only when it is small, honest and window-labelled. Decide the rule in advance.

  • What exactly do you lose by stratifying on region instead of adjusting for it?
    You lose any estimate or test of region's effect — there is no coefficient at all. You also lose efficiency when strata are many or thin, because comparisons occur only within a stratum. And proportional hazards is still assumed for every remaining covariate inside each stratum, so stratifying on region does nothing for a treatment effect that fades.
  • How would you report an effect that is strong in the first month and null afterwards?
    Report piecewise hazard ratios with the windows named — for example 0.55 over days 0 to 30 and 0.98 from day 31 onward — rather than one pooled number. State that the windows were pre-specified. A single hazard ratio here understates the early benefit, overstates the later one, and shifts with the length of follow-up.
  • Why should the proportionality check and its remedy be pre-specified?
    Because the remedy changes the headline. Choosing between stratifying, splitting follow-up, and keeping the pooled estimate after seeing residuals lets the analyst select the most flattering result, and window boundaries picked post hoc are an unreported multiple-comparison problem. Fixing the rule in advance makes the reported uncertainty mean what it claims.
  • When would you consider an accelerated failure time model instead?
    When several covariates violate proportional hazards at once, suggesting the multiplicative-hazard form is a poor fit rather than one variable misbehaving. An accelerated failure time model has covariates stretch or compress the time scale, `log T = b'x + error`, and its coefficients read as time ratios. The cost is committing to a distribution for the error term.

saying these in an interview costs you the question

  • Stratifies on the variable whose effect is the headline result
  • Thinks stratifying fixes proportionality for all covariates
  • Creates dozens of thin strata without noticing the efficiency loss
  • Picks follow-up windows after inspecting the residual plot
  • Expects a hazard ratio for the stratifying variable
  • Quotes an averaged hazard ratio without naming its follow-up window

context