skip to content

In a Cox model, how do you handle a covariate that changes mid-follow-up, such as an upgrade to premium?

level: seniorimportance: should knowfreq 36%

answer

  1. one row is not one subject any more
  2. cut follow-up at the change point
  3. start and stop times per interval
  4. the event flag rides the last interval
  5. backdating the upgrade creates immortal time

basics

~20 s

Split the subject's follow-up into intervals at the moment the covariate changes, so one customer becomes two rows carrying start and stop times plus the value in force. The partial likelihood uses whichever value was current at each event time.

solid answer

~40 s

Use episode splitting, sometimes called the counting-process layout. A customer observed from day 0 to day 200 who upgrades on day 60 becomes two rows: `(0, 60, premium = 0)` and `(60, 200, premium = 1)`, with the event indicator attached only to the final interval. The subject is still one subject — the rows just describe which risk sets they belong to with which covariate value. Because the partial likelihood evaluates each risk-set member's covariates at the current event time, it automatically compares the customer as non-premium before day 60 and as premium after. The failure mode to avoid is labelling the customer premium for their entire history: that is immortal time bias, because they had to survive to day 60 to upgrade at all, and it makes premium look spuriously protective.

go deeper

for a junior

Know that a Cox model can accept covariates that change during follow-up, and that this requires restructuring the data into intervals rather than adding another column.

for a middle

Be able to lay out the split rows explicitly, with start and stop times, the covariate value in force for each interval, and the event flag on the final interval only.

for a senior

Show that you spot immortal time bias in a proposed analysis and can explain its direction, and that you know prediction now needs a landmark analysis or a model of the covariate path.

for a principal

Own the design decision: whether the question is really about current state at all, or whether a landmark or baseline-only analysis answers the business question with far less room for silent data errors.

## Why the standard layout is not enough A baseline Cox model attaches one fixed covariate vector to each subject. Plenty of real predictors are not fixed: a customer upgrades a plan, a user turns on a feature, a device changes firmware. Forcing such a variable into a single baseline value throws away when it changed and, worse, can invert the answer. ## Episode splitting The fix is mechanical. Represent each subject as a set of non-overlapping intervals `(start, stop, event, covariates)` covering their follow-up, cut at every moment a time-varying covariate changes value. A customer observed from day 0 to day 200 who upgrades to premium on day 60 and cancels on day 200 becomes: - `(0, 60, event = 0, premium = 0)` - `(60, 200, event = 1, premium = 1)` The event indicator is 1 only on the interval that actually ends in the event. Every other interval ends because the covariate changed, not because anything happened to the subject. This is not a way of turning one subject into two independent subjects. The partial likelihood is built from risk sets at event times, and each interval simply declares which risk sets the subject belongs to and with what covariate value. At an event on day 45 the customer appears with `premium = 0`; at an event on day 120 they appear with `premium = 1`. The exponentiated coefficient is then read as the hazard ratio between being premium right now and not being premium right now, comparing subjects at the same follow-up time. ## The trap: immortal time bias The tempting shortcut is to label a customer premium for their whole observation because they upgraded at some point. This is wrong in a specific, predictable direction. To upgrade on day 60, the customer had to still be a customer on day 60 — those first 60 event-free days are guaranteed by the definition of the group. Assigning them to the premium arm donates risk-free time to premium and makes upgrading look protective even if it does nothing at all. The interval layout removes the bias by construction, because the first 60 days are correctly attributed to the non-premium state. ## Time-varying covariate versus time-varying coefficient These are routinely confused and interviewers probe the distinction. - A **time-varying covariate** is a value `x(t)` that changes while the coefficient `b` stays fixed. The hazard is `h0(t) * exp(b * x(t))`. Proportionality in the usual sense is not violated; the comparison is simply between current states. - A **time-varying coefficient** is `b(t)` — the effect itself strengthens or fades with follow-up time, which is the proportional-hazards violation, modelled by interacting a fixed covariate with a function of time. Saying the words correctly is half the answer. ## What gets harder **Prediction.** With a fixed covariate you can draw a survival curve for a profile straight away. With `x(t)`, a survival curve requires the whole future covariate path, which you do not have at prediction time. Two common responses: model the covariate process as well, or use a landmark analysis — pick a landmark time, keep only subjects still under observation then, freeze their covariate state as of that moment, and model survival forward from there. **Interpretation of the covariate itself.** If the covariate can react to the subject's evolving state — a customer who is already disengaging is less likely to upgrade — then the current-state comparison is on shaky ground, and the interval layout fixes the bookkeeping without fixing that. **Data plumbing.** The splitting has to be exact. Intervals that overlap, gaps in coverage, or a covariate value recorded slightly after the change point are all silent errors that no diagnostic will catch. Reconcile the total person-time before and after splitting as a sanity check. ## The short answer Split at the change point, carry the value that was true in each interval, put the event on the last interval only, and never backdate a state the subject did not yet have.

  • What is the difference between a time-varying covariate and a time-varying coefficient in a Cox model?
    A time-varying covariate is a value that changes while its coefficient stays fixed — the hazard is `h0(t) * exp(b * x(t))`, and proportionality is not violated. A time-varying coefficient means the effect itself changes with follow-up time, which is the proportional-hazards violation and is modelled by interacting a fixed covariate with a function of time.
  • Why does labelling a customer as premium from day zero bias the estimate?
    It creates immortal time bias. To upgrade on day 60 the customer had to remain a customer through day 60, so those 60 event-free days are guaranteed by group membership. Crediting them to the premium state hands premium risk-free follow-up time and makes upgrading look protective even when it has no effect.
  • Why is prediction awkward once a covariate is time-varying?
    Drawing a survival curve requires the covariate's entire future path, which is unknown at prediction time. Either model the covariate process alongside survival, or run a landmark analysis: fix a landmark time, keep subjects still under observation then, freeze their state as of that moment, and model survival forward from the landmark.

It is like billing a phone plan that changes tariff mid-month. You do not retroactively charge the whole month at the new rate; you split the month at the switch date and bill each stretch at the rate that was actually in force.

saying these in an interview costs you the question

  • Assigns the post-change value to the whole follow-up
  • Treats each split interval as an independent subject
  • Puts the event indicator on every interval
  • Confuses a changing covariate with a changing coefficient
  • Expects a simple survival curve per group after splitting
  • Leaves gaps or overlaps between a subject's intervals

context