In a subscription churn analysis, how should you record a user who signed up 21 days ago and is still subscribed?
answer
- the row is not missing, it is bounded
- two fields: a time and a flag
- you know only that T exceeds it
- still counts in the denominator until it leaves
basics
~20 sRecord the user as right-censored at 21 days: the subscription lasted at least 21 days, and the eventual cancellation time is unknown. Deleting the row biases survival downward; coding it as a cancellation invents an event that never happened.
solid answer
~40 sIt is a right-censored observation. Every row in a time-to-event dataset carries two fields, a follow-up time and an event indicator, so this user is `(21, 0)` while a user who cancelled on day 14 is `(14, 1)`. Right-censoring is partial information, not absent information: you know the true lifetime exceeds 21 days, which is a real constraint the estimator can use. Kaplan-Meier consumes it directly — the user sits in the risk set for every event time up to day 21, then leaves the denominator without ever contributing an event. The two naive alternatives both bias the result: dropping censored rows throws away all the exposure of the newest cohort and understates survival, and coding the user as churned at day 21 fabricates an event and understates it even harder.
go deeper
Be ready to define right-censoring in one sentence and show the two-column layout: a follow-up time plus a 0/1 event indicator. Know that a still-active user is a 0, not a dropped row.
Explain what censoring does mechanically — the subject stays in the risk-set denominator until its censoring time, then leaves without producing a step — and name the direction of bias from dropping or mis-coding those rows.
Show you check whether censoring is informative before trusting the curve. Say how observation actually ends in your pipeline and whether that reason correlates with churn risk, and flag thin risk sets in the tail.
Own the definition of the time origin, the event, and the analysis cutoff for the whole organisation. Inconsistent choices across teams make survival numbers non-comparable far more often than the estimator does.
## The quantity being estimated Time-to-event analysis estimates the **survival function** `S(t) = P(T > t)`, the probability that the event of interest — cancellation, failure, relapse — has not happened by time `t`. `T` is the subject's true event time measured from a well-defined time origin (here, signup). `S(0) = 1` and `S` is non-increasing. The practical difficulty is that at the moment you run the analysis, most subjects have not had the event yet. Their `T` is not observed. What is observed is a follow-up time and a flag saying whether that time is an event or the end of observation. ## Right-censoring A subject is **right-censored** at time `c` when observation stops at `c` with the event not yet having occurred. The only thing known is `T > c`. Three common sources: - **Administrative censoring**: the analysis cutoff arrives. A user who signed up 21 days before the cutoff and is still subscribed is censored at 21 days purely because the calendar ran out. This is the most common form in product data, and it is a benign one — the cutoff date has nothing to do with any individual's propensity to cancel. - **Loss to follow-up**: the subject leaves the study for reasons unrelated to the event. - **Competing termination**: observation ends because something else ended it (an account deleted by an administrator rather than cancelled by the user). Contrast with **left-censoring** — the event is known to have happened *before* observation began, but the exact time is unknown (`T < c`) — and with **interval-censoring**, where the event is known only to fall between two visits. Right-censoring is the case interviewers mean when they say "censored" without qualification. ## Data layout A survival dataset row is `(time, event)`: - user cancelled on day 14 → `(14, 1)` - user signed up 21 days ago, still active at cutoff → `(21, 0)` The `time` column is *not* "time to cancel" for everyone; it is `min(true event time, censoring time)`, and the indicator says which of the two it was. Getting this pair right is most of the setup work in a churn analysis. ## Why the two naive fixes are wrong **Dropping censored rows.** The surviving subscribers are exactly the ones with long lifetimes, so keeping only rows with an observed cancellation and averaging their durations estimates the mean lifetime *among people who have already cancelled*. That number is systematically too small, and it gets smaller the more recent your cohort is, because young accounts can only appear in the retained set if they cancelled fast. It also silently changes the population you are describing. **Coding censored rows as events.** Recording `(21, 1)` says the user cancelled on day 21. Every still-active subscriber is then a fabricated cancellation at their current tenure, which drives the estimated survival curve to zero by the end of follow-up no matter how loyal the base is. **Calling it a missing value.** Censoring is not an absent measurement to be imputed. `T > 21` is information; a sound estimator uses the inequality rather than guessing a point value. ## How Kaplan-Meier uses the censored row The estimator works through the ordered event times. At each one it asks: of the subjects still under observation and event-free just before this instant, what fraction had the event? A censored subject contributes to that denominator for every event time up to its censoring time — it is genuine evidence that someone survived that long — and then quietly leaves. It never creates a step in the curve. That is precisely how the `T > 21` constraint is honoured without inventing a cancellation date. ## The assumption underneath All of this rests on **independent (non-informative) censoring**: a subject censored at time `c` should be representative of everyone still at risk at `c`. Administrative cutoffs generally satisfy this. What breaks it is censoring correlated with the event — for example, if accounts are removed from the dataset when a payment starts failing, and payment failure precedes cancellation, then censored subjects are worse-than-average risks and the estimated survival will be optimistic. State this assumption out loud in an interview; it is the part candidates most often skip.
- What is the difference between right-censoring and left-censoring?Right-censoring means observation stopped before the event, so the true time is known only to be greater than the recorded time. Left-censoring means the event had already happened before observation started, so the true time is known only to be less than the recorded time. Right-censoring is the default case in churn and reliability data; left-censoring shows up with detection thresholds and retrospective enrolment.
- When does censoring stop being harmless?When it is informative — when the reason observation ended is related to the risk of the event. Administrative cutoffs are fine because the calendar does not care who is about to cancel. But if accounts vanish from the dataset once billing starts failing, and billing failure precedes cancellation, the censored subjects are higher-risk than those left at risk, and every survival estimate comes out optimistic.
- How does a censored subject affect the width of the estimate late in follow-up?It removes one from the risk set at its censoring time, so every later event divides by a smaller denominator. The steps get taller and the confidence interval widens. The tail of a survival curve computed from a handful of remaining subjects is unstable, which is why plots normally carry a number-at-risk row underneath.
It is like a rope you stopped measuring at 21 metres because the tape ran out. You do not know the length, but you know it is longer than 21 metres, and averaging only the ropes that ended before the tape ran out gives a short answer.
saying these in an interview costs you the question
- Codes a still-active subscriber as a cancellation at the cutoff date
- Deletes censored rows so the dataset looks complete
- Treats censoring as a missing value to be imputed
- Averages only observed cancellation times and calls it mean lifetime
- Never mentions that censoring must be unrelated to the event risk