How does building a survival cohort only from customers who already survived 90 days distort the estimate?
answer
- selection happened before observation started
- early failures are absent, not censored
- denominators hold subjects that cannot fail
- each subject needs an entry time
- or restart the clock at day 90
basics
~20 sThe cohort is left-truncated: nobody in it could have cancelled before day 90, yet a signup clock puts them in the early risk sets anyway. Those denominators fill with subjects that cannot fail, so early survival is overstated.
solid answer
~50 sThis is left-truncation, also called delayed entry. Selecting customers who were still active at day 90 means every early cancellation in the original population is absent by construction. If you then measure time from signup and treat everyone as at risk from day 0, the risk sets covering days 0 to 90 are padded with subjects who had zero chance of producing an event there, so `d_i / n_i` is too small and the curve is optimistic exactly where the real churn happens. There are two correct handlings. Either let each subject enter the risk set only at its entry time, so it contributes to denominators from day 90 onward and not before — Kaplan-Meier and the risk-set machinery support this directly — or move the time origin to day 90 and say plainly that you are estimating survival *conditional on having reached 90 days*. Both are defensible; silently mixing the selection with a signup clock is not.
go deeper
Know the vocabulary: truncation removes subjects from the sample entirely, while censoring keeps a subject whose event has not been seen yet. That distinction alone answers most screening versions of this.
Explain the mechanics — which risk sets get padded, and why a padded denominator makes the estimated conditional failure probability too small — and describe the three-field entry, exit, event layout.
Demonstrate that you would catch this before running anything, by asking how the cohort table was assembled and whether membership required surviving a period, then choose and justify delayed entry or a shifted time origin.
Set the rule that cohorts are built from signup logs rather than snapshots of active accounts, and make sure conditional and unconditional survival curves are never presented on the same axis without labels.
## Truncation is not censoring Censoring and truncation both stop you from seeing a full event time, but they differ in what is in the dataset at all. - **Right-censoring**: the subject is in the sample, observed for a while, and the event has not happened yet. You know `T > c` for that individual. - **Left-truncation**: the subject is in the sample *only because* it had not had the event by its entry time. Individuals who failed before that threshold are absent entirely — you do not know they existed, and no row records them. - **Left-censoring**, a third and distinct thing, means the event is known to have happened before observation started but its exact time is unknown. The subject is present; the time is bounded above. The critical difference: censoring costs you the tail of one subject's information, while truncation changes which subjects exist in the sample. Truncation is a selection on the outcome, and selection on the outcome is the more dangerous of the two. ## The concrete failure Suppose a cohort is assembled from the customer table as "everyone who had an active subscription on the first of the month", and each subject's duration is measured from their signup date. A customer who signed up 200 days ago contributes a row with 200 days of tenure. Fine so far. But now build risk sets by tenure: At tenure day 10, who is in the denominator? Under the naive treatment, every selected customer — including all the ones who signed up long ago. Yet those long-tenured customers could not possibly have produced an event at tenure day 10 within this dataset, because if they had, they would not have been active on the first of the month and would never have been selected. The denominator is inflated with subjects who are, for that stretch of the time axis, immortal. `1 - d_i / n_i` is too close to 1, the early curve is too flat, and the analysis reports that the product retains far better in its first weeks than it does. The distortion is strongest exactly where subscription products have their real risk — the first days after signup — and it grows with the length of the survival requirement used to build the cohort. ## Fix one: delayed entry into the risk set Give every subject an **entry time** as well as an exit time and an event indicator: the row becomes `(entry, exit, event)`. The risk set at event time `t_i` is then everyone with `entry < t_i <= exit`. A customer selected at tenure 90 contributes nothing to the denominators before 90 and everything after. The product-limit formula is unchanged — only the definition of `n_i` widens — and the resulting curve estimates survival in the full population, using each subject over exactly the stretch of the time axis where it was genuinely observable. One consequence to check: risk sets early in the time axis may become very small or empty, because nobody was under observation there. A survival estimate over a stretch where the risk set is one or two subjects is not worth reporting, and with an empty risk set the estimate simply is not defined that early. Delayed-entry handling makes that honest rather than hiding it. ## Fix two: move the time origin Redefine `t = 0` as the moment of selection — day 90 — and estimate survival from there. Everyone in the cohort is genuinely at risk from that instant, and the estimator needs no special handling. The cost is interpretive: the curve now answers "given that a customer reached 90 days, how long do they last?" rather than "how long does a new customer last?". That is a perfectly useful question, and often the one the business meant. It just must be labelled, because the two curves are not comparable and the conditional one always looks better. ## How this shows up in practice - A cohort pulled from a "current customers" table rather than from a signup log. - A dataset built by joining an event log with a snapshot of active accounts. - Any analysis where the requirement to be in the sample is "still alive at the time we pulled the data", combined with a clock that starts earlier than the pull. - Backfilled histories where the event log only goes back a year, so anyone whose lifetime started earlier enters late. ## What an interviewer wants to hear Name the mechanism (left-truncation, delayed entry), state the direction of the bias (early survival overstated, the early part of the curve too flat), and give both remedies with their interpretive consequences. The strongest answers add the diagnostic: ask how the cohort table was built and whether membership in it required surviving anything.
- How does the dataset schema change to support delayed entry?Each row carries three fields instead of two: an entry time, an exit time, and an event indicator. The risk set at any event time is everyone whose entry is before it and whose exit is at or after it. The product-limit formula is untouched; only the definition of the denominator widens to respect entry.
- Which direction does the bias run if left-truncation is ignored?Survival is overstated, most severely in the early part of the time axis. Risk sets there are padded with subjects who could not have produced an event, so the estimated conditional failure probability at each early event time is too small and the curve is too flat exactly where the real risk concentrates.
- When is restarting the clock at the selection time the better choice?When the business question is genuinely conditional — how long an already-established customer lasts, how a renewal cohort behaves — and when the early risk sets under delayed entry would be too thin to support an estimate anyway. Label the result as conditional survival so nobody compares it against a curve that starts at signup.
It is like measuring how long cars last by surveying a parking lot: nothing that already broke down could have driven there, so the first years of the lifetime curve look flawless.
saying these in an interview costs you the question
- Calls the excluded early failures censored observations
- Starts the clock at signup while selecting on being active later
- Says the bias is small because the sample is large
- Imputes the unobserved pre-entry period for each subject
- Compares a conditional curve against one starting at signup