How do you compute a Kaplan-Meier survival estimate by hand from a table of event and censoring times?
answer
- chain conditional probabilities, do not add
- denominator changes at every event time
- censored rows shrink it without a step
- product of (1 - d over n)
- steps get taller as follow-up thins
basics
~20 sAt each event time, divide the events by the number still at risk just before it and multiply the surviving fractions: S(t) = product of (1 - d_i / n_i). Censored subjects shrink the risk set without creating a step.
solid answer
~50 sKaplan-Meier is a product-limit estimator. Order the distinct times at which events occur. At each event time `t_i`, let `n_i` be the number of subjects still under observation and event-free just before `t_i`, and `d_i` the number of events at `t_i`. The conditional probability of getting through that instant is `1 - d_i / n_i`, and the survival estimate is the running product `S(t) = prod over t_i <= t of (1 - d_i / n_i)`. Censored subjects never produce a step; they simply leave the risk set after their censoring time, so later factors have smaller denominators and the steps grow taller as follow-up thins. By convention a subject censored at exactly an event time is still counted in that event's risk set. The curve is flat between event times and right-continuous, and it is an estimate of `S(t) = P(T > t)`, not a smooth model.
go deeper
Be ready to say the curve is a step function that drops only when an event occurs and stays flat otherwise, and that censored subjects do not cause drops.
You should be able to write S(t) as a product of (1 - d_i / n_i) and walk a small table row by row, tracking how the risk set shrinks at both events and censorings.
Talk about trust: shrinking risk sets, Greenwood's variance, why tail comparisons are fragile, and why a number-at-risk row belongs under every plot you show a stakeholder.
Decide when a non-parametric curve is the right deliverable at all versus a summary the business can act on, and set the conventions — time origin, event definition, follow-up window — that make curves comparable across teams.
## The idea Kaplan-Meier estimates `S(t) = P(T > t)` without assuming any shape for the distribution of event times. It does this by chaining conditional probabilities: surviving to time `t` means getting past the first event time, and then the second given you got past the first, and so on. Each link is estimated from the subjects who were actually being watched at that moment, which is what lets censored observations contribute honestly. ## The formula Let `t_1 < t_2 < ... < t_k` be the distinct times at which at least one event occurs. At `t_i`: - `n_i` = the **risk set** size: subjects still under observation and still event-free *just before* `t_i`. - `d_i` = the number of events at `t_i`. Then ``` S(t) = product over all t_i <= t of ( 1 - d_i / n_i ) ``` Before the first event time the estimate is 1. Between event times it is flat. It drops only at event times, which is why the plot is a step function rather than a curve. ## Where censoring enters Censored subjects appear only in the denominators. A subject censored at time `c` sits in `n_i` for every event time up to and including `c`, then disappears. It creates no step of its own, because nothing happened to it. Two conventions matter and are worth stating in an interview: - If a censoring time coincides with an event time, the censored subject is counted **in** that event's risk set (censoring is treated as occurring just after the event). - If the largest observed time is a censoring time, the curve simply stops at its current height; it is not extended to zero. ## A worked ten-subject example Ten subscribers, days from signup, `+` marks censored: ``` 5, 8, 8+, 12, 15+, 20, 22+, 30, 30+, 40+ ``` | event time | n_i | d_i | 1 - d_i/n_i | S(t) | |---|---|---|---|---| | 5 | 10 | 1 | 9/10 = 0.900 | 0.900 | | 8 | 9 | 1 | 8/9 = 0.889 | 0.800 | | 12 | 7 | 1 | 6/7 = 0.857 | 0.686 | | 20 | 5 | 1 | 4/5 = 0.800 | 0.549 | | 30 | 3 | 1 | 2/3 = 0.667 | 0.366 | Walk the denominators. Ten are at risk at day 5. After that event nine remain. At day 8 there is one event and one censoring; the censored subject is counted in the risk set of nine, and both leave afterwards, so seven are at risk at day 12. The censoring at 15 drops the risk set to five by day 20; the censoring at 22 leaves three by day 30, where again one event and one censoring occur together. The final subject is censored at 40, so the estimate ends at 0.366 and is not defined beyond day 40. Notice the steps: `10%`, then `11%`, then `14%`, then `20%`, then `33%` of the remaining height. Same one event each time, but a shrinking denominator makes each one matter more. ## Median survival Read off the smallest time where the estimate falls to 0.5 or below. In the table, survival is 0.549 after day 20 and 0.366 after day 30, so the estimated median is **30 days**. If the curve never reaches 0.5 within follow-up, the median is not estimable from the data. ## Uncertainty and the cumulative hazard The usual variance for the estimate is **Greenwood's formula**: ``` Var(S(t)) ~= S(t)^2 * sum over t_i <= t of d_i / ( n_i * (n_i - d_i) ) ``` Intervals built directly from it can stray outside `[0, 1]` in the tail, so implementations usually transform first (for example on the log-minus-log scale) before back-transforming. Plots should carry a number-at-risk row, because a step computed from `n_i = 3` deserves far less trust than one computed from `n_i = 300`. The companion quantity is the **cumulative hazard**, estimated by **Nelson-Aalen** as `sum of d_i / n_i` over event times up to `t`. The hazard function itself, `h(t)`, is the instantaneous event rate among those still at risk, and the two are linked by `S(t) = exp(-H(t))` where `H` is the cumulative hazard. Kaplan-Meier and Nelson-Aalen are close numerically when the per-step fractions are small. ## Common mistakes Adding the fractions instead of multiplying them; using the original sample size as the denominator at every event time; letting censored observations create steps; and forgetting that with no censoring at all Kaplan-Meier collapses to the plain empirical survival curve `1 - (number of events by t)/n`, which is a useful sanity check.
- What happens if a censoring time is exactly equal to an event time?The standard convention counts the censored subject in the risk set for that event, treating the censoring as occurring an instant later. It matters only in small samples, but stating the convention shows you know the estimator is defined on ties rather than silently guessing. Software follows this convention by default.
- Why does the curve become unreliable in the tail?Because the risk set shrinks. Late in follow-up each event divides by a handful of remaining subjects, so a single event can drop the estimate by a third, and Greenwood's variance grows accordingly. That is why survival plots carry a number-at-risk row and why tail comparisons between groups should be made cautiously.
- What does Kaplan-Meier reduce to when nothing is censored?The empirical survival curve: the fraction of the original sample whose event time exceeds t. With no censoring every risk set is just the number of subjects who have not yet had the event, and the telescoping product collapses to 1 minus the cumulative event proportion. It is a good check on a hand computation.
saying these in an interview costs you the question
- Adds the per-step fractions instead of multiplying them
- Uses the original sample size as the denominator at every event time
- Lets a censored observation create a downward step
- Draws the curve down to zero past the last censoring time
- Reads the curve as a fitted model rather than a step estimate