skip to content

Why use a log-rank test on time-to-cancel instead of comparing day-30 cancellation rates between two onboarding flows?

level: seniorimportance: should knowfreq 42%

answer

  1. compares curves, not one milestone
  2. observed minus expected at every event time
  3. risk sets carry the censored subjects
  4. signed contributions cancel when curves cross
  5. gives a p-value, not an effect size

basics

~20 s

The log-rank test compares whole survival curves, accumulating observed minus expected events at every event time and keeping partially observed users through the risk sets. A single day-30 rate discards timing and every account younger than 30 days.

solid answer

~50 s

A day-30 rate is one point on a curve. It cannot separate a flow whose users leave in the first week from one whose users leave on day 29, and it forces you to drop every account younger than 30 days or mislabel it as retained. The log-rank test uses all of it. At each time an event occurs it computes how many cancellations to expect in each group if the two survival distributions were identical, given each group's share of the risk set at that moment, then sums observed minus expected across event times and refers the standardised total to a chi-square distribution with one degree of freedom for two groups. Recent signups contribute to every risk set up to their censoring time and then drop out, so nobody is discarded and no event is invented. Only the ordering of event times is used, so it is non-parametric.

go deeper

for a junior

Recall that the log-rank test compares two whole survival curves rather than a single milestone, and that it handles subjects whose outcome has not been observed yet.

for a middle

Explain the machinery: expected events per group at each event time from the at-risk split, summed observed minus expected, and the fact that only the order of event times is used.

for a senior

Show you would plot first, spot crossing curves, check that censoring behaves similarly in both arms, and pair the p-value with a landmark or restricted-mean difference a stakeholder can act on.

for a principal

Decide when a whole-curve comparison is the right decision instrument at all, fix the analysis window and any weighting in advance, and stop teams from shopping among tests and landmarks for a significant result.

## What the test asks The log-rank test is a hypothesis test for the null that two groups have **identical survival distributions** over the whole follow-up period. It is not a test of a point estimate or a single milestone; it is a comparison of two curves. ## The mechanism Walk the pooled event times in order. At event time `t_i`: - `n_i` subjects are at risk overall, of which `n_Ai` belong to group A; - `d_i` events occur, of which `d_Ai` are in group A. Under the null hypothesis, the group membership of the events at `t_i` is just a random split of `d_i` among the `n_i` subjects at risk, so the expected number of group-A events is ``` E_Ai = d_i * ( n_Ai / n_i ) ``` with a variance given by the hypergeometric distribution of that split. Sum across event times to get `O_A = sum d_Ai`, `E_A = sum E_Ai`, and `V = sum of the per-time variances`. The statistic ``` (O_A - E_A)^2 / V ``` is referred to a chi-square distribution with one degree of freedom when there are two groups (`k - 1` degrees of freedom for `k` groups). A large positive `O_A - E_A` means group A produced more cancellations than an identical-curves world predicts. Two structural features fall out of this. First, censoring is handled through the risk sets: a subject who is still active contributes to `n_i` and `n_Ai` at every event time up to its censoring time and then leaves, contributing no event. Nobody is dropped and nobody is coded as a cancellation. Second, only the **ordering** of the event times matters, not their numeric values — the test is non-parametric and assumes no distributional shape. ## Why a single-day rate is weaker **It discards timing.** Two flows can produce identical day-30 cancellation percentages while one bleeds users in the first 72 hours and the other loses them slowly across the month. Those are different product problems with different fixes, and the point comparison cannot separate them. The log-rank statistic accumulates the discrepancy at every event time, so an early divergence registers even if the totals converge later. **It forces a censoring decision you do not want to make.** To compute a day-30 rate you need each account to have 30 days of history. Accounts younger than that must either be excluded — which throws away your most recent and often most relevant users, and delays every read by a month — or counted as retained, which is a fabrication. Survival analysis takes them as right-censored and uses the exposure they do have. **It wastes information and therefore power.** Every cancellation carries a time, and the point comparison collapses all of them into a single binary. Given the same data, the whole-curve comparison detects a real difference with a smaller sample. ## Where the log-rank is weak - **Crossing curves.** The test sums signed discrepancies. If flow A is worse early and better late, the positive and negative contributions cancel and the total can land near zero even though the curves are visibly different. Always plot before testing; a crossing pattern calls for a landmark comparison or a split-window analysis instead of a single p-value. - **Differences concentrated where nobody is left.** A gap that opens only in the thin tail carries little weight because the risk sets there are tiny. - **Unequal censoring patterns.** The test assumes censoring is unrelated to the risk of the event and behaves similarly in both groups. If one flow's users disappear from the dataset for reasons tied to their likelihood of cancelling, the comparison is contaminated. - **It returns no effect size.** The output is a p-value. Pair it with something a stakeholder can act on: the two Kaplan-Meier curves with intervals, a difference in survival at a pre-specified landmark, or a difference in restricted mean event-free days over a fixed horizon. ## Weighted variants The standard log-rank weights every event time equally. Weighted variants exist that emphasise early event times — the Gehan-Breslow generalised Wilcoxon test weights each time by the size of the risk set, so differences observed while many subjects remain count for more. Choosing the weighting after seeing which one gives a smaller p-value is a multiplicity problem; fix it in advance based on where you believe a difference should appear. ## In an interview Say what the null is, describe the observed-minus-expected accumulation over risk sets, explain how censored subjects participate, and then name the two honest caveats: crossing curves and the absence of an effect size. Candidates who describe the log-rank as "a t-test for survival times" have missed both the censoring handling and the rank-based construction.

  • What does the log-rank test do when the two survival curves cross?
    It loses power. The statistic sums signed observed-minus-expected contributions across event times, so a group that is worse early and better late produces terms that cancel, and the total can sit near zero despite an obvious visual difference. Plot the curves first; when they cross, report pre-specified landmark comparisons or split the follow-up rather than leaning on one p-value.
  • How do censored subjects participate in the log-rank calculation?
    Through the risk sets. A subject still active at time t counts in both the overall and the group-specific at-risk counts for every event time up to its censoring time, which is what determines the expected split of events there. After that it leaves the denominators and never contributes an event. That is how recent signups inform the test without being dropped or mislabelled.
  • The log-rank returns p = 0.01. What else do you report?
    An effect size and the picture. Show both Kaplan-Meier curves with confidence intervals and a number-at-risk row, and quantify the gap with a pre-specified landmark difference such as survival at 30 days, or a difference in restricted mean event-free days over a fixed horizon. A p-value alone tells a stakeholder that something differs but not whether it matters.

saying these in an interview costs you the question

  • Describes the log-rank as a t-test on survival times
  • Ignores that censored subjects contribute through the risk sets
  • Reads a large p-value as proof the curves are identical when they cross
  • Reports the p-value with no curve and no effect size
  • Drops accounts with under 30 days of history to compute a rate

context