skip to content

Generalized Linear Models

A link function sits between the linear predictor and the mean so the model can handle binary and count outcomes, fitted by maximum likelihood. Logistic and Poisson are the cases interviewers expect.

on this pageshow

explore

questions

20

In a subscription churn analysis, how should you record a user who signed up 21 days ago and is still subscribed?

level: juniorimportance: must knowfreq 78%

answer

  1. the row is not missing, it is bounded
  2. two fields: a time and a flag
  3. you know only that T exceeds it
  4. still counts in the denominator until it leaves

basics

~20 s

Record the user as right-censored at 21 days: the subscription lasted at least 21 days, and the eventual cancellation time is unknown. Deleting the row biases survival downward; coding it as a cancellation invents an event that never happened.

solid answer

~40 s

It is a right-censored observation. Every row in a time-to-event dataset carries two fields, a follow-up time and an event indicator, so this user is `(21, 0)` while a user who cancelled on day 14 is `(14, 1)`. Right-censoring is partial information, not absent information: you know the true lifetime exceeds 21 days, which is a real constraint the estimator can use. Kaplan-Meier consumes it directly — the user sits in the risk set for every event time up to day 21, then leaves the denominator without ever contributing an event. The two naive alternatives both bias the result: dropping censored rows throws away all the exposure of the newest cohort and understates survival, and coding the user as churned at day 21 fabricates an event and understates it even harder.

go deeper

for a junior

Be ready to define right-censoring in one sentence and show the two-column layout: a follow-up time plus a 0/1 event indicator. Know that a still-active user is a 0, not a dropped row.

for a middle

Explain what censoring does mechanically — the subject stays in the risk-set denominator until its censoring time, then leaves without producing a step — and name the direction of bias from dropping or mis-coding those rows.

for a senior

Show you check whether censoring is informative before trusting the curve. Say how observation actually ends in your pipeline and whether that reason correlates with churn risk, and flag thin risk sets in the tail.

for a principal

Own the definition of the time origin, the event, and the analysis cutoff for the whole organisation. Inconsistent choices across teams make survival numbers non-comparable far more often than the estimator does.

## The quantity being estimated Time-to-event analysis estimates the **survival function** `S(t) = P(T > t)`, the probability that the event of interest — cancellation, failure, relapse — has not happened by time `t`. `T` is the subject's true event time measured from a well-defined time origin (here, signup). `S(0) = 1` and `S` is non-increasing. The practical difficulty is that at the moment you run the analysis, most subjects have not had the event yet. Their `T` is not observed. What is observed is a follow-up time and a flag saying whether that time is an event or the end of observation. ## Right-censoring A subject is **right-censored** at time `c` when observation stops at `c` with the event not yet having occurred. The only thing known is `T > c`. Three common sources: - **Administrative censoring**: the analysis cutoff arrives. A user who signed up 21 days before the cutoff and is still subscribed is censored at 21 days purely because the calendar ran out. This is the most common form in product data, and it is a benign one — the cutoff date has nothing to do with any individual's propensity to cancel. - **Loss to follow-up**: the subject leaves the study for reasons unrelated to the event. - **Competing termination**: observation ends because something else ended it (an account deleted by an administrator rather than cancelled by the user). Contrast with **left-censoring** — the event is known to have happened *before* observation began, but the exact time is unknown (`T < c`) — and with **interval-censoring**, where the event is known only to fall between two visits. Right-censoring is the case interviewers mean when they say "censored" without qualification. ## Data layout A survival dataset row is `(time, event)`: - user cancelled on day 14 → `(14, 1)` - user signed up 21 days ago, still active at cutoff → `(21, 0)` The `time` column is *not* "time to cancel" for everyone; it is `min(true event time, censoring time)`, and the indicator says which of the two it was. Getting this pair right is most of the setup work in a churn analysis. ## Why the two naive fixes are wrong **Dropping censored rows.** The surviving subscribers are exactly the ones with long lifetimes, so keeping only rows with an observed cancellation and averaging their durations estimates the mean lifetime *among people who have already cancelled*. That number is systematically too small, and it gets smaller the more recent your cohort is, because young accounts can only appear in the retained set if they cancelled fast. It also silently changes the population you are describing. **Coding censored rows as events.** Recording `(21, 1)` says the user cancelled on day 21. Every still-active subscriber is then a fabricated cancellation at their current tenure, which drives the estimated survival curve to zero by the end of follow-up no matter how loyal the base is. **Calling it a missing value.** Censoring is not an absent measurement to be imputed. `T > 21` is information; a sound estimator uses the inequality rather than guessing a point value. ## How Kaplan-Meier uses the censored row The estimator works through the ordered event times. At each one it asks: of the subjects still under observation and event-free just before this instant, what fraction had the event? A censored subject contributes to that denominator for every event time up to its censoring time — it is genuine evidence that someone survived that long — and then quietly leaves. It never creates a step in the curve. That is precisely how the `T > 21` constraint is honoured without inventing a cancellation date. ## The assumption underneath All of this rests on **independent (non-informative) censoring**: a subject censored at time `c` should be representative of everyone still at risk at `c`. Administrative cutoffs generally satisfy this. What breaks it is censoring correlated with the event — for example, if accounts are removed from the dataset when a payment starts failing, and payment failure precedes cancellation, then censored subjects are worse-than-average risks and the estimated survival will be optimistic. State this assumption out loud in an interview; it is the part candidates most often skip.

  • What is the difference between right-censoring and left-censoring?
    Right-censoring means observation stopped before the event, so the true time is known only to be greater than the recorded time. Left-censoring means the event had already happened before observation started, so the true time is known only to be less than the recorded time. Right-censoring is the default case in churn and reliability data; left-censoring shows up with detection thresholds and retrospective enrolment.
  • When does censoring stop being harmless?
    When it is informative — when the reason observation ended is related to the risk of the event. Administrative cutoffs are fine because the calendar does not care who is about to cancel. But if accounts vanish from the dataset once billing starts failing, and billing failure precedes cancellation, the censored subjects are higher-risk than those left at risk, and every survival estimate comes out optimistic.
  • How does a censored subject affect the width of the estimate late in follow-up?
    It removes one from the risk set at its censoring time, so every later event divides by a smaller denominator. The steps get taller and the confidence interval widens. The tail of a survival curve computed from a handful of remaining subjects is unstable, which is why plots normally carry a number-at-risk row underneath.

It is like a rope you stopped measuring at 21 metres because the tape ran out. You do not know the length, but you know it is longer than 21 metres, and averaging only the ropes that ended before the tape ran out gives a short answer.

saying these in an interview costs you the question

  • Codes a still-active subscriber as a cancellation at the cutoff date
  • Deletes censored rows so the dataset looks complete
  • Treats censoring as a missing value to be imputed
  • Averages only observed cancellation times and calls it mean lifetime
  • Never mentions that censoring must be unrelated to the event risk

context

open as a page

In logistic regression, what is the difference between odds and probability?

level: juniorimportance: must knowfreq 78%

basics

~10 s

Probability is p, a number between 0 and 1. Odds is p/(1-p), the event weighed against its complement, and runs from 0 to infinity. A probability of 0.8 is odds of 4, or 4-to-1.

open as a page

Why model event counts with Poisson regression instead of ordinary linear regression?

level: juniorimportance: must knowfreq 70%

basics

~20 s

Counts are non-negative integers, usually skewed, and their spread grows with their level. Poisson regression models the log of the expected count, so fitted values stay positive and each predictor acts multiplicatively rather than adding a fixed number of events.

open as a page

How do you compute a Kaplan-Meier survival estimate by hand from a table of event and censoring times?

level: middleimportance: must knowfreq 70%

basics

~20 s

At each event time, divide the events by the number still at risk just before it and multiply the surviving fractions: S(t) = product of (1 - d_i / n_i). Censored subjects shrink the risk set without creating a step.

open as a page

In a Cox proportional hazards model of churn, what does a hazard ratio of 1.8 on an annual-plan flag mean?

level: middleimportance: must knowfreq 70%

basics

~10 s

Annual-plan customers cancel at 1.8 times the instantaneous rate of the reference group, at every time point the model considers. A hazard ratio compares rates among those still subscribed; it is not a probability.

open as a page

How do you turn a logistic regression coefficient of 0.693 into an odds ratio?

level: middleimportance: must knowfreq 70%

basics

~20 s

Exponentiate the coefficient: exp(0.693) is about 2, so each one-unit rise in that predictor multiplies the odds of the outcome by 2, with the other predictors held fixed. Negative coefficients exponentiate to odds ratios below 1.

open as a page

A Poisson regression coefficient is 0.18. How do you interpret it on the count scale?

level: middleimportance: must knowfreq 75%

basics

~10 s

Exponentiate it. exp(0.18) is about 1.20, so a one-unit increase in that predictor multiplies the expected count by roughly 1.2 — about 20% more events per unit of exposure, holding the other predictors fixed.

open as a page

How would you check the proportional hazards assumption for a treatment indicator in a Cox model?

level: seniorimportance: must knowfreq 58%

basics

~20 s

Plot the scaled Schoenfeld residuals for treatment against time and test their slope; a trend means the hazard ratio is not constant. Confirm with log-minus-log survival curves, which should stay parallel, and by fitting a treatment-by-time interaction.

open as a page

When a Kaplan-Meier curve never drops to 0.5, what does 'median survival not reached' mean?

level: middleimportance: should knowfreq 46%

basics

~20 s

It means follow-up ended before half the cohort had the event, so no time exists at which estimated survival is 0.5 or lower. The median is not estimable from the data, and extrapolating past the last observed time invents an answer.

open as a page

Why does fitting a Cox proportional hazards model never require specifying the baseline hazard?

level: middleimportance: should knowfreq 54%

basics

~20 s

Coefficients come from a partial likelihood built only from who fails first: at each event time it compares the failing subject's risk score against everyone still at risk. The baseline hazard is a common factor in that ratio and cancels.

open as a page

Why does a linear probability model fitted with OLS fail on a binary outcome?

level: middleimportance: should knowfreq 52%

basics

~20 s

A straight line on a 0/1 outcome has no ceiling or floor, so it predicts impossible values like -0.12 and 1.30. The outcome's variance p(1-p) also changes with the predictors, so the errors are heteroscedastic by construction.

open as a page

Why does a Poisson regression on counts observed over unequal exposure need an offset?

level: middleimportance: should knowfreq 58%

basics

~20 s

Units observed longer accumulate more events for no interesting reason. Adding the log of exposure as an offset, with its coefficient fixed at 1, makes the model describe the event rate per unit of exposure instead of the raw count.

open as a page

Why use a log-rank test on time-to-cancel instead of comparing day-30 cancellation rates between two onboarding flows?

level: seniorimportance: should knowfreq 42%

basics

~20 s

The log-rank test compares whole survival curves, accumulating observed minus expected events at every event time and keeping partially observed users through the risk sets. A single day-30 rate discards timing and every account younger than 30 days.

open as a page

In a Cox model, how do you handle a covariate that changes mid-follow-up, such as an upgrade to premium?

level: seniorimportance: should knowfreq 36%

basics

~20 s

Split the subject's follow-up into intervals at the moment the covariate changes, so one customer becomes two rows carrying start and stop times plus the value in force. The partial likelihood uses whichever value was current at each event time.

open as a page

Your churn model reports an odds ratio of 2.0 — how do you explain that to a product manager?

level: seniorimportance: should knowfreq 48%

basics

~20 s

An odds ratio of 2 doubles the odds, not the probability. From a 1% baseline that lands near 2%, but from a 40% baseline it lands near 57%, not 80%. Report an average marginal effect in percentage points instead.

open as a page

Page views per session have variance about eight times the mean. What breaks in a Poisson regression?

level: seniorimportance: should knowfreq 55%

basics

~20 s

That is overdispersion. The Poisson likelihood assumes the variance matches the mean, so it understates uncertainty: coefficient estimates stay roughly right, but standard errors come out far too small, intervals too narrow and significance too easy to claim.

open as a page

How does building a survival cohort only from customers who already survived 90 days distort the estimate?

level: seniorimportance: nice to knowfreq 32%

basics

~20 s

The cohort is left-truncated: nobody in it could have cancelled before day 90, yet a signup clock puts them in the early risk sets anyway. Those denominators fill with subjects that cannot fail, so early survival is overstated.

open as a page

In a logistic model, what does a coefficient of 18 with a standard error of 4000 tell you?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

Separation is the usual cause: a predictor splits the outcome perfectly, so the likelihood has no finite maximum and the estimate runs away while its standard error explodes. Look for a zero cell, then use a penalized likelihood such as Firth's.

open as a page

When are excess zeros in a count model a reason to use a zero-inflated model?

level: seniorimportance: nice to knowfreq 32%

basics

~10 s

Only when zeros plausibly come from two processes: units that could never have an event, and units that could but did not. Overdispersion alone often explains a heavy pile of zeros without any mixture.

open as a page

In a Cox model where region violates proportional hazards, when do you stratify rather than model a time-varying effect?

level: principalimportance: nice to knowfreq 32%

basics

~20 s

Stratify when region is a nuisance adjuster you never need an estimate for: each stratum gets its own baseline hazard while the other coefficients stay shared. Model a time-varying effect instead when region's changing effect is itself the finding.

open as a page