skip to content

In a trial/paid/churned chain with churn absorbing, how do you compute the expected months until churn?

level: seniorimportance: should knowfreq 52%

answer

  1. strip out the absorbing rows first
  2. one equation per transient state
  3. expected visits add up to expected time
  4. invert I minus Q, sum the row

basics

~20 s

Take the submatrix Q of transitions among the non-absorbing states and solve (I - Q) t = 1, where 1 is a vector of ones. The trial entry of t is the expected number of months before the user churns.

solid answer

~50 s

Write the chain in canonical form: transient states first, absorbing states last, so `P` splits into `Q` (transient to transient), `R` (transient to absorbing) and an identity block. With monthly rates `trial -> trial 0.5, -> paid 0.3, -> churned 0.2` and `paid -> paid 0.9, -> churned 0.1`, the transient block is `Q = [[0.5, 0.3], [0, 0.9]]`. The **fundamental matrix** `N = (I - Q)^-1 = [[2, 6], [0, 10]]` has entries `N[i][j]` = expected months spent in state `j` before absorption, starting from `i`. Its row sums are the expected times to absorption: 8 months starting from trial, 10 months starting from paid. The trial row also tells you the split — about 2 months in trial and 6 months paid. `I - Q` is invertible exactly because the chain leaves the transient states with probability 1, so `Q^n` decays to zero.

go deeper

for a junior

Be ready to recognise an absorbing state as one with a self-transition probability of one, and to say that expected time to absorption comes from the transient block of the matrix.

for a middle

Set up the canonical form, write (I - Q) t = 1, and do the two-by-two inverse by hand, explaining what each entry of the fundamental matrix counts.

for a senior

Connect the arithmetic to the business question: expected months paid versus months alive, conversion probability from N times R, and whether the flat-hazard assumption survives contact with the data.

for a principal

Own the state-space design decision — how finely to split tenure, whether churn should be absorbing at all when users resubscribe, and how much model complexity the available data can support.

## Absorbing chains A state is **absorbing** when `P[i][i] = 1`: once the chain enters it, it never leaves. A chain is an **absorbing chain** when it has at least one absorbing state and every other state can reach one. Non-absorbing states are called **transient**, because the chain leaves each of them permanently with probability 1. Order the states with transient ones first and absorbing ones last. Then the transition matrix splits into blocks: ``` P = [ Q R ] [ 0 I ] ``` `Q` holds transient-to-transient probabilities, `R` holds transient-to-absorbing probabilities, and the bottom rows say absorbing states stay put. Note that rows of `Q` sum to less than 1 for at least some states — the missing mass is the chance of being absorbed this step. ## The fundamental matrix Raising `P` to the `n`th power keeps the block structure with `Q^n` in the corner, and `Q^n -> 0` because absorption is certain. That makes the geometric series converge: ``` N = I + Q + Q^2 + Q^3 + ... = (I - Q)^-1 ``` `N` is the **fundamental matrix**, and `N[i][j]` is the expected number of visits to transient state `j` before absorption, starting from `i` (counting the starting visit). Summing across a row counts every transient step the chain takes, which is exactly the **expected time to absorption**: ``` t = N * 1 equivalently (I - Q) t = 1 ``` The second form is often easier by hand: it is one linear equation per transient state, and it encodes 'one step now, plus whatever the state you land in still costs'. ## Worked lifecycle example Monthly states `trial`, `paid`, `churned`, with `churned` absorbing: - `trial`: stay 0.5, `paid` 0.3, `churned` 0.2 - `paid`: stay 0.9, `churned` 0.1 Transient block and its complement: ``` Q = [ 0.5 0.3 ] I - Q = [ 0.5 -0.3 ] [ 0.0 0.9 ] [ 0.0 0.1 ] ``` The determinant is `0.5 * 0.1 = 0.05`, so ``` N = (I - Q)^-1 = [ 2 6 ] [ 0 10 ] ``` Reading it off: starting in `trial`, the user spends on average 2 months in `trial` and 6 months in `paid`, for **8 months total** before churning. Starting in `paid`, 10 months. The single-state check is reassuring: from `paid`, absorption happens with probability 0.1 each month, and the mean of that geometric count is `1 / 0.1 = 10`. It is worth noticing that 8 is less than 10. A trial user is not worse off than a paid user in some paradoxical sense — most trial users churn without ever converting, which drags the average lifetime down. ## Absorption probabilities With two or more absorbing states, the matrix `B = N R` gives `B[i][k]` = probability of eventually being absorbed in state `k` starting from `i`. To ask 'does a trial user ever convert?', make `paid` absorbing as well and recompute. Now the only transient state is `trial`, with `Q = [0.5]` and `R = [0.3, 0.2]`, so `N = 1 / 0.5 = 2` and ``` B = N R = [0.6, 0.4] ``` A trial user converts with probability 0.6 and churns without converting with probability 0.4. Intuition check: conditional on leaving `trial` in a given month, the split is `0.3 : 0.2`, i.e. 60/40 — the stay probability cancels out. ## Assumptions this model imposes A constant per-month transition probability out of a state means the time spent there is geometric, so the hazard of churning is flat: a user in month 2 of `paid` is exactly as likely to churn as a user in month 30. Real subscriptions almost never behave that way — churn spikes at onboarding and again at renewal. Two honest repairs: split the state by tenure (`paid-new`, `paid-renewed`, `paid-established`) so the hazard can vary across states, or move to a model that lets the sojourn time follow an arbitrary distribution. Also check whether `churned` is truly absorbing; if a meaningful share of users resubscribe, an absorbing state overstates how final churn is, and the answer becomes a long-run share question rather than an absorption question. ## Failure modes If some transient state cannot reach any absorbing state, absorption is not certain, `Q^n` does not decay, and `I - Q` is singular — a numerical blow-up that is really a modelling error. And keep the two outputs straight: row sums of `N` are **times** in months, while entries of `B` are **probabilities**. Reporting one as the other is the most common way this calculation goes wrong in a review.

  • How do you get the probability that a trial user ever converts rather than churning first?
    Make `paid` absorbing too and compute `B = N R`. With `trial -> trial 0.5, -> paid 0.3, -> churned 0.2`, the only transient state is `trial`, so `N = 1 / 0.5 = 2` and `B = [0.6, 0.4]`: a 60% chance of converting. The intuition is that conditional on leaving trial, the split is 0.3 to 0.2.
  • What does a single entry of the fundamental matrix mean on its own?
    `N[i][j]` is the expected number of months spent in transient state `j` before absorption, starting from state `i`, counting the initial visit. Row sums give total expected time to absorption. That per-state breakdown is often the more useful number — expected months paid, not just expected months alive.
  • What churn assumption does this model impose, and when does it break?
    A fixed monthly exit probability makes the time in a state geometric, so the churn hazard is flat over tenure. Real subscriptions spike at onboarding and at renewal. Repair by splitting the state into tenure buckets so each carries its own hazard, or by moving to a model with arbitrary sojourn-time distributions.
  • What happens if a transient state cannot reach any absorbing state?
    Absorption is no longer certain, `Q^n` does not decay to zero, and `I - Q` is singular, so the inverse does not exist. Numerically you see a blow-up or a huge condition number. The real fix is structural: either that group of states is genuinely closed, or transitions are missing from the estimated matrix.

saying these in an interview costs you the question

  • Inverts the full transition matrix instead of I minus Q
  • Reports absorption probabilities as if they were expected times
  • Forgets that the current step counts toward the total
  • Assumes a flat churn hazard without checking tenure effects
  • Simulates thousands of users when a two-by-two inverse is exact

context