What does a subscription service give up by training on a usage-decline proxy instead of the cancellation outcome?
answer
- the outcome is late and already decided
- a proxy must be defined as precisely
- quiet is not the same as leaving
- product change moves the target
- keep the real outcome to measure the gap
basics
~20 sIt gives up the guarantee that the target is the business outcome. The proxy matures in weeks and fires while an account is still winnable, but it flags seasonal dips, misses steady users who leave over price, and moves whenever the product changes.
solid answer
~50 sTwo things make the real outcome awkward as a target: it takes about six weeks to mature, and by the time an account executes a cancellation the decision is usually already made, so a perfect prediction arrives too late to act on. A proxy such as `usage in the trailing week below 20% of the account's trailing 8-week median, two weeks running` matures in a fortnight, has a far higher base rate, and exists for accounts that never leave - all of which make it trainable and actionable. The cost is that you are now optimising something adjacent to the thing you care about. The proxy catches holidays as well as disengagement, misses low-usage accounts that cancel purely on price, and shifts when the product changes what usage means. The real outcome must still be recorded so the gap between proxy and outcome stays measurable.
go deeper
Recall that a proxy target is a stand-in for the real outcome, chosen because it arrives sooner, and that it is not the same thing as the outcome.
Explain the trade concretely: faster maturity and a denser positive class, paid for with seasonal false positives and blindness to steady users who leave on price.
Show that you keep the real outcome recorded alongside the proxy, and that you read the gap between them in both directions on a schedule.
Own the strategic risk: a proxy adopted for convenience and never revisited quietly becomes the goal, and the business ends up optimising engagement while revenue leaks.
## Why anyone would model something other than the outcome The cancellation outcome is the thing the business cares about, so choosing anything else needs a reason. Two of them are usually good enough. - **Maturity.** With a 30-day horizon and a 14-day grace period the outcome takes 44 days to become trustworthy, so training always reasons about a six-week-old world. - **Actionability.** The cancellation is the *execution* of a decision, not the decision. An account that reaches the cancellation flow has typically already chosen; a retention offer at that moment fights an argument that is over. A target that fires while the account is still ambivalent is worth more even if it is less exactly right. There is also a volume argument: at a low single-digit monthly rate, positives are scarce, and a target with a higher base rate produces a denser table. ## What the proxy looks like, concretely A usage proxy has to be as precisely specified as the real label, or it is just a feeling. For example: an account is positive if its usage in the trailing 7 days falls below 20% of its own trailing 8-week median, for two consecutive weeks. That definition fixes the unit it is measured on, the comparison baseline (the account's own history rather than a global threshold), and the persistence requirement that stops a single quiet week firing. | Property | Cancellation outcome | Usage-decline proxy | |---|---|---| | Maturity | about 44 days | about 14 days | | Base rate per account-month | low single digits | substantially higher | | Exists for accounts that stay | no | yes | | Means what the business means | exactly | approximately | | Stability over time | anchored to billing | shifts when the product changes | | Actionable while it fires | rarely | usually | ## What you give up 1. **Both error directions get worse in ways the outcome does not.** Accounts go quiet for holidays, a seasonal workload or a role change and come back; the proxy calls them positive. Other accounts use the product steadily right up to renewal and then leave on price; the proxy never fires for them, and those are often the most valuable to save. 2. **The proxy is partly under your own control.** Because it is defined on product telemetry, a redesign that changes what counts as usage moves the target itself. The real outcome is anchored to billing events and does not drift when a screen is reorganised. 3. **The gap is invisible unless you keep the real label.** A proxy target only stays honest if the cancellation outcome is still recorded and the relationship between the two is re-examined - what share of proxy positives actually leave, and what share of leavers ever fired the proxy. Dropping the real label because the proxy is quicker removes the only evidence the proxy still stands for anything. 4. **Acting on the target changes what it predicts.** Once retention intervenes on proxy positives, the treated accounts stop leaving, and the proxy-to-outcome relationship measured before the intervention no longer describes the world the model now lives in. The mechanics of that feedback belong to the retraining design, but the exposure is created here, when the target is chosen. ## Making the choice defensible The answer in a design round is rarely 'proxy instead of outcome'. It is usually: model the proxy where the speed genuinely buys an intervention, keep the cancellation outcome as the recorded truth, state the gap between them as a quantity someone owns, and say under what condition you would switch back. A target chosen for convenience and never revisited is how a team ends up two years later optimising engagement while revenue leaks. ## Where this goes wrong - Adopting the proxy because it is easier to compute, without ever writing down what it is standing in for. - Leaving the proxy threshold as a round number nobody has revisited since the first notebook. - Dropping the real outcome from the warehouse once the proxy is in production. - Treating proxy positives as churners in business reporting, so the reported rate no longer matches billing. - Assuming a proxy is automatically cleaner than the real label, when it carries its own noise from seasonality and product change.
- How would you check that a usage-decline proxy still stands for the real outcome?Keep labelling the cancellation outcome in parallel, on matured rows only, and read the relationship in both directions: of accounts that fired the proxy, what share ended within the horizon, and of accounts that ended, what share ever fired it. Re-read those two shares on a schedule, and specifically after any product change that alters what usage means, because the proxy definition silently moves with it.
- Does a higher base rate make the proxy the easier target to work with?Denser positives help, but the proxy brings its own noise: seasonal quiet periods, accounts with genuinely intermittent workloads, and threshold effects around the cut-off. It is a different problem, not a strictly easier one, and the extra positives are only useful to the extent they correspond to accounts actually at risk.
A clinic that watches for missed follow-up appointments rather than waiting for a discharge letter learns sooner and can still act, but it also flags everyone who was simply on holiday.
saying these in an interview costs you the question
- A proxy target is cleaner because it has no billing edge cases
- Once the proxy works, the real cancellation label can be dropped
- Low usage means an account is leaving
- The proxy threshold is a modelling detail, not a target definition
- The proxy-to-outcome relationship holds even after acting on the proxy