What counter-metrics would you track for a test that increases ad density on a page?
answer
- assume it won for the worst reason
- revenue up, users quietly worse off
- unsubscribes, ad-blockers, return visits
- slow signals need holdbacks
- rare events are badly powered
basics
~10 sTrack what a revenue win costs users: unsubscribe and churn rate, ad-blocker installs, return-visit frequency, complaints, and page latency. Short-run ad revenue almost always rises, so revenue alone settles nothing.
solid answer
~50 sA counter-metric is the metric you expect to move badly if the treatment wins by extracting value rather than creating it, and more ads is the textbook case: short-run revenue almost always goes up, so revenue alone cannot tell you whether the change was good. I would track unsubscribe or churn rate, ad-blocker install rate, return-visit frequency, session length or pages per session, complaint and support-ticket volume, and page latency, since more ad slots usually mean more requests and a slower page. Two design points matter more than the list. First, several of these are slow — churn and ad-blocker adoption barely move in a two-week window — so the experiment needs enough run time or a holdback to see them at all. Second, each needs a threshold agreed before launch, or the revenue win will simply be declared to outweigh whatever was observed.
go deeper
Be ready to say what a counter-metric is and name two for an ads change — unsubscribes and return visits — and to explain why revenue alone cannot settle the question.
Explain the reasoning that generates the list: assume the treatment won by extracting value, then name what was extracted. Cover latency as a technical counter-metric and the per-user versus per-session pitfall.
Demonstrate that you handle detection lag and low power — holdbacks, leading indicators, and stating whether the interval actually rules out a harmful amount rather than reporting non-significance.
Own the exchange rate: how much short-run revenue the organisation will trade for how much user-experience degradation, who decides when a counter-metric breaches, and how that policy stops slow erosion across many launches.
### Why counter-metrics exist A counter-metric is a guardrail chosen specifically because it is the *predictable casualty* of the treatment succeeding. The logic runs backwards from a cynical story: assume the treatment wins on its target metric for the worst possible reason. What would you see? Whatever that is, measure it. Raising ad density is the canonical example because the short-run mechanics are almost guaranteed. More ad slots means more impressions, and more impressions means more revenue this week nearly regardless of quality. The cost is paid by users, arrives later, and lands on metrics nobody looks at if they are not on the scorecard. An experiment that reports only revenue on this change is not measuring the tradeoff; it is measuring one half of it. ### The list, and the reason behind each entry **Unsubscribe and churn rate.** The most direct statement of "users left because of this". For a subscription product this is the sharpest counter-metric available. Its weakness is timing: cancellations often require a renewal date to arrive, so the signal lags the change by weeks. **Ad-blocker install or activation rate.** A specific, high-signal behavioural response — a user going out of their way to defeat the change. It is rare enough to be noisy, but a movement here is hard to explain away. **Return-visit frequency and sessions per user.** Users rarely announce dissatisfaction; they just come back less often. This is usually the earliest broad signal of a degraded experience, and it compounds, which is why a small movement matters more than its size suggests. **Engagement depth: pages per session, scroll depth, session duration.** Distinguishes "the page got worse" from "users abandoned the product". If depth falls while visits hold, the page experience degraded; if visits fall too, the damage is broader. **Complaints, negative feedback and support-ticket volume.** Low-volume and lagging, but a movement is very hard to dismiss in a launch review, which gives it disproportionate weight in the decision. **Page latency and layout stability.** More ad slots mean more network requests and more late-arriving content. This is a technical cost of the same change and belongs on the scorecard alongside the behavioural ones. **Ad quality signals: click-through rate per slot, and revenue per impression.** If total revenue rises while revenue per impression falls sharply, you are diluting inventory rather than growing it, and the trend continues in the wrong direction as density rises further. ### The measurement problems you must raise The hardest issue is **detection lag**. Churn, ad-blocker adoption and habit erosion are slow, low-base-rate behaviours. In a two-week test they can be genuinely present and still statistically invisible. Two responses are standard: run the experiment longer, or keep a **long-run holdback** — a fraction of users who never receive the change — so that the accumulated effect can be read months later against a clean baseline. The second issue is **sensitivity**. Rare events have high relative variance, so a counter-metric may be far less powered than the revenue metric it is supposed to check. This makes a naive reading dangerous: revenue crosses its bar, churn shows "no significant change", and the change ships on an asymmetry of statistical power rather than on evidence of safety. The fix is to state a degradation threshold that matters and check whether the interval on the counter-metric is actually tight enough to rule it out — if it is not, the honest report is "we could not tell", not "no harm". The third is **selection**. Measuring only sessions from users who visited during the test window silently drops the users the change drove away; a per-user metric over the whole assigned population avoids that, while a per-session metric can hide it. ### Turning it into a decision Counter-metrics are only useful if the tradeoff is written down before launch. A workable form is an explicit exchange rate: this much revenue gain is acceptable for at most this much degradation in retention or complaints, with any breach going to a named decision owner rather than to whoever argues hardest in the review. The same structure applies to other extraction-shaped changes — an aggressive upsell flow needs refund rate and billing-support volume beside its conversion win, and a notification-volume increase needs opt-out rate beside its click win. The pattern generalises: whenever a change can succeed by taking rather than creating, name the thing being taken and put a number on it in advance.
- Churn is far too slow to move in a two-week test. How do you still protect against it?Keep a long-run holdback: a slice of users who never get the change, read months later against the treated population. Pair it with faster leading indicators — return-visit frequency, session depth, complaint volume — that move within the test window and correlate with eventual churn. The holdback measures the real cost; the leading indicators let you react before it lands.
- How would counter-metrics differ for an aggressive upsell flow instead of an ad-density change?The extraction is monetary rather than attentional, so the casualties change: refund and chargeback rate, billing support-ticket volume, subscription downgrade rate, and completion of the flow being interrupted. Upsell conversion rising while refunds and support tickets rise with it means the flow is pressuring users into purchases they reverse, which is a net loss once handling cost is counted.
- The revenue metric is significantly up and the churn counter-metric shows no significant change. Is that a green light?Not on its own. Churn is a rare, slow event, so its confidence interval is usually wide enough to contain a commercially serious degradation. The question to answer is whether the interval excludes the amount of churn that would matter — not whether the test returned a non-significant result, which mostly reflects lower power on that metric.
saying these in an interview costs you the question
- Reports only revenue and calls the ad test a win
- Picks counter-metrics after seeing the results
- Reads a non-significant churn result as proof of no harm
- Ignores that churn cannot move within the test window
- Measures engagement per session, hiding users who left