In an A/B readout, how does an intent-to-treat analysis differ from a triggered-only analysis?
answer
- as assigned versus as exposed
- one preserves randomisation by construction
- the other needs matched trigger records
- check trigger rates match across arms
- business number versus feature number
basics
~20 sIntent-to-treat counts every assigned user and answers what shipping does to the whole population. A triggered analysis keeps only users who met the trigger condition in both arms, answering what the change does to those who encounter it.
solid answer
~50 sThe two readouts answer different questions on the same experiment. Intent-to-treat analyses every user in the arm the platform assigned them to, whether or not they ever reached the change; it preserves randomisation by construction and gives the business impact of shipping. A triggered analysis restricts both arms to users who met the trigger condition, for example everyone who opened the settings page, and gives the effect on the people the change can actually reach. The triggered number is far more sensitive because it drops users who add only noise, but it is only a randomised comparison when the trigger is recorded identically in both arms and the treatment does not change who meets it. If treatment makes the trigger more likely, the two triggered sets are no longer comparable and intent-to-treat is the honest fallback. Choose which one is primary before launch, and report both.
go deeper
Be ready to state that intent-to-treat keeps every assigned user as assigned, while a triggered read keeps only users who met the exposure condition in both arms.
Explain why intent-to-treat preserves randomisation automatically and what a triggered read must satisfy before it is trustworthy: matched logging and a trigger the treatment cannot move.
Demonstrate the diagnosis: compare trigger rates across arms, and say what you do when they differ rather than quietly reporting the sharper number.
Own the policy that the primary readout and the powering assumption are fixed before launch, so no team gets to choose its estimand after the numbers arrive.
## The same data, two questions One experiment supports more than one readout, and the difference is not statistical hair-splitting: the two numbers answer questions different people are asking. **Intent-to-treat (ITT)** analyses every assigned user in the arm they were assigned to, no exceptions, regardless of whether they ever reached the change. It answers: *if we ship this, what happens to the population?* **Triggered analysis** restricts both arms to the users who met the trigger condition, and compares those two sets. It answers: *for the people who actually encounter this change, what does it do?* ## Why ITT is the safe default Randomisation guarantees that, in expectation, the two assigned groups are alike on everything, observed and unobserved. Analysing exactly those groups keeps that guarantee intact. No decision made after assignment, by the user, the product or the analyst, can break it, because nobody is being included or excluded on the basis of anything that happened after randomisation. That is also why ITT is described as conservative. Because it averages over users the change could not touch, it dilutes the effect toward zero. On a narrowly triggered feature that dilution can be enormous, and a real win reads flat. ## Why the triggered read is worth having Dropping users who cannot be affected removes variance without removing signal. That is a genuine, sometimes order-of-magnitude, sensitivity gain, and it is the number a product team needs to know whether the design works. But it is a comparison you have to earn. Three conditions: 1. **The trigger is recorded on both arms.** Control code must evaluate the same condition at the same point and log that the user qualified, without changing anything the user sees. Without that, the only available comparison is triggered treatment users against all control users, which mixes an exposed group with a general population and is not a valid contrast at all. 2. **The trigger is not affected by the treatment.** If the treatment makes people more likely to reach the settings page, the triggered set in treatment contains people the triggered set in control does not, and the two are no longer exchangeable. The randomisation that made the comparison trustworthy applied to assignment, not to who ends up qualifying afterwards. 3. **The trigger is fixed before launch.** A trigger definition adjusted after seeing outcomes is just another researcher degree of freedom. ## Diagnosing condition 2 The practical check is trigger-rate parity: compute the share of assigned users who triggered in each arm. Under a trigger the treatment cannot influence, those shares should be close, within what sampling noise allows. A material gap is evidence that exposure itself is downstream of the change, and it should be treated as a reason to lean on ITT and to redefine the trigger toward something determined before treatment can act, such as an eligibility check evaluated on entry rather than an interaction deep in the flow. ## Reporting both Mature practice reports both numbers side by side and labels them: - Triggered effect: what the change does to those who see it, with its interval. - ITT effect: what shipping does to the whole population, with its interval. - The trigger rate itself, which is the bridge between them. All three belong in the readout. Quoting the triggered number alone inflates the apparent value of the launch by roughly the reciprocal of the trigger rate. Quoting ITT alone can bury a feature that works well for the people who use it. ## The discipline that matters most Decide, before the experiment starts, which readout the ship decision hangs on and how the experiment is powered. Both readouts are legitimate; picking between them after the results land is not analysis, it is selection. The pre-registered choice is usually the triggered read as the primary decision metric for feature efficacy, with ITT reported as the business impact, and with the trigger-rate parity check standing as the gate that decides whether the triggered read is trustworthy at all.
- When does a triggered analysis stop being a randomised comparison?When the treatment itself changes who meets the trigger condition. If the new design makes users more likely to reach the surface, the triggered set in treatment contains people with no counterpart in control, and the two sets are no longer alike. The practical detector is comparing trigger rates across arms; a material gap means you fall back to intent-to-treat and redefine the trigger to something settled before treatment can act.
- Which readout should the ship decision be pre-registered on?Whichever question the decision actually turns on, chosen before launch. Feature efficacy decisions usually hang on the triggered read, which is also what the experiment should be powered for, with intent-to-treat reported alongside as the business impact. What matters is that the choice is written down first; picking the readout after seeing both numbers is selection dressed up as analysis.
- Is intent-to-treat ever the wrong number to report?It is never invalid, but it can be uninformative on its own. On a feature reaching 3% of users, the intent-to-treat effect is tiny and its interval will usually contain zero, so reporting it alone tells you almost nothing about whether the design works. Report it as the population impact, and pair it with the triggered effect and the trigger rate so the reader can see both.
saying these in an interview costs you the question
- Calls intent-to-treat biased because it includes unexposed users
- Compares triggered treatment users against the whole control arm
- Picks whichever of the two readouts looks better afterwards
- Assumes the triggered effect generalises to everyone
- Never checks whether trigger rates match across the two arms