skip to content

How would you design an encouragement experiment when you cannot randomize the treatment itself?

level: principalimportance: nice to knowfreq 30%

answer

  1. randomize the invitation, not the treatment
  2. uptake is voluntary but assignment is not
  3. precision hinges on how much uptake moves
  4. a persuasive nudge is a dangerous nudge
  5. decide which population the decision needs

basics

~20 s

Randomize the invitation rather than the treatment, leave uptake voluntary, then use the nudge as an instrument for actual use. The nudge must move uptake substantially and must not affect the outcome by any other route.

solid answer

~50 s

Randomly assign an encouragement, an invitation, a reminder, a temporary discount, rather than the treatment itself, and let people choose. Randomization buys independence. The design then rests on two judgment calls. **Strength**: the nudge must shift uptake enough to serve as a denominator, because precision degrades roughly with the square of the compliance rate, so a nudge that moves uptake by 5 points needs far more sample than one that moves it 40. **Exclusion**: the nudge must not change the outcome by itself. This is where encouragement designs usually die, because a persuasive email that explains the benefits can move behaviour, mood, or attention independently of the treatment. Keep the encouragement content-thin and channel-neutral. Then decide up front what the resulting complier effect will authorise: it answers what happens if you keep nudging, not what happens if you ship to everyone by default.

go deeper

for a junior

Know the core trick: you randomize who gets invited rather than who gets the treatment, and adoption stays the person's own choice.

for a middle

Be able to explain why the invitation works as an instrument, which two assumptions randomization does and does not buy you, and why uptake strength drives the precision of the result.

for a senior

Demonstrate operational judgment: pilot the first stage before committing, keep the nudge content thin to protect exclusion, and watch for contamination or follow-up processes that create second pathways.

for a principal

Own the estimand-versus-decision alignment before anyone builds anything, and be willing to tell leadership that this design cannot forecast a default-on rollout no matter how clean the execution.

## When the design applies Sometimes you cannot force the treatment on anyone. Users choose whether to adopt a feature, patients choose whether to take a drug, teams choose whether to use a tool. Forcing it may be technically impossible, contractually forbidden, or ethically unacceptable. What you often can randomize is a **nudge**: who gets invited, reminded, offered a trial, shown a prompt. Randomizing the nudge and using it as an instrument for actual uptake is the encouragement design. The randomization delivers the independence assumption outright, which is the assumption that is hardest to earn in observational instrumental-variables work. What you still have to earn is the other two. ## Judgment call one: how hard to push The estimate divides the nudge's effect on the outcome by the nudge's effect on uptake. That denominator is the compliance rate, and it sets your precision. Because the standard error of the rescaled effect is roughly the standard error of the nudge effect divided by the compliance rate, sample requirements grow with the inverse square of compliance. Moving uptake from 5 percent to 9 percent gives a 4-point denominator; moving it from 10 to 50 gives a 40-point denominator, and the second design is roughly a hundred times more efficient on the same population. So the design work is mostly uptake engineering. Pick a nudge channel people actually see. Offer something with real pull. Consider stacking the nudge, several reminders rather than one, if that is acceptable. Run a small pilot whose only purpose is to measure the first stage before committing the full sample, because a design with a 4-point first stage is not worth running. ## Judgment call two: keep the nudge inert The exclusion restriction says the nudge may touch the outcome only by causing uptake. Every persuasive element you add to raise uptake also threatens this. An email that explains why the feature improves productivity may make people more productive by suggestion or by directing attention, whether or not they adopt. A discount changes the money available for other things. A reminder re-engages a dormant user with the product generally. The design discipline is to make the encouragement as thin as possible while still moving uptake: an announcement rather than an argument, a neutral pointer rather than a benefits pitch. These two goals fight each other, which is the essential tension of the design. Say out loud in the plan which side you erred on and why. Also ensure the nudge is not observed by the control group, and that no operational process treats nudged and unnudged units differently for other reasons. A support team that follows up on invited users has created a second pathway. ## What the design will and will not tell you Two quantities come out. The effect of the nudge itself needs no rescaling and no extra assumptions beyond randomization; if your actual decision is whether to run the campaign, this is the number you want and the instrumental machinery is optional. The rescaled effect of uptake requires exclusion and monotonicity, and applies only to the people the nudge moved. That second population is the crux of the planning conversation. If the plan is to ship the treatment to everybody by default, the never-takers are in scope and this design says nothing about them. If the plan is to keep nudging, the compliers are exactly the people future campaigns will move, and the local effect is the operationally correct number. Settle this before running, not after seeing the result. ## Pre-registration checklist 1. Target first stage, with a pilot estimate, and a stated minimum below which the analysis will not be rescaled at all. 2. Exact nudge content, fixed in advance, with an argument for why it cannot affect the outcome directly. 3. The outcome window, chosen so it starts after uptake could plausibly occur. 4. The estimand the decision needs, and an explicit statement of whether the compliers match that population. 5. A plan for reporting the encouragement effect regardless of what the rescaled estimate does. ## Failure modes worth naming A nudge that moves nobody produces an unusable denominator, and no analysis can rescue it. A nudge so persuasive that it changes behaviour on its own breaks exclusion, and the resulting number is not a treatment effect. Contamination, where uninvited users learn about the feature from invited colleagues, breaks the comparison in the other direction by shrinking the measured first stage while possibly moving control outcomes. The best encouragement designs are boring nudges with large uptake effects, and they are rarer than they sound.

  • Your pilot shows the nudge moves uptake by only 4 percentage points. What do you do?
    Do not run the full design as planned. Precision scales roughly with the inverse square of compliance, so a 4-point first stage would need an implausible sample. Redesign the nudge for reach and pull, or accept reporting only the encouragement effect itself, which needs no rescaling. Running anyway and rescaling by a tiny denominator produces a number that is both biased and dishonestly precise.
  • How does making the encouragement more persuasive trade against the design's validity?
    Persuasion raises uptake, which improves precision, but every persuasive element is a candidate direct pathway to the outcome and so threatens the exclusion restriction. A benefits pitch can change behaviour among people who never adopt at all. The resolution is to buy uptake through reach, salience and timing rather than through argument, and to state the tradeoff explicitly in the plan.
  • Leadership wants to know the effect of shipping the feature to everyone. Does this design answer that?
    Not directly. It estimates the effect among people the nudge moved, which excludes those who would never adopt voluntarily and those who already had. Extrapolating to a default-on rollout requires assuming effects do not vary with adoption propensity, which is often the least plausible assumption available. Say so before running, and if the decision truly is about a default rollout, argue for a rollout experiment instead.

saying these in an interview costs you the question

  • Reports the rescaled estimate without checking the first stage
  • Writes a persuasive nudge without considering direct effects
  • Presents the complier effect as a default-rollout forecast
  • Ignores contamination between nudged and unnudged units
  • Skips a pilot measurement of how much uptake moves

context