skip to content

Why can clicks-per-user as an experiment's OEC reward clickbait content?

level: seniorimportance: should knowfreq 52%

answer

  1. recorded before value is known
  2. promise versus delivery
  3. measure becomes a target
  4. optimisation finds the cheap path
  5. count only clicks that paid off

basics

~20 s

Clicks measure attention captured, not value delivered. A sensational headline earns the click before the user can tell whether the content is worth it, so a variant that misleads scores as a win on the criterion while making the product worse.

solid answer

~50 s

A click happens *before* the user knows whether the content was any good, so clicks-per-user records the promise, not the delivery. Any variant that makes the promise louder — overstated headlines, withheld information, misleading thumbnails — raises the criterion without improving anything the user experienced. The deeper point is Goodhart's law: a proxy that was a decent passive indicator of satisfaction stops being one once teams start optimising against it, because optimisation searches exactly the space of ways to move the number without moving the underlying thing. The fix belongs in the criterion itself. Make the metric require evidence that the click paid off — a click that survives some minimum engagement, a satisfaction-adjusted or dissatisfaction-discounted version, or an outcome measured after the click rather than at it. Then validate the revised criterion against longer-horizon behaviour: gaming it should now require actually helping the user.

go deeper

for a junior

Know that a click is recorded before the user can judge the content, so counting clicks alone can reward a misleading headline.

for a middle

Explain the mechanism: the proxy correlates with satisfaction only while nobody optimises against it, and optimisation searches exactly the ways to move the number without moving the thing.

for a senior

Demonstrate the fix and its price. Describe redefining the criterion so a click must pay off, what that costs in sensitivity, and how you would confirm the new criterion survives optimisation pressure.

for a principal

Be ready to argue when the org should accept a gameable but sensitive criterion under supervision versus mandate a costlier one, and how incentives around shipped-win counts make proxy erosion an organisational problem, not just a metric bug.

## Where the gap opens A click is recorded at the moment a user decides a thing looks worth opening. Whether it *was* worth opening is determined afterwards — by whether they read it, finished it, came back, or bounced away irritated. Clicks-per-user therefore measures the strength of the promise a piece of content makes, not the value it delivers. For a passive observer this distinction may not matter much: across a large sample, better content probably does get clicked more, so clicks correlate with satisfaction. The correlation is what makes the metric tempting. It is also what makes it dangerous, because a correlation observed under normal conditions is not a guarantee that survives deliberate optimisation. ## Goodhart's law, stated precisely Goodhart's law: when a measure becomes a target, it ceases to be a good measure. The mechanism is not mysterious. There are two families of ways to raise clicks — make the content genuinely better, and make the promise more aggressive than the content justifies. Under passive observation, only the first is being exercised much, so the correlation with satisfaction holds. Once the metric is the criterion that decides launches, teams and ranking models search *both* families, and the second is usually cheaper. The population of changes being evaluated shifts toward exactly the ones that break the correlation. This is why the objection is not answered by *but we checked, clicks correlate with satisfaction*. The correlation was measured on the old distribution of content. The criterion changes that distribution. ## The concrete failure In a ranking or headline experiment where clicks-per-user is the declared criterion, the winning variant may be one that promotes sensational or withholding headlines. The criterion registers a clean, statistically significant win. The users, meanwhile, click, discover the content does not match the promise, and leave — and if the criterion is read at the click, none of that shows up. Worse, the failure often shows up *positively* in the short run. A user who is misled once may click again on the next equally sensational item before learning to distrust the surface, so early readings of the metric can look even better than the true effect. The damage accrues over a longer horizon than the experiment observes. ## Fixing the criterion The response is to change what the criterion counts, so that the cheap way to move it stops working. **Require the click to pay off.** Instead of counting all clicks, count only those where the user did something afterwards that indicates the content delivered — stayed past a threshold, scrolled meaningfully, completed the item, took a subsequent action. This does not eliminate gaming, but it raises the price: to move the metric you now have to hold attention, not just capture it. **Penalise the failure mode explicitly.** A criterion can subtract or down-weight interactions with the signature of dissatisfaction, such as an immediate return to the previous surface. This makes a misleading promise actively costly on the criterion rather than merely uncounted. **Measure after the click, not at it.** The strongest version moves the criterion downstream entirely, to an outcome that only occurs when the user was served well. The cost is that downstream outcomes are rarer and slower, which reduces sensitivity — which is the standard tension, and it has to be paid consciously. **Validate the revised criterion against a longer horizon.** After redefining, check that treatment effects on the new criterion agree with treatment effects on longer-run behaviour across past experiments. If the revised metric still moves for changes that later look bad, the redefinition did not work. A separate line of defence — declaring auxiliary metrics whose degradation can veto a launch — belongs to a different part of the metric system and is not covered here. What *is* covered here is that the criterion carrying the decision should not be one whose cheapest path upward is deception. ## Judgment an interviewer wants to hear The strong answer distinguishes three things: that clicks are recorded before value is known (the mechanical gap), that optimisation pressure predictably exploits that gap (Goodhart), and that the remedy is a criterion redefinition with a stated cost in sensitivity. A candidate who only says *use a better metric* has identified the problem without owning the tradeoff — the better metric is invariably rarer and noisier, and the team has to decide whether it can still resolve the effects it needs to see. One more piece of judgment worth showing: the gaming is often not intentional. Nobody has to plan to mislead users. If the criterion rewards aggressive promises, then honest optimisation — a ranking model trained toward it, a team iterating on headlines — will find that direction on its own. Assuming bad faith is required for Goodhart to bite is a misunderstanding of the mechanism.

  • The team says clicks correlate strongly with user satisfaction in historical logs. Why is that not a defence?
    The correlation was measured while nobody was optimising against clicks. Once clicks decide launches, the changes being evaluated shift toward those that raise clicks most cheaply, and the cheapest route is an inflated promise rather than better content. The historical relationship describes the old distribution of content, not the one the criterion will produce.
  • You redefine the criterion to count only clicks with meaningful post-click engagement. What did that cost?
    Sensitivity, mostly. Qualified clicks are rarer than raw clicks and carry an extra thresholding decision, so the metric is noisier per user and resolves smaller effects less well — the experiment may need longer or more traffic. The threshold itself is also a design choice that can be tuned toward a desired result, so it should be fixed in advance and justified.
  • Does an OEC only get gamed when someone is acting in bad faith?
    No, and assuming so is the common mistake. A ranking model optimised toward the criterion, or a team honestly iterating to move it, will find the cheap path without anyone intending to mislead. Goodhart's law describes what optimisation pressure does to a proxy, not what dishonest people do. Design the criterion so the cheapest way up is genuinely helping the user.

It is like grading a restaurant on how many people walk in the door. A big photo of a dish nobody actually serves fills the room once, and the score never records the people who left after the first bite.

saying these in an interview costs you the question

  • Claiming historical correlation makes a proxy safe to optimise
  • Assuming gaming requires deliberate bad intent
  • Treating clicks as a direct measure of satisfaction
  • Proposing a better metric with no account of lost sensitivity
  • Believing a short experiment would surface the trust damage

context