skip to content

Why are app-store star ratings a biased estimate of average user satisfaction?

level: middleimportance: should knowfreq 42%

answer

  1. nobody drew this sample
  2. who bothers to click rate
  3. the silent middle never posts
  4. the prompt decides who is asked
  5. it estimates a conditional mean

basics

~20 s

Reviewers volunteer. Posting a rating takes effort, and the people willing to spend it are disproportionately delighted or furious, while the indifferent majority stays silent. The published average estimates the mean rating among reviewers, not among users.

solid answer

~40 s

Ratings are a **volunteer sample**: nobody drew the reviewers, they selected themselves. Effort filters the population, and the motivation to spend that effort is correlated with the thing being measured, which is exactly the condition for selection bias. The result is the familiar J-shaped pile-up at one and five stars with a thin middle — strong feelings get posted, mild ones do not. So a 4.6-star average estimates `E[rating | the user chose to review]`, not `E[satisfaction]` across users. A second, sharper problem is the solicitation mechanism: if the app prompts for a rating after a successful action, the frame itself is conditioned on success, and changing that prompt moves the average without anything about the product changing. To measure satisfaction, survey a random sample of users and treat non-response explicitly.

go deeper

for a junior

Be ready to say that reviewers choose themselves, that strong opinions are far more likely to be posted than mild ones, and that the average therefore describes reviewers rather than users.

for a middle

Explain the mechanism precisely: the published average estimates the mean rating conditional on choosing to review, and the J-shaped distribution follows from effort filtering by intensity. Name the prompt as a second selection step.

for a senior

Show the operating instinct: before trusting any user-sentiment number, ask who was invited, what triggered the invitation, and what fraction answered, and refuse release-over-release comparisons when the solicitation rule changed.

for a principal

Take a position on whether public ratings belong in the company's metric set at all, and define what replaces them — a sampled satisfaction survey with a stated frame, alongside behavioural measures that need no volunteers.

## Volunteer samples A **volunteer** or **self-selected** sample is one where membership is decided by the unit, not by the researcher. Public product ratings, open web polls, banner-recruited surveys and call-in votes are all volunteer samples. They are the cheapest data in existence and, for estimating a population mean, close to worthless — because the decision to participate is made by the same person whose opinion you are trying to measure, using that opinion as an input. Formally, let `V = 1` mark a user who leaves a rating and `Y` be the rating they would give. The observed average estimates the conditional mean `E[Y | V = 1]`. This equals the population mean `E[Y]` only if participation is independent of the opinion. For reviews it plainly is not: annoyance and delight both raise the probability of posting, and mild contentment lowers it. ## The J-shaped distribution Public ratings across many platforms show a characteristic shape: a large mass at the top of the scale, a secondary bump at the bottom, and very little in between. Two forces produce it: 1. **Effort filtering by intensity.** Writing a review costs a few minutes. Users pay that cost when they feel strongly. A user whose experience was "fine" has no motivation to spend it. 2. **Purchase self-selection upstream.** People who chose to install an app already expected to like it, which lifts the top of the distribution relative to the population of all potential users. The practical consequence is that the *mean* of a J-shaped, self-selected distribution is unstable and hard to interpret; a handful of one-star posts moves it substantially, and the movement reflects who was angry enough to post rather than a shift in satisfaction. ## The solicitation mechanism is part of the frame The sharper, more interview-relevant issue is that ratings are usually *prompted*. If the in-app prompt fires after a completed purchase, a finished workout or a successful sync, then the sampling frame is not "users" but "users who just had a good moment". That is a designed selection on the outcome. Three consequences follow: - The average is inflated by construction, and the size of the inflation depends on the trigger, not on the product. - **Version-to-version comparisons break** if the prompt logic changed between releases. A rise from 4.3 to 4.6 after a release that also moved the prompt tells you about the prompt. - Users who churned before ever reaching the trigger cannot appear at all — the least satisfied segment has zero probability of being sampled, which is undercoverage in its purest form. ## What to do instead - **Draw the sample yourself.** Prompt a random subset of users, chosen by the analytics system rather than by their own enthusiasm, and prompt them at a time that is not conditioned on a success event. - **Measure and report the response rate**, not just the number of responses. A 3% response rate on a random draw is still a volunteer sample of a random invitation, and should be reported that way. - **Compare respondents to everyone on observables.** You usually know tenure, plan, platform and engagement level for every user, respondent or not. If respondents differ sharply on those, say so; if the differences run through variables you observe, they can be weighted back toward the user base. - **Triangulate with behaviour.** Retention, repeat usage and support-contact rates are measured on the whole population and do not require anyone to volunteer. They answer a different question than satisfaction, but they are not selection-filtered in the same way. - **Use ratings for what they are good for.** The *text* of reviews is a rich source of specific defects and feature requests. Treating reviews as a bug-report channel is sound; treating their arithmetic mean as a satisfaction metric is not. ## Interview framing The strongest answers do three things: state precisely which conditional mean the published number estimates, identify the *two* selection steps (who installs and reviews at all, and who the prompt reaches), and propose a concrete alternative design rather than just criticising the metric. A weak answer stops at "reviews are biased because unhappy people complain more" — half the story, and it misses the prompt mechanism entirely.

  • How would you get an unbiased satisfaction estimate instead?
    Draw a random sample of users from your own user table and invite them, rather than letting users nominate themselves. Avoid triggering the invitation on a success event, record the response rate, and compare respondents with non-respondents on attributes you already hold, such as tenure and engagement. Where those attributes differ, weight the respondents back toward the user base and report that you did.
  • The average rating jumped after a release that also changed when the rating prompt fires. What do you conclude?
    Nothing about satisfaction. Changing the prompt changes who is asked, and a shift in the selection rule is a sufficient explanation for the shift in the average. To separate the two you would need a rating channel whose solicitation was held constant across the releases, or a randomly sampled survey run identically before and after.
  • Do more reviews make the average more trustworthy?
    They make it more precise, not less biased. The extra reviews come from the same self-selecting population, so the estimate converges on the reviewers' mean rather than the users' mean. Volume is a cure for noise and never a cure for a selection rule.

saying these in an interview costs you the question

  • Treats review counts as a random sample of users
  • Says a large number of reviews makes the average representative
  • Assumes the silent majority matches the reviewers
  • Compares versions while ignoring a changed rating prompt
  • Blames only angry users, missing the prompt-based selection

context