How do you estimate an overall mean when power users were deliberately oversampled?
answer
- the sample is not the population
- inverse of the selection probability
- weighted mean, not plain average
- unequal weights cost effective sample size
basics
~20 sWeight each respondent by the inverse of their selection probability and report the weighted mean, sum(w_i * y_i) / sum(w_i). Oversampled power users get small weights and ordinary users large ones, so the estimate reflects the real population mix.
solid answer
~50 sIf power users were sampled at 1 in 10 and everyone else at 1 in 1,000, the design weight is the inverse inclusion probability: 10 for a power user, 1,000 for everyone else. The unweighted average answers a question nobody asked — the mean of a fictional population that is mostly power users — so you report the weighted mean `sum(w_i * y_i) / sum(w_i)`, which is unbiased for the population mean under the design. The price is variance: unequal weights inflate it, and Kish's rule gives an effective sample size of `(sum of w)^2 / sum of w^2`, which can fall far below the number of responses when a few weights dominate. In practice you trim extreme weights, trading a little bias for a large variance reduction, and you make sure no downstream consumer ever averages that table unweighted.
go deeper
Know that oversampling a group makes the raw average over-represent it, and that the fix is a weighted average rather than dropping responses.
Be able to compute a design weight as one over the inclusion probability and apply the weighted-mean formula to a simple two-group example.
Show that you check the effective sample size, watch for a few dominant weights, and decide whether to trim them before making any precision claim.
Own the decision to oversample at all: it buys a reportable segment but obliges every downstream consumer to honour weights forever, so weigh that against a simpler self-weighting design.
## Why the raw average is wrong Deliberate oversampling is a legitimate design choice. Power users are 1% of the base but drive most of the revenue, and you want to say something reportable about them, so you sample them at 1 in 10 while sampling everyone else at 1 in 1,000. The survey comes back with, say, 1,000 power users and 1,000 ordinary users. Averaging those 2,000 responses gives the mean of a population that is 50% power users. No such population exists. The number is not noisy or imprecise — it is an accurate estimate of the wrong quantity, and it will stay wrong as the sample grows. ## Design weights The repair is built into the design. Each unit `i` had a known probability `pi_i` of being selected, and its **design weight** is `w_i = 1 / pi_i`. A power user selected with probability 1/10 carries weight 10 — it represents itself and nine unsampled peers. An ordinary user selected with probability 1/1,000 carries weight 1,000. Weights have a direct interpretation: they are the number of population units each respondent stands for, and they should sum to approximately the population size. That sum is a useful sanity check — if your weights add up to something far from the known base, the inclusion probabilities you used are wrong. For a population **total**, the weighted estimator is `sum of w_i * y_i`. For a population **mean**, divide by the sum of the weights: `ybar_w = sum(w_i * y_i) / sum(w_i)`. Dividing by the weight total rather than by a fixed population count is the standard choice when the realised sample composition varies; it is what makes the estimator behave well in practice. Note what a weighted mean is *not*: it is not a correction for who declined to answer, and it does not compensate for units missing from the frame. Design weights repair exactly one thing — the unequal selection probabilities you yourself chose. ## Unequal weights cost precision A design weight repairs bias but spends variance. Intuitively, when a small number of respondents each speak for a thousand people, the estimate hinges on those few responses, and one unusual answer moves the number. Kish's approximation makes this concrete. The **effective sample size** of a weighted sample is `n_eff = (sum of w_i)^2 / (sum of w_i^2)` It equals the actual sample size when all weights are equal, and drops as the weights spread out. With 2,000 responses split as above, the weight distribution is extremely lopsided — half the sample carries weight 10 and half carries 1,000 — and the effective sample size is a small fraction of 2,000. Any precision statement should be based on `n_eff`, never on the raw count of responses and certainly never on the *sum* of the weights, which is a population estimate rather than a sample size. This is the tension inherent in oversampling: you bought a reportable power-user segment, and you paid for it in the precision of every population-level number. ## Trimming and capping When a handful of weights dominate, practitioners **trim** them: cap weights at some threshold, or at a multiple of the median weight, and optionally redistribute the trimmed mass so the total still matches the population. Trimming introduces a small bias toward the sample's own composition while cutting variance substantially, and it often lowers mean squared error overall. It is a judgment call, so it must be documented: the trimming rule changes what the published number means. ## Weighted for what, exactly A useful discipline is to state which population each number describes. - **Population mean across all users**: weights required. - **Mean within the power-user segment alone**: no weights needed *if* every unit inside that segment had the same selection probability. Within a stratum the design is self-weighting, so a plain average is correct. This is the point of oversampling in the first place — it makes the segment estimate solid. - **Any cross-tab that mixes segments**: weights required, and the effective sample size for each cell should be checked before publishing a number for it. ## Operational reality The most common failure is not statistical but organisational. The weights are computed once, live in a column, and then somebody exports the table into a dashboard and takes a plain average. Practical defences: never publish a raw response table without the weight column attached and documented; where possible, publish pre-aggregated weighted figures rather than the row-level data; and record the inclusion probabilities at collection time, since weights that have to be reconstructed months later usually cannot be. Interviewers ask this because it separates candidates who treat a dataset as a rectangle of equally important rows from candidates who ask how each row got there. If you cannot say what probability produced a row, you cannot say what population your average describes.
- Why would you ever cap or trim large sampling weights?Because a handful of huge weights lets a few respondents dominate the estimate, driving variance up and making the result fragile. Trimming caps the largest weights at a threshold, which biases the estimate slightly toward the sample's own composition but often lowers mean squared error substantially. The rule must be documented, since it changes what the published number means.
- What does the effective sample size tell you about a weighted survey?It converts weighted responses into the number of equally weighted responses carrying the same information, approximated by the squared sum of the weights divided by the sum of squared weights. With heavy oversampling of a small group, 2,000 responses might carry the information of several hundred, and every precision claim should rest on that figure rather than on 2,000.
- Do you need the weights to report on the power-user segment by itself?Not if every unit inside that segment had the same selection probability — within a stratum the design is self-weighting, so a plain average of those respondents is already correct. Weights matter when you combine segments into a population figure. Publish segment estimates and population estimates as clearly separate numbers.
If you invited one in ten of your enterprise customers and one in a thousand of your free users, each enterprise respondent speaks for ten people and each free respondent for a thousand. The overall average has to count them that way.
saying these in an interview costs you the question
- Reports the raw average from a deliberately oversampled survey
- Thinks oversampling by itself makes the overall estimate more accurate
- Treats the sum of the weights as the effective sample size
- Invents weights without knowing the selection probabilities
- Trims extreme weights without documenting the rule