Is the mean of 1-5 Likert satisfaction responses a legitimate summary?
answer
- what kind of scale the options form
- are the gaps between options equal
- relabel the options, keep the order
- report the distribution, not one number
basics
~20 sStrictly no: a 1-5 Likert scale is ordinal, so the gaps between adjacent options are not known to be equal and the mean has no defensible unit. Report it only beside the full response distribution.
solid answer
~50 sA 1-5 Likert item is an ordinal scale: the numbers are codes for ordered categories, and nothing guarantees that the step from `strongly disagree` to `disagree` is the same size as the step from `neutral` to `agree`. Taking a mean silently asserts equal spacing, which is an empirical claim about how people read the wording, not a formality. The sharpest test is to recode the same responses as 1, 2, 3, 4, 10 — the data and their order are identical, but the mean moves a long way, so the mean is a property of your coding, not of the responses. What the ordinal scale does license is counts, proportions, order comparisons and the median. I would report the full frequency table, plus a top-two-box proportion as the headline, and treat any mean I quote as a rough tracking number with the distribution shown next to it.
go deeper
Be ready to say that Likert responses are ordered categories rather than measured quantities, and that the numbers 1 to 5 are codes. Knowing that proportions and counts are always safe will carry the answer.
Explain the equal-spacing assumption that the mean smuggles in, and show the recoding argument that separates statistics belonging to the data from statistics belonging to the coding.
Show the reporting judgment: what actually goes on the survey dashboard, why a top-two-box proportion beats a mean as a headline, and how you would spot a polarised distribution hiding behind a stable average.
Own the measurement standard for the organisation. Decide whether survey KPIs are defined as proportions or as scale means, what disclosure accompanies them, and how you keep those definitions stable so quarter-over-quarter comparisons remain meaningful.
## The setup A Likert item offers a small set of ordered response options — for example `strongly dissatisfied`, `dissatisfied`, `neutral`, `satisfied`, `strongly satisfied` — which are stored as 1 through 5 for convenience. The question is whether that storage convenience licenses arithmetic. ## Why it is ordinal The respondent chose a *category*. The categories are ordered, so we know 4 sits above 3, which sits above 2. What we do not know is how far above. The psychological distance between `neutral` and `satisfied` — the step where a person crosses from indifference into approval — may be much larger, or much smaller, than the distance between `satisfied` and `strongly satisfied`. Respondents also differ from each other in how they use the endpoints; some never pick 1 or 5 at all. None of that is captured by the integers. An interval scale requires equal, meaningful gaps. A Likert item gives you order and nothing more, so by the rules of measurement scales the mean is not defined on it. ## The decisive demonstration Because an ordinal scale is defined only up to an order-preserving relabelling, any legitimate statistic must be unchanged when you relabel while keeping the order. Try it. Take ten responses: four 5s, three 4s, two 3s and one 1. - Coded 1, 2, 3, 4, 5, the mean is `(4*5 + 3*4 + 2*3 + 1*1) / 10 = 39 / 10 = 3.9`. - Recode the options as 1, 2, 3, 4, 10 — still strictly increasing, so the ordinal content is identical. The mean becomes `(4*10 + 3*4 + 2*3 + 1*1) / 10 = 59 / 10 = 5.9`. The responses did not change. The mean did. Now check a statistic the scale actually supports: the median response is the fourth option in both codings, and the proportion answering in the top two options is 0.7 in both codings. Those are properties of the data; the mean is a property of your spreadsheet. ## Why people compute it anyway The practice is widespread, and the arguments for it are not empty: - It compresses a distribution into one trackable number, which is what a weekly dashboard wants. - Movement in the mean does usually reflect movement in the distribution, so as a *trend* signal it is often serviceable. - Multi-item scales are a stronger case: when several related items are averaged into one composite intended to measure a single underlying construct, the composite takes many more distinct values and behaves much more like an interval quantity than any single item does. Analysing composites as interval data is standard practice in survey research and is far more defensible than averaging one item. The honest position is therefore not "never compute it" but "know what you are assuming, and never let it travel alone". ## What the mean hides One number cannot distinguish a population that is uniformly lukewarm from one that is sharply split. Five respondents all answering 3, and a group where half answer 1 and half answer 5, can both produce a mean of 3 while describing completely different situations — and the second one is the situation you needed to act on. Reporting the mean without the response counts throws away exactly the structure that makes survey data worth collecting. ## What to report instead - **A frequency and proportion table** over all five options: each option with its count and its share. This is fully licensed by an ordinal scale and preserves the whole distribution. - **Top-two-box and bottom-two-box proportions**: the share choosing the two most positive options, and the share choosing the two most negative. These are proportions of clearly-defined groups, they survive any recoding, and they are directly interpretable by non-technical readers. - **The median response**, which needs only order. - **The sample size**, always, since survey subgroups get small fast. ## A ratio trap to avoid Even if you accept the mean, do not divide two of them. "Our score rose from 3.5 to 4.2, a 20 percent improvement" imports a true zero that the scale does not have — the coding could have started at 0 or at 10 with the same responses, and the percent change would differ. Report the movement in the units of the scale, or better, report the movement in the top-two-box proportion, which is a genuine proportion and supports ratio statements. ## How to answer this in an interview State the scale, state the assumption that averaging makes, give the recoding demonstration in one sentence, then say what you would actually put in the report. Interviewers are looking for a candidate who neither parrots "you can never average ordinal data" nor shrugs and averages it without noticing — the credible answer is that you know the assumption, you disclose it, and you always ship the distribution beside it.
- What would you report instead, given the scale is ordinal?A frequency and proportion table across all five options, a top-two-box and bottom-two-box proportion as the headline figures, the median response, and the sample size. Those are all licensed by order alone, they survive any recoding of the options, and they preserve the shape of the distribution rather than collapsing it.
- Why does the arbitrary numeric coding of the options matter so much?Because an ordinal scale carries no more than order, any order-preserving recoding describes the same data. Recode 1, 2, 3, 4, 5 as 1, 2, 3, 4, 10 and the mean jumps while the median and the top-two-box proportion do not. A statistic that moves under a legitimate recoding is measuring your coding, not the responses.
- When is averaging Likert responses more defensible?When several items written to measure one construct are combined into a composite score. The composite takes many more distinct values, individual item quirks partly cancel, and it is conventionally analysed as interval data. State the assumption of equal spacing explicitly, and still publish the item-level distributions.
- Two products both average 3.0 out of 5. What have you not learned?The shape of either distribution. One product may have almost everyone answering 3, the other may be split between 1s and 5s — the same mean, entirely different problems, and only the polarised one has an angry segment to address. Compare the full response tables, not the single numbers.
Averaging Likert codes is like averaging the finishing positions in a race: the ranks are real, but the arithmetic quietly assumes every gap between places was the same length.
saying these in an interview costs you the question
- Treats Likert codes as real quantities with a unit
- Reports a mean with no response distribution beside it
- Assumes the step to strongly agree equals every other step
- Says 4.2 out of 5 is 20 percent better than 3.5
- Insists ordinal data can never be summarised at all