How is a System Usability Scale score calculated in a usability study, and why does a score of 72 not mean '72 percent usable'?
answer
- ten statements, five-point agreement
- odd positive, even negative
- reverse, sum, multiply
- commonly cited average near 68
- perceived usability, not task success
basics
~20 sSystem Usability Scale scoring: odd items give response minus 1, even items 5 minus response; the ten-item sum times 2.5 gives 0 to 100. It is not a percentage; 72 is just above the commonly cited average of about 68.
solid answer
~50 sThe **System Usability Scale (SUS)** is a ten-statement questionnaire given once after a session, each statement rated from 1 (strongly disagree) to 5 (strongly agree). Odd statements are positively worded and contribute the response minus 1; even statements are negatively worded and contribute 5 minus the response. The ten contributions sum to 0 to 40, and multiplying by 2.5 gives a score from 0 to 100. The range invites a percentage reading, but it is a scaled sum of ratings, not a share of users or tasks, and the scores are not centred on 50: the commonly cited average across published studies is about 68. So 72 means 'somewhat above average' perceived usability. I would report it with a confidence interval, compare it with the product's own earlier score, and read it alongside completion rates, because SUS measures perception, not performance, and does not say what to fix.
code
pseudocode · 8 linesfunction susScore(responses): // ten ratings, each 1..5, in questionnaire order
total = 0
for i in 1..10:
if i is odd:
total = total + (responses[i] - 1) // positively worded statement
else:
total = total + (5 - responses[i]) // negatively worded statement, reversed
return total * 2.5 // 0..100, not a percentagego deeper
Recall that SUS is a ten-statement questionnaire giving a 0 to 100 perceived-usability score, and that the score is not a percentage.
Walk through the scoring, odd items minus one, five minus even items, sum times 2.5, and read results against the commonly cited average of about 68.
Show that you use SUS as a trend next to behavioural metrics, report confidence intervals for small samples and keep wording standard so comparisons hold.
Place SUS in a measurement strategy: when a standardised perceived-usability score earns its cost across products, and when task-level metrics answer the business question better.
## What the System Usability Scale is The **System Usability Scale (SUS)** is a standardised questionnaire, created in the mid-1980s, that yields a single score for **perceived usability**. Participants answer it once, usually right after the tasks of a usability session and before any open discussion, about the product as a whole. It is popular because it is short, free to use, technology-agnostic and backed by decades of published comparison data. It has **ten statements**, each rated on a **five-point agreement scale** from 1 (strongly disagree) to 5 (strongly agree). The statements **alternate in tone**: - **odd items** (1, 3, 5, 7, 9) are **positively** worded, for example that the system was easy to use; - **even items** (2, 4, 6, 8, 10) are **negatively** worded, for example that the system was unnecessarily complex. The alternation discourages participants from ticking the same column all the way down, and it is also the main source of scoring mistakes. ## Scoring it 1. For each **odd** item, subtract 1 from the response, so 1 to 5 becomes 0 to 4. 2. For each **even** item, subtract the response from 5, so 5 to 1 becomes 0 to 4. This reverses the negative items, so a higher contribution always means better perceived usability. 3. Add the ten contributions. The total runs from 0 to 40. 4. Multiply by **2.5**. The score now runs from **0 to 100**. 5. For a study, average the individual scores and report a confidence interval, because samples are usually small. A participant who answers 4, 2, 4, 1, 5, 2, 4, 2, 4, 3 contributes 3 + 3 + 4 + 3 + 3 = 16 from the odd items and 3 + 4 + 3 + 3 + 2 = 15 from the even items, 31 in total, which becomes **77.5**. ## Why 72 is not '72 percent usable' The 0 to 100 range invites a percentage reading, but nothing in the scale measures a proportion of users, tasks or anything else. - **It is not a percentage of anything.** No participant is '72 percent satisfied'; the score is a scaled sum of ten ratings. - **The distribution is not centred on 50.** Across large collections of published studies, the **commonly cited average is about 68**. A score of 50 is well below average, not middling. - **Interpretation is relative.** A 72 reads as 'somewhat above the commonly cited average'. Practitioners translate scores into percentile ranks or grade-like bands derived from benchmark data, or, better still, compare against the product's own previous score. | Score (common interpretation) | Meaning | |---|---| | Below about 50 | Clearly below average; serious perceived usability problems are likely | | Around 68 | Roughly average across published studies | | Around 80 and above | Among the better-rated products in benchmark data | These band edges are conventions drawn from benchmark data, not thresholds defined by the questionnaire itself. ## What SUS can and cannot tell you - **It measures perception, not performance.** A participant can fail tasks and still rate the product generously, especially after a friendly moderated session. Report SUS alongside completion rate and observed problems, never instead of them. - **It does not diagnose.** A low score says the product feels hard to use; it does not say which screen or flow. The sessions and severity-rated findings answer that. - **It is read as one total.** Treating a single statement's rating as a diagnosis of a specific feature over-interprets one data point from each participant. - **Small samples wobble.** With eight participants the confidence interval around the mean can easily span ten or more points, so a move from 70 to 74 between two small rounds may be noise. ## Administering it well - Give it **after the tasks and before the debrief**, so the facilitator's discussion does not colour the ratings. - Keep the wording standard. Small, widely used substitutions, such as 'app' or 'website' in place of 'system', are common practice; rewriting statements breaks comparability with benchmark data. - Do not make every item positive without changing the scoring key to match. A mismatched key silently corrupts every score. - For a per-task signal, use a single post-task ease rating instead; SUS is a whole-product measure. In a music-streaming app study, SUS is most useful as a trend line: the team scores the listening experience before a navigation redesign and again after it, with the same wording and a similar participant profile, and reads the change alongside the task-level evidence of what got easier and what did not.
- How would you compare System Usability Scale scores before and after a redesign?Keep the wording, the participant profile and the session format the same, then compare the means with confidence intervals rather than as bare point scores. With small samples a few points of change can be noise, so pair the SUS trend with task-level evidence such as completion rates and the problems observed.
- Why do the System Usability Scale statements alternate between positive and negative wording?Alternation discourages participants from ticking the same column without reading and makes careless responding easier to spot. The cost is scoring errors: forgetting to reverse the even items is the most common mistake, and it silently corrupts every score in the study.
saying these in an interview costs you the question
- A System Usability Scale score of 72 means 72 percent of users are satisfied.
- A System Usability Scale score of 50 is an average result.
- You can simply add up the raw System Usability Scale responses.
- A System Usability Scale score tells you which screen has the problem.
- A high System Usability Scale score proves participants completed their tasks.
- Rewording the statements freely does not affect comparison with benchmarks.