After usability testing a music app's playlist-sharing flow, you have 30 findings; how do you rate their severity and decide what to fix first?
answer
- observations are not yet problems
- frequency, impact, persistence
- one sighting can still be critical
- severity apart from fix effort
- re-test what you fixed
basics
~20 sGroup usability observations into underlying problems, then rate each by frequency, impact on the task and persistence on an agreed scale. Keep severity separate from fix effort, fix harmful problems first, and re-test the fixes.
solid answer
~50 sFirst I would **synthesise**: thirty notes usually collapse into a dozen problems once observations with a shared cause are grouped. Each problem gets its evidence (how many participants, on which task, a clip or quote) and a **severity** rating on a scale the team agreed beforehand, judged by **frequency**, **impact** (does it block the task, cause a costly error or only slow people) and **persistence** (a one-time hurdle or a recurring one), plus how critical the task is. A single sighting can still be the top issue: one participant unknowingly making a private playlist public is severe on impact alone. Two or three people rate independently and reconcile. Only then do I weigh **effort**: problems that harm users' data, money or privacy are fixed before release as a rule, high-severity cheap fixes go next, costly ones get an owner and a plan, and the fixes are re-tested in the next round.
go deeper
Recall that findings are grouped into problems and rated by how often they occur, how badly they hurt the task and whether they keep recurring.
Explain a severity scale with agreed definitions, why a single sighting can still be severe, and why raters should work independently before reconciling.
Show how you turn thirty observations into a short, evidence-backed list, keep severity apart from effort, and schedule re-testing of the fixes.
Frame severity as the organisation's shared currency for usability debt: one scale across teams, and a clear policy on what must be fixed before release.
## From observations to findings Thirty notes from a usability round are rarely thirty problems. The first step is **synthesis**: group observations that share an underlying cause. Five participants each hesitating at the share options, two of them asking 'which one copies the link?', is one finding about ambiguous share options, not seven. A useful finding records: - the **problem**, stated as what the design does to people ('the share options do not say whether the recipient can edit the playlist'); - the **evidence**: how many participants were affected, on which task, with a short quote or clip; - the **severity** rating; - a **recommendation**, or at least the direction of a fix. ## Rating severity A widely used approach judges each finding on three factors: 1. **Frequency**: how many participants hit it, and how common the situation is in real use. 2. **Impact**: how hard it is to overcome. Does it block the task, cause a costly error, or only slow people down? 3. **Persistence**: is it a one-time learning hurdle, or will users keep hitting it every time? Many teams add **task criticality**: a problem on the core listening or sharing flow matters more than one on a rarely used setting. The factors combine into a scale. A common one runs from 0 to 4: | Rating | Label | Example from a playlist-sharing flow | |---|---|---| | 4 | Catastrophe | A participant shared a private playlist and unknowingly made it public | | 3 | Major | Three of five could not find how to share with one specific friend and gave up | | 2 | Minor | Two participants hesitated before recognising the copy-link option, then used it | | 1 | Cosmetic | One participant remarked that the share icon looked dated | | 0 | Not a problem | A stated disagreement with the design that had no effect on behaviour | Three-level scales (critical, serious, minor) work as well; what matters is that the team agrees on the definitions **before** rating. ## Where severity judgements go wrong - **Counting only frequency.** A problem seen once can still be the most severe. The privacy exposure above happened to one participant out of five, but its consequence, a private playlist visible to everyone, makes it critical, and one sighting in five says little about how rare it really is. - **Rating by volume of complaint.** Participants who talk a lot generate many notes. Rate behaviour and consequence, not how strongly someone voiced an opinion. - **Letting the designer of a screen rate its findings alone.** Ownership biases ratings in both directions. Have two or three people rate independently, then discuss the differences. - **Mixing severity with fix effort.** 'It's a one-line wording change, so it's minor' confuses how bad the problem is with how cheap the fix is. ## From severity to a fix order Severity says how bad a problem is; **prioritisation** also weighs effort, risk and strategy. Keeping the two steps separate stops an expensive but severe problem from quietly disappearing. 1. **Fix harmful problems before release.** Problems that expose users' data, money or privacy are not backlog items, even when the fix is expensive. 2. **Order the rest by severity against effort.** High-severity, low-effort findings, such as a misleading label in the share options, go first; high-severity, high-effort ones get an owner and a plan rather than a silent deferral. 3. **Batch low-severity items** into a polish list and revisit them when the flow is next changed. 4. **Re-test the fixes.** A fix is a new design; the next round confirms it solved the problem and did not create another. ## Presenting the findings - Lead with the handful of high-severity findings, each with its evidence. A report that lists thirty equal bullets hides the three that matter. - Include **positive findings**, what worked well, so the team does not 'fix' something that is fine. - State the limits: a five-person round shows that problems exist; it does not measure how many listeners they affect. - Link each finding to its task and participant count so readers can judge the evidence themselves. Severity rating is what turns a usability test into a product decision. Without it, the loudest finding wins; with it, the team can defend why the privacy problem ships fixed and the dated icon waits.
- Two raters disagree on a usability finding's severity. What do you do?Ask each to explain the evidence they weighted: frequency, impact, persistence and task criticality. Disagreement usually exposes different assumptions about how often the situation arises in real use or how costly the error is. Agree the definitions, re-rate, and record the reasoning so later readers can see why the rating stands.
- How do you handle a severe usability finding that is expensive to fix?Keep the severity as rated and make the cost visible: give it an owner, a plan and, where possible, an interim mitigation such as clearer wording or a confirmation step. Downgrading it to fit the budget hides risk from the people who decide what ships.
saying these in an interview costs you the question
- A problem seen by only one participant cannot be severe.
- Severity is how many participants complained about something.
- A problem with a cheap fix should be rated as low severity.
- Every finding in a usability report deserves equal weight.
- The designer of a screen should rate its findings alone.