skip to content

In a usability test where three of five participants completed a music app's offline-download task, how do you report completion rate, time on task and errors honestly?

level: seniorimportance: should knowfreq 38%

answer

  1. counts before percentages
  2. a wide interval from five people
  3. small samples show problems exist
  4. skewed times need a median
  5. define an error before counting

basics

~20 s

Report five-person usability metrics as counts with an interval: 3 of 5 completed, plausibly 23 to 88 percent. Summarise skewed task times with a median or geometric mean, define errors beforehand, and pair each number with its observed cause.

solid answer

~50 s

I would say '3 of 5 completed', not '60 percent', and add that a small-sample (adjusted Wald) 95 percent interval runs from roughly 23 to 88 percent, so the rate is barely pinned down. What five sessions do show is that the failure is not rare: two of five failing is hard to square with a problem only a few listeners hit. For **time on task** I would report successful attempts separately from failures and use a median or geometric mean, because task times are right-skewed, and I would not compare think-aloud timings with a silent benchmark. For **errors** I would use a definition written before the sessions and distinguish slips from mistakes. Then I would lead with why the two failures happened. If the business needs a reliable rate, say for a launch gate, that calls for a larger summative study after the fixes.

go deeper

for a junior

Recall the three core usability metrics, completion rate, time on task and error rate, and that each is defined before the sessions start.

for a middle

Explain why task times are skewed, why failed attempts are reported separately, and what the five-users heuristic assumes about problem frequency.

for a senior

Report small samples honestly with counts and intervals, argue that five sessions prove a problem exists rather than size it, and call for a summative study when a rate matters.

for a principal

Decide when the organisation needs formative rounds versus summative benchmarks, and set reporting norms so small-sample percentages never drive launch decisions alone.

## The metrics and what each answers | Metric | Definition | What it tells you | Small-sample caution | |---|---|---|---| | **Completion rate** | Share of participants who reached the task's predefined success state | Whether people can do the task at all | A yes-or-no outcome from five people has a very wide confidence interval | | **Time on task** | Time from task start to success, or to giving up | Efficiency, and where people hesitate | Right-skewed; think-aloud inflates it | | **Error rate** | Errors per task attempt, or share of participants making at least one | Where the design misleads people | Needs a written definition of what counts as an error | All three depend on decisions made **before** the sessions: the success criterion for each task, when the clock starts and stops, and what counts as an error. Deciding afterwards invites the team to count near-misses as successes. ## Completion rate: say '3 of 5', then give the range Three of five participants set up offline listening; two did not. Reporting '60 percent completion' implies a precision the data does not have. A 95 percent **adjusted Wald** interval, commonly recommended for small-sample completion rates, runs from roughly **23 percent to 88 percent**. The honest statement is: > '3 of 5 participants completed the task. With this sample, the true completion rate among similar listeners could plausibly be anything from about a quarter to nearly nine in ten.' The same data supports a firmer claim about the **failures**: two failures in five give a failure-rate interval of roughly 12 to 77 percent. It is unlikely that almost nobody fails. That is the asymmetry of small samples: they are good at showing that a problem **exists and is not rare**, and poor at pinning down **how big** it is. ## Why five participants, and when that is wrong The widely cited **'five users' heuristic** comes from a problem-discovery model associated with Nielsen and Landauer: if each problem is hit by an average of about 31 percent of users, five participants will see roughly 85 percent of such problems at least once. It is a sensible rule for **formative** testing, meaning iterative rounds meant to find and fix problems, with known limits: - it assumes problems are fairly common; a problem that affects 5 percent of users will usually be missed; - it assumes one fairly homogeneous user group; distinct groups (new listeners versus long-time subscribers, phone versus in-car use) each need their own participants; - it says nothing about estimating metrics; **summative** studies that benchmark a completion rate or compare versions need much larger samples, often dozens per group depending on the precision required; - its value comes from **iteration**: test with five, fix, test with five more. Who those participants are and how they are recruited is a research-planning question in its own right; the point here is what their numbers can and cannot support. ## Time on task: skewed, and only comparable like for like - **Use the median or geometric mean**, not the arithmetic mean. Task times are **right-skewed**: most people finish in similar times and a few take much longer, which drags the arithmetic mean up. For small samples the geometric mean is commonly recommended as the better estimate of the typical time. - **Separate successful and failed attempts.** Averaging the time of someone who gave up after four minutes with the times of people who succeeded mixes two different outcomes. Report time on successful attempts and carry failures in the completion rate. - **Do not compare across protocols.** Think-aloud sessions run slower than silent ones, so an unmoderated silent benchmark and a moderated think-aloud round are not comparable timings. - With five participants, time on task is mostly **descriptive**, for example 'two of the three successes took over three minutes, searching in settings', not a benchmark. ## Error rate: define it before counting An **error** is an action that departs from a path to success: a wrong tap that leads nowhere, offline listening switched on for the wrong playlist, a setting changed by mistake. Write down which actions count, and distinguish: 1. **Slips**: the intention was right but the execution failed, such as a mis-tap on a small control. 2. **Mistakes**: the intention was wrong because the participant misunderstood the design, such as looking for offline listening in the device's own settings. Mistakes usually point at labelling or mental-model problems; slips point at layout and control size. Report error rate as errors per attempt or as the share of participants with at least one error, and say which. ## Reporting honestly - Lead with counts ('3 of 5'), then the interval, then the observation that explains the numbers, for example that both failures looked for a save option on individual songs rather than on the playlist. - Put each metric next to the problem it illustrates; the product decision comes from the problem, not from the percentage. - If the business needs a reliable rate, for a launch gate or a comparison between versions, run a larger summative study, often unmoderated, after fixing what this round found. A headline of '60 percent success' is not wrong arithmetic; it is the wrong claim. The defensible claim is smaller and more useful: a meaningful share of listeners will fail to set up offline listening, and here is why.

  • When would you run a larger usability study instead of another five-person round?
    When the decision needs a number: a launch gate on completion rate, a comparison between two designs, or a benchmark to track over time. Those are summative questions and need samples large enough for useful confidence intervals, often dozens per group. Finding and fixing problems stays with small iterative rounds.
  • Do you include failed attempts when reporting time on task in a usability test?
    Report them separately. Mixing the time of someone who gave up with the times of people who succeeded blends two outcomes. Common practice reports time on successful attempts, carries failures in the completion rate, and notes how long people persisted before giving up when that is informative.
  • How do you tell slips from mistakes when counting usability-test errors?
    A slip is the right intention executed wrongly, like a mis-tap on a small control; a mistake is a wrong intention caused by misunderstanding, like looking for offline listening in the device's settings. Mistakes point at labels and mental models, slips at layout and control size, so they lead to different fixes.

saying these in an interview costs you the question

  • Three of five completing gives a 60 percent completion rate we can publish.
  • Five users are enough to find every usability problem.
  • Time on task should be the arithmetic mean of all attempts, failures included.
  • Results from only five participants tell you nothing at all.
  • Think-aloud timings can be compared directly with an earlier silent benchmark.