skip to content

In a usability test, what is wrong with the task 'Tap Download to save your Road Trip playlist offline', and how would you rewrite it?

level: middleimportance: must knowfreq 55%

answer

  1. the task hands over the answer
  2. on-screen words inside the task
  3. goal and situation, not path
  4. success criterion written in advance
  5. post-task questions can lead too

basics

~20 s

The usability task names the on-screen label and the action, so participants match words instead of discovering the path. Rewrite it as a realistic goal with no interface vocabulary, such as making the playlist playable on a flight without internet.

solid answer

~40 s

The task 'Tap Download to save your Road Trip playlist offline' gives away the answer: 'Download' is the button's own label and 'tap' names the interaction, so participants scan for a matching word and the test stops measuring whether real listeners would find the feature. It also gives no reason, so participants cannot judge for themselves when they are done. I would rewrite it as a situation and a goal in everyday words: 'You're flying tomorrow and the plane has no internet. Make sure you can listen to your Road Trip playlist during the flight.' Before the sessions I would write down the success criterion, the playlist marked as available without a connection, and I would pilot the task and keep post-task questions neutral, because 'That was easy, wasn't it?' leads just as badly.

go deeper

for a junior

Recall that a usability task states a goal and a reason in everyday words, never the button, menu or label the participant is supposed to use.

for a middle

Explain why a label in the task inflates completion, how to write a goal-based rewrite with a predefined success criterion, and how post-task questions can lead as well.

for a senior

Show that you audit tasks against the screens, pilot them, counterbalance their order and would refuse to act on a confident result produced by a leading script.

for a principal

Treat task wording as research quality control across teams: shared task review and piloting standards stop confidently wrong evidence from reaching product decisions.

## Why task wording decides the result In a **usability test**, a **task scenario** is the instruction a participant tries to carry out, such as finding a song or setting up offline listening. Whatever the task tells the participant, they no longer have to work out from the interface. A task that names the path measures reading, not usability, and it produces the most dangerous kind of result: a confident, near-perfect completion rate for a flow that real users cannot find. Take the task 'Tap Download to save your Road Trip playlist offline'. It has three problems: 1. **It names the on-screen label.** 'Download' is the exact word on the button, so the participant can scan for a matching word instead of working out where the capability lives. Real listeners may think of it as 'saving', 'keeping' or 'making it work on the plane'. 2. **It names the interaction.** 'Tap' plus the label turns a discovery problem into an instruction to follow. The question the team needed answered, 'would people find this?', is answered by the task itself. 3. **It gives no motivation.** Nothing tells the participant why they would do this, so they cannot bring their own judgement about when the job is finished. ## Rewriting it A better version describes a **goal** and a **situation**, in the participant's words, and leaves the path to them: > 'You're flying tomorrow and the plane has no internet. Make sure you'll be able to listen to your Road Trip playlist during the flight.' This version: - uses no interface vocabulary ('download', 'offline' and 'tap' are all gone); - gives a realistic reason, so the participant can decide for themselves when they are done; - has one goal, so success and failure are unambiguous; - still has an **observable success criterion**, written down before the sessions: the playlist is marked as available without a connection. Writing it down in advance stops the team deciding afterwards that a near-miss 'basically counts'. ## Rules for unbiased task scenarios | Rule | Leading version | Neutral version | |---|---|---| | Describe goals, not steps | 'Open Library, then tap the plus icon' | 'Put together some songs for your morning run' | | Avoid on-screen labels | 'Find the Sleep Timer' | 'You fall asleep listening; make the music stop by itself after a while' | | One goal per task | 'Find a podcast, follow it and share an episode' | Three separate tasks, each with its own success state | | Give a realistic reason | 'Make a playlist' | 'Your friend is hosting a party and asked you for songs' | | Do not signal difficulty | 'This one might be tricky, but…' | Present every task the same way | | Observe, do not poll | 'Would you use collaborative playlists?' | Watch whether they complete a collaborative-playlist task | There is a trade-off to name: a scenario must still be specific enough that the participant knows what 'done' means. 'Explore the app' is unbiased but unmeasurable. The skill is being specific about the **goal** while saying nothing about the **path**. Common domain words that users share with the interface, such as 'playlist' in a music app, can stay when the task is not about finding that label. ## Bias beyond the task text - **Post-task questions.** 'That was easy, wasn't it?' or 'Did you like the new download button?' lead as badly as a leading task. Use a balanced rating, such as a single post-task rating from very difficult to very easy, which is common practice, and open follow-ups like 'Tell me about how that went'. - **Task order.** Early tasks teach the interface for later ones. If two tasks share a path, rotate or counterbalance their order across participants. - **Moderator reactions.** Saying 'great' after a correct tap tells the participant they are on track. Stay neutral in both directions. - **Realistic content.** Where possible, let participants work with their own taste in music or a realistic seeded account; a library full of unfamiliar songs adds confusion that is not the design's fault. ## Checking your tasks before the study 1. Read each task and underline every word that appears on the screens being tested. Replace each one or justify it. 2. Write the success criterion next to each task. 3. Pilot the tasks with one or two people outside the team and watch where they ask 'what do you mean?'. 4. For an unmoderated study, where nobody can clarify anything, pilot until no one needs to ask. A task that gets this right yields evidence the team can act on. A leading task yields a confident, useless number, and the team ships the very problem the test was meant to find.

  • How specific can a usability task be before it starts leading?
    Specific about the goal and the situation, silent about the path. Naming the playlist, the deadline and the reason is fine; naming the tab, the button or the menu is not. A practical check is to underline every word in the task that appears on the tested screens and either replace it or justify it.
  • When is it acceptable to use an interface word in a usability task?
    When the word is also the user's own vocabulary, such as 'playlist' in a music app, and the study is not asking whether people find or understand that label. If the question is whether listeners recognise a label, the task must avoid it; otherwise a shared domain word can stay.
  • How do you stop earlier usability tasks from teaching the interface for later ones?
    Order tasks the way real use would flow, and where two tasks share a path, rotate or counterbalance their order across participants so learning does not always favour the same task. Note likely learning effects in the report rather than treating later successes as independent evidence.

A leading usability task is like an exam question that contains its own answer: everyone scores full marks, and the score tells you nothing about what the students actually know.

saying these in an interview costs you the question

  • Naming the button in a usability task keeps sessions short and does no harm.
  • A 100 percent completion rate proves the feature is easy to find.
  • Asking participants whether they would use a feature is a usability task.
  • Success criteria can be decided after watching the sessions.
  • Unbiased tasks have to be vague, like 'explore the app'.