skip to content

In information architecture, what does tree testing measure that card sorting cannot, and how do you read its results?

level: middleimportance: should knowfreq 40%

answer

  1. card sorting in reverse
  2. text-only hierarchy, no visuals
  3. find-it tasks, clicked top down
  4. success, directness, first click
  5. baseline tree versus revision

basics

~20 s

Tree testing checks whether people can find items in a text-only hierarchy by clicking down through labels to complete find-it tasks. It measures findability — success, directness and first clicks — which the grouping exercise of card sorting never tests.

solid answer

~50 s

Card sorting shows how people **group** items; tree testing shows whether they can **find** them in a proposed hierarchy. Participants see only a text version of the tree — no visual design, no search, no navigation widgets — and get tasks such as “You booked tickets online and want to print them here. Where would you go?” They click down level by level and choose where they expect the answer. The main measures are **success rate** (did they end at a correct node), **directness** (did they get there without backtracking), **first click** (did they start down the right branch, which is closely associated with eventual success) and **time**. Reading results means studying **paths**, not only scores: which wrong label attracted people and where they turned back. Because it strips out visual design, a tree test isolates structure and labels, so a baseline tree can be compared with a revision before anything is built.

go deeper

for a junior

Recall that tree testing checks whether people can find items in a text-only hierarchy, while card sorting checks how they group items.

for a middle

Explain success, directness and first click, why the tree is stripped of visual design, and how to write tasks that do not leak the label.

for a senior

Show how you read paths and competing labels, compare a baseline against a revision task by task, and decide what to rename or cross-list.

for a principal

Decide where cheap structural validation belongs in a product's delivery process, and what evidence threshold a restructure must meet before release.

## Grouping versus finding **Card sorting** asks people to group content with everything visible at once. Real navigation works the other way round: a person starts at the top, sees only the labels at that level, and must predict which one leads to what they want. **Tree testing** — sometimes described as card sorting in reverse — measures exactly that. It answers the question card sorting cannot: **can people find things in this hierarchy, using these labels, starting from the top?** ## How a tree test works 1. **Build the tree** as plain text: the top-level labels, their children, and so on, down to the items — no visual design, no icons, no search. 2. **Write tasks** that describe a need without using the target's label, for example “Your annual pass has expired and you want to extend it today.” 3. **Define correct answers** for each task — sometimes more than one node is acceptable. 4. **Run it**, usually unmoderated, with each participant attempting a set of tasks and clicking down (and back up) until they choose a destination. 5. **Analyse** per task, then across the tree. Because the test ignores layout, colour and interaction, it isolates the two things IA controls: **structure** and **labels**. ## The measures | Measure | What it tells you | Kiosk reading | |---|---|---| | Success rate | Share of participants ending at a correct node | Low for “print online booking” means that item is effectively hidden | | Directness | Share who reached the answer without backtracking | Low directness with high success: people get there only after wrong turns | | First click | Which top-level label people chose first | Many first clicks on “Tickets” for a member task: that label is a magnet | | Time | How long the task took | Long times flag hesitation between similar labels | | Paths | The full route each person took | Shows exactly where the structure misleads | **First click** is watched closely because people who start down the right branch usually succeed, while a wrong first click often ends in failure or a long detour. Top-level labels are therefore the most expensive to get wrong. ## Reading the results - **Per task, not only overall**: an average success rate can hide one task that fails badly. - **Follow the wrong paths**: the label that wrongly attracted people tells you what to rename or where to cross-list an item. - **Separate direct and indirect success**: indirect success means the structure works as a fallback but its labels mislead; on a walk-up kiosk, people who backtrack often give up in real life. - **Look for competing labels**: two siblings drawing similar shares of first clicks are not distinct enough. - **Compare trees**: run the same tasks on the current tree as a **baseline** and on the proposed one, and check which tasks improved and which got worse. ## Writing tasks that don't leak the answer - Describe the situation and goal, not the destination. - Avoid the exact words of the target label; a task containing “Membership renewal” just tests word-matching. - Cover the most important and most troubled items rather than every leaf of the tree. ## Limits - **No search, no visual cues**: real interfaces add both, so a tree test is a harsher condition than real use and shows the structure with nothing propping it up. - **No content**: people choose where they *expect* something to be, not whether the page there satisfies them. - **Tasks are artificial**: they are given, not self-motivated, so they show findability more than motivation. A useful rhythm on a restructure is: open card sort to generate a structure, closed sort to check placement, tree test to check findability, then a first-click test on designed screens once visual design exists, so each method answers the question it is good at and no more. Within those limits a tree test is one of the cheapest ways to validate a structure before building it, and it works identically whether the hierarchy will become kiosk screens, website pages or the views of a native mobile app.

  • Why is the first click watched so closely in a tree test?
    People who start down the correct branch usually succeed, while those who start wrong often fail or take long detours. First clicks show which top-level labels fail to signal their contents — the labels whose mistakes cost the most, because every task passes through them.
  • What does high success with low directness tell you?
    People eventually find the item, but only after wrong turns. The structure works as a fallback while the labels mislead. On a walk-up kiosk, where people have little patience and someone is often waiting behind them, indirect success in a test tends to become abandonment in real use.
  • How do you use a tree test to compare two structures?
    Run the same tasks against the current tree as a baseline and against the proposed one, with comparable participants. Compare success, directness and first clicks task by task, and look for tasks that got worse as well as the average improvement, since a restructure often fixes some items by hiding others.

saying these in an interview costs you the question

  • A tree test should use the full visual design so it feels realistic.
  • A high overall success rate means every part of the tree works.
  • Tasks can reuse the target label's exact wording to avoid confusion.
  • Tree testing and card sorting answer the same question.
  • Only the final destination matters; paths and backtracking are noise.