Testing Findability: Card-Based Classification Evaluation and Tree Testing

Donna Spencer's card-based classification evaluation tests a hierarchy in isolation from the actual user interface. Rather than sorting cards into categories, participants respond to realistic task scenarios, questions like "what US state lost the most acreage to wildfires in 2020?", and the researcher lays out the category cards level by level as the participant says where they'd look next, drilling down through the real hierarchy. Because it tests the structure alone, it complements usability testing, which is better suited to evaluating the interface itself. Sessions are quick, typically 10 to 15 minutes covering 10 to 15 scenarios, with 20 or more participants recommended.

Tree testing is the online, automated version of the same idea, sometimes called tree jacking. A tool presents the hierarchy's top level, the participant clicks down through it to complete a task, and the tool records exactly which path they took. First clicks matter enormously: if a participant's first click matches the intended top-level category, task success runs around 87 percent; if it doesn't, success drops to around 46 percent. Because it's automated and online, tree testing scales easily to large samples, collecting data from 500 participants costs little more than collecting it from 50, and it's well suited to comparing two alternative hierarchies against each other directly.

Both methods ask the same underlying question, can people find their way through this hierarchy, but one puts a researcher in the room to watch it happen, and the other lets you ask it at a scale no room could hold.

Exercise

The scenario: A government agency has redesigned the information hierarchy for its public-facing website and needs to evaluate findability before launch. Their users are spread across a wide geographic area, the budget doesn't support extensive in-person sessions, and the timeline is tight, but the team specifically wants to understand not just whether users succeed, but why they make the navigation choices they do.

Given the geographic spread, budget, and timeline constraints, which method should the team choose as their primary evaluation approach?
Try It With Your Data: Tree Test Analysis

Fill in your own details below; the prompt updates as you type. When it's ready, copy it into Claude or whatever AI tool you use.

What this prompt is made of (RICE breakdown)
R

Role

Casts the AI as: You're a UX researcher analyzing tree testing results for navigation optimization.

I

Instructions

Names 5 specific outputs to produce, so the response comes back structured rather than a general summary.

C

Context

Purpose, Previous research and Known issues tell the AI what your specific situation is, not a generic one.

E

Expected format

Comprehensive analysis with actionable recommendations

Paste or describe the tree you tested
Whole number, e.g. 8
Whole number, e.g. 40
One task per line, success rate, directness, common paths, first clicks
Why you ran this test.
What research led to this tree
Problems you already suspect
Assembled prompt