Donna Spencer's card-based classification evaluation tests a hierarchy in isolation from the actual user interface. Rather than sorting cards into categories, participants respond to realistic task scenarios, questions like "what US state lost the most acreage to wildfires in 2020?", and the researcher lays out the category cards level by level as the participant says where they'd look next, drilling down through the real hierarchy. Because it tests the structure alone, it complements usability testing, which is better suited to evaluating the interface itself. Sessions are quick, typically 10 to 15 minutes covering 10 to 15 scenarios, with 20 or more participants recommended.
Tree testing is the online, automated version of the same idea, sometimes called tree jacking. A tool presents the hierarchy's top level, the participant clicks down through it to complete a task, and the tool records exactly which path they took. First clicks matter enormously: if a participant's first click matches the intended top-level category, task success runs around 87 percent; if it doesn't, success drops to around 46 percent. Because it's automated and online, tree testing scales easily to large samples, collecting data from 500 participants costs little more than collecting it from 50, and it's well suited to comparing two alternative hierarchies against each other directly.
Both methods ask the same underlying question, can people find their way through this hierarchy, but one puts a researcher in the room to watch it happen, and the other lets you ask it at a scale no room could hold.
Tree testing, run remotely and automatically, is the better fit. A geographically spread user base, a limited budget, and a tight timeline all point away from in-person sessions. Tree testing's ability to scale to large samples at low marginal cost, entirely online, directly solves the geographic and budget constraints that would make card-based classification evaluation harder to run at the scale this project needs.
The trade-off is real, though: tree testing is unmoderated by default, so it captures what path a participant took and whether they succeeded, but not why they chose that path at each decision point. Card-based classification evaluation, by contrast, has a researcher present, able to ask directly why a participant is choosing a particular category as they drill down, capturing exactly the reasoning tree testing's automated clicks can't record on their own.A reasonable complement here is running a small number of moderated card-based classification evaluation sessions alongside the larger tree test, or adding a brief moderated tree-testing option for a handful of participants. Either approach lets the team pair tree testing's scale and efficiency with a smaller amount of the qualitative "why" that only a researcher asking follow-up questions in real time can capture.
Fill in your own details below; the prompt updates as you type. When it's ready, copy it into Claude or whatever AI tool you use.
Casts the AI as: You're a UX researcher analyzing tree testing results for navigation optimization.
Names 5 specific outputs to produce, so the response comes back structured rather than a general summary.
Purpose, Previous research and Known issues tell the AI what your specific situation is, not a generic one.
Comprehensive analysis with actionable recommendations