Labels should be tested with real, representative users, not just during initial design, but again after launch, since a labeling system that made sense on paper can still fail once real people encounter it in context. One simple technique is asking participants to describe what they expect to find at a link's destination before they click it, then comparing that expectation to what's actually there.
Two methods covered earlier in this course are particularly effective for label testing specifically. Card sorts, open, modified-Delphi, or closed, reveal how people understand structure, categories, and the words that should describe them. Free-listing surfaces a domain's vocabulary and boundaries directly from participants, without the researcher's own assumptions shaping what gets tested in the first place.
A label that made perfect sense in a design review can still fail its first real test, the only way to know for certain is asking the people who were never in that room.
1: Free-listing. With no labels proposed yet and a need to understand citizens' own vocabulary from scratch, free-listing is built for exactly this moment, surfacing a domain's language and boundaries directly from participants before any structure is imposed on them.
2: Closed card sort. Testing whether citizens would correctly group content under an already-proposed set of labels is precisely what a closed sort evaluates, it uses predefined categories and checks whether real users' groupings match the team's intent.
3: Search-log analysis on the live site. Understanding whether real, live users are turning to search over browsing requires actual usage data from the running site, logs of what people are typing when navigation alone hasn't gotten them where they need to go.
A closed card sort can confirm whether citizens group content the way the team intended, but it can't tell the team whether the label wording itself is the best possible choice, a citizen might sort content correctly under a label while still finding that label mildly confusing, non-native, or not quite how they'd phrase it themselves, simply because the closed sort constrains them to the categories already provided rather than asking what they'd call something in their own words. An open card sort, or circling back to free-listing with the same participants, would be better suited to answering that separate question about word choice itself, since both let participants generate language rather than just react to language the team already chose.