Defining Optimal Labels for Categories

Once candidate terms have been gathered from multiple sources, the real work begins: reconciling them into one coherent vocabulary. This means collecting every candidate into a single list, resolving inconsistencies in spelling, punctuation, and case, and, where several sources offer different words for the same underlying concept, deciding which one becomes the preferred term, with the others recorded as variants that authors and indexers should avoid using going forward.

This process should also surface gaps: does the resulting vocabulary actually cover every important concept the content needs, or are there topics with no proper term at all yet? At the same time, restraint matters, adding a term for every edge case bloats the system and causes the whole vocabulary to lose focus. Terminology isn't static either; as usage evolves, categories that go unused should be reconsidered, either folded into something more relevant or relabeled to match how people are talking about the topic now.

The hardest part of this work usually isn't finding candidate terms: it's what to do when the term your own organization prefers and the term your actual users search for turn out to be two different things.

Exercise

The scenario: A UX team has completed label discovery research for a financial services information space, gathering candidate terms from four sources. The raw list includes: from user interviews, "retirement savings account"; from competitive analysis, "IRA"; from search-log analytics, "individual retirement account" and a common misspelling, "retirment account"; and from subject-matter experts inside the company, the internal product name "Future Fund Vehicle."

Given this mix of candidate terms, which one should become the preferred term?