Developing a Controlled Vocabulary

Building a controlled vocabulary follows a sequence, even though the steps often overlap in practice. It starts with research: conducting and analyzing user research such as card sorting, learning from search-log analytics to see the actual vocabulary people already use, and analyzing existing content and interviewing content owners to understand the domain. From there, you define a design strategy, what type of controlled vocabulary the situation actually calls for, often captured in a metadata matrix that helps prioritize which vocabularies are worth developing at all.

Only after that groundwork do you gather candidate terms from all these sources into a terminology matrix, decide on preferred terms, identify their variant, broader, narrower, and related terms, and then implement, test, and document the result. Skipping straight to defining preferred terms without the research groundwork is one of the most common ways this process goes wrong, the vocabulary ends up reflecting what the team assumed rather than what users and the domain actually require.

A controlled vocabulary built before the research is really just an organization's internal jargon, formalized and given the appearance of authority it hasn't actually earned yet.

Exercise

The scenario: A national news organization is developing a controlled vocabulary for its digital archive of thirty years of journalism across politics, business, sport, culture, and science. The archive grows daily. Multiple journalists, editors, and archivists will be responsible for tagging content, and the vocabulary needs to support both internal search and public-facing topic navigation.

Which of the following should happen first, before the team starts defining any preferred terms?