Building a controlled vocabulary follows a sequence, even though the steps often overlap in practice. It starts with research: conducting and analyzing user research such as card sorting, learning from search-log analytics to see the actual vocabulary people already use, and analyzing existing content and interviewing content owners to understand the domain. From there, you define a design strategy, what type of controlled vocabulary the situation actually calls for, often captured in a metadata matrix that helps prioritize which vocabularies are worth developing at all.
Only after that groundwork do you gather candidate terms from all these sources into a terminology matrix, decide on preferred terms, identify their variant, broader, narrower, and related terms, and then implement, test, and document the result. Skipping straight to defining preferred terms without the research groundwork is one of the most common ways this process goes wrong, the vocabulary ends up reflecting what the team assumed rather than what users and the domain actually require.
A controlled vocabulary built before the research is really just an organization's internal jargon, formalized and given the appearance of authority it hasn't actually earned yet.
Conducting search-log analysis and content analysis first is correct. The development process begins with research, understanding the archive's actual scope and the vocabulary people already use, before any preferred terms get defined. Relying on one senior editor's personal judgment, however experienced, substitutes one person's internal mental model for the broader research the process calls for, and risks producing a vocabulary that reflects how that editor thinks rather than how the full range of journalists, archivists, and public users actually search and describe topics. Implementing directly into the CMS without this groundwork risks building the entire technical structure around assumptions that research would have corrected far more cheaply upfront.
The challenge of thirty years of evolving language and daily new content affects vocabulary maintenance most directly: this is fundamentally a documentation and governance question, not a one-time research task. The source material is explicit that developing a controlled vocabulary requires ongoing effort as a domain evolves, and this news archive is about as demanding a case as that gets: political terminology shifts, new beats emerge, and yesterday's breaking-news topic becomes tomorrow's historical category.A reasonable ongoing approach is treating vocabulary maintenance as a recurring editorial responsibility rather than a project with an end date, periodically reviewing search logs for new or shifting terms, giving journalists and archivists a clear, documented process for proposing new preferred terms as new topics emerge, and keeping the vocabulary's documentation current so future team members understand not just what the preferred terms are, but why specific terminology decisions were made.