Why Classification Is Hard, and What Metadata Does About It

Eleanor Roach's two principles of categorization explain why classification feels intuitive when it's right and jarring when it's wrong: cognitive economy means categories should give people maximum information for minimum mental effort, and perceived world structure means people expect digital categories to map onto the same kind of structure they perceive in the physical world. Morville and Rosenfeld identify four concrete challenges that make classification genuinely difficult in practice: ambiguity, since language itself is imprecise and category boundaries blur; heterogeneity, when content varies too much in type or granularity for one scheme to fit cleanly; differences in perspective, since people bring different mental models shaped by their own knowledge and experience; and internal politics, where departments disagree about categorization for reasons that have nothing to do with users.

Metadata is the tool that makes classification operational. Three types matter most: descriptive metadata, title, topic, keywords, describes what a content object actually is, and has the greatest impact on navigation and search. Structural metadata describes a content object's internal structure, enabling the same content to be assembled and displayed differently in different contexts. Administrative metadata, owner, dates, status, supports the business processes of managing that content, like workflow and publishing.

Every "which category does this belong in" argument on a real project is usually one of these four challenges wearing a disguise, recognizing which one you're facing tells you what kind of resolution is even possible.

Exercise

The scenario: A large law firm has an internal knowledge base containing case notes, research memos, precedent documents, and client correspondence. Users consistently fail to find relevant documents even when they exist. Four specific problems have been identified.

1. The category "Litigation" means something different to the corporate team than to the disputes team, and each group insists their definition is correct, causing documents to be filed inconsistently.
2. Case notes, video-recorded depositions, scanned handwritten documents, and structured billing records all need to sit in the same knowledge base, but no single classification approach fits all of them well.
3. Junior associates categorize documents by the specific legal task they were working on; senior partners expect documents categorized by client and matter number instead.
4. The firm's two largest practice groups are each refusing to adopt a shared classification scheme, insisting their own group's existing system stay in place, mainly because neither wants to lose control over how their documents are organized.