Search-Log Analytics

A site's internal search log is a direct record of the actual words people use to describe what they want, arguably one of the most honest sources of vocabulary data available, since users type these terms with no awareness anyone is analyzing them. Lou Rosenfeld describes five approaches to analyzing this data: pattern analysis, which looks at trends in search behavior, vocabulary, and timing; failure analysis, which focuses on queries that return zero or poor results: these are top priority, since they represent confirmed, immediate findability failures; session analysis, which follows a single user's sequence of queries to see how their understanding evolves within one visit; audience analysis, which segments search behavior by user group; and goal-based analysis, which ties search performance directly to business KPIs.

Good search-log analysis balances two competing goals: precision, the percentage of returned results that are actually relevant, and recall, the percentage of all relevant content that the search actually surfaces. Learnings from search logs feed directly back into the rest of the IA work covered in this chapter, aligning controlled vocabulary and labels with users' real terms, identifying missing content, and informing where navigation or faceted search needs improvement.

A zero-result search query is one of the rare moments where a user has told you, in their own words, exactly what they wanted and exactly where you failed to give it to them.

Exercise

The scenario: A search log analysis for a professional membership organization's website has produced these findings: the query "renew membership" is searched 400 times a month and returns zero results, because the actual page is titled "Annual Subscription Continuation"; the query "conference schedule" is searched 150 times a month and returns three moderately relevant results; and the query "code of ethics" is searched 20 times a month and returns zero results because no such page currently exists on the site.

With limited resources to act on all three findings at once, which should the team prioritize fixing first?