A content inventory catalogs what content actually exists, every page, document, and content object, and where it lives in the current structure. It's a necessary first step before any redesign, migration, or merger, because you can't restructure or improve what you haven't fully accounted for. Before starting, you need to decide the inventory's scope and granularity: whole pages, or the individual chunks and components pages are made of.
A typical inventory spreadsheet captures descriptive attributes, page ID, title, content type, topics, location, alongside governance attributes like author, owner, publication date, and status. For very large sites, doing the entire inventory in one pass can be impractical; a "rolling content inventory," working through one section at a time and cycling back around, is often more sustainable than trying to capture everything in a single exhausting effort.
A content inventory is tedious in exactly the way that pays off later, every hour spent cataloging now is an hour someone doesn't have to spend guessing what exists once the redesign is underway.
Automating the bulk of the inventory, then adding manual detail selectively, is the realistic approach. Twelve years of accumulated articles makes a fully manual, page-by-page inventory impractical within eight weeks. The source material specifically recommends automating what you can, using a crawling tool or exporting structured data directly from the CMS, then layering manual detail on top only where it adds real value, rather than treating every single article as equally worth the team's limited hours.
Skipping the inventory is the option to rule out clearly. For a CMS migration specifically, the source material is direct that a thorough inventory is essential, without one, the team has no reliable way to know what exists, what's safe to leave behind, and what would break silently during the move, including the unknown broken links already present.A reasonable page ID scheme, following the numbering convention described in the source material: the home page as 0.0, top-level sections like "About Us" as 1.0, 2.0, and so on, individual pages within a section as 1.1, 1.2, individual news articles as their own numbered entries within an "Articles" section, and an embedded PDF linked within a specific article numbered relative to that article, for example, 3.4-PDF1. This lets the team distinguish a static page, an article, and an embedded document from one another at a glance, and trace exactly which article a given PDF belongs to.
Fill in your own details below; the prompt updates as you type. When it's ready, copy it into Claude or whatever AI tool you use.
Casts the AI as: You're a content analyst specializing in identifying redundant and duplicate content.
Names 4 specific outputs to produce, so the response comes back structured rather than a general summary.
Content domain, Known history and User impact tell the AI what your specific situation is, not a generic one.
Grouped findings by severity with specific merge recommendations
Fill in your own details below; the prompt updates as you type. When it's ready, copy it into Claude or whatever AI tool you use.
Identify outdated content based on update dates, product version history, and staleness signals. Prioritize content refresh based on user impact and staleness severity.
Fill in your own details below; the prompt updates as you type. When it's ready, copy it into Claude or whatever AI tool you use.
Use this prompt when: During content audit to find redundancy.