A content audit is a systematic review of everything published on a website, judged against what it is supposed to achieve. Every page is inventoried, measured on traffic, rankings, engagement, accuracy, and business relevance, and then assigned a decision: keep it, improve it, merge it into something stronger, or retire it. The output is not a report. It is a list of actions with owners.
Most sites need one more than their owners expect. Content accumulates over years, through staff changes, campaigns nobody remembers, and pages created because someone asked. The result is usually a site where a minority of pages do nearly all the work and the rest quietly drag down the average.
Why audit rather than just publish more
Content decays. Facts age, tools get renamed, competitors publish something better. A page that ranked two years ago often needs updating rather than replacing, and updating it is usually cheaper and faster than writing a new one.
Your pages compete with each other. When several pages target the same question, search engines have to choose, and they often choose badly or split the signals between them. Consolidating them into one stronger page is one of the most reliable wins available.
Thin pages at scale drag the site. Google evaluates quality partly at site level. A large number of shallow pages can suppress the performance of the good ones around them.
Stale content costs trust. Outdated prices, retired services, and old statistics undermine confidence in everything else, including with AI systems that draw on your pages to describe you.
Crawl attention is finite. Effort spent fetching low-value URLs is effort not spent on the pages you care about.
You cannot plan without knowing what you have. Almost every content strategy engagement starts here, because deciding what to remove usually matters more than deciding what to add.
When to run one
Before setting a content strategy. Before a redesign or migration, so you are not rebuilding pages that should not survive. After a traffic drop you cannot explain. When a site passes a few hundred pages and nobody can say what is on it. And on a routine cycle after that, annually for most sites, more often for large or fast-moving ones.
What to collect
One row per URL, with enough columns to make a decision without opening every page:
Identity: URL, title, content type, publication and last-updated dates, author.
Performance: sessions and engaged sessions from GA4, impressions, clicks, and average position from Search Console, and conversions or assisted conversions where you can attribute them.
Search context: the primary query the page actually ranks for, which is frequently not the one it was written for, and whether another of your pages competes for it.
Technical status: status code, indexability, canonical target, and number of internal links pointing to it.
Substance: word count, whether the facts are current, and what the page is actually for.
The data comes from four places: a crawl (Screaming Frog or similar) for the technical columns, Search Console for search performance, GA4 for behavior and conversions, and a CMS export for dates and authorship. Joining them on URL is the whole technical exercise.
The four decisions
Keep
The page performs, the facts are current, and it serves a clear purpose. Leave it alone and note when to review it again.
Improve
The page has potential it is not realizing: ranking on page two, covering the subject partially, missing the direct answer up top, or simply out of date. This is usually the largest and most valuable bucket, because improving a page with existing history and links outperforms publishing a new one. Pages ranking just below the fold of page one are the obvious place to start.
Merge
Several pages cover nearly the same ground. Pick the strongest as the survivor, usually the one with the best existing rankings and links, fold the genuinely useful material from the others into it, and 301 redirect the retired URLs to it. Update any internal links that pointed to the old pages so they go straight to the survivor rather than through a redirect.
Retire
No traffic, no rankings, no business purpose, nothing worth salvaging. Redirect to the closest relevant page if one exists, because the URL may still have links or be in someone’s bookmarks. If nothing relevant exists, let it return a 404 or 410 rather than redirecting it to the homepage, which helps nobody and looks like a soft error to search engines. Noindex is the right tool for pages that should stay reachable but do not belong in search, such as thank-you pages and thin archives.
A worked example from our own site
Our glossary reached 140 entries, most of them one-paragraph definitions written to be filled in later. That is thin content at scale on our own site, and it was holding back the entries we had properly developed. The audit sorted all 140 into buckets: a set to expand first on the highest-value terms, a larger set to expand in batches over time, the ones already finished, four to merge into surviving entries, and a group to noindex because expanding them was never going to be worth the effort. The noindex decision is the one people resist, and it is usually the right call. A section that is three-quarters thin drags down the quarter that is excellent, and putting the unfinished entries out of the index until their batch ships means the section’s quality signal reflects the finished work.
The lesson generalizes: an audit that only ever says “improve” is not an audit. Deciding what will never be worth finishing is the part that changes outcomes.
Prioritizing the work
An audit of a large site produces more actions than anyone can do at once. Sequence by effort against return: pages ranking on page two that need small improvements, then merges that resolve pages competing with each other, then factual updates on pages that still get traffic, then structural work grouping content into topic clusters, then retirements. Do the consolidation before commissioning anything new, or the new content lands on top of the problem it was meant to solve.
Common mistakes
Judging on traffic alone. A page with fifty visits a month that converts a third of them is not underperforming.
Auditing and not acting. A spreadsheet nobody works through is a cost with no return.
Deleting without redirecting. Retired URLs with links or traffic should point somewhere sensible.
Redirecting everything to the homepage. Irrelevant redirects are treated as soft 404s and frustrate visitors.
Ignoring pages outside the blog. Service pages, landing pages, and old campaign URLs need auditing too, and usually matter more commercially.
Judging too soon after a change. Give updates time to be recrawled and reassessed before drawing conclusions.
Treating it as one-off. Content decays continuously, so the audit has to recur.
Making it routine
The heavy audit is the first one. After that, keep the inventory current, set review dates on pages whose facts age, and check a slice of the site each quarter rather than repeating the whole exercise. Building the update cycle into the content plan from the start is what keeps a library from becoming a liability, and it is a standing part of how we approach topical authority.
Our SEO audit services produce exactly this: every page mapped against search intent with a keep, improve, merge, or retire decision and a prioritized order of work, alongside the technical review that runs beside it. If you have more content than anyone can account for, book a discovery call.