How do you run a content inventory before you hire a writer?
You build one spreadsheet with every published URL, attach the performance data you already have, and sort it into four piles: keep, refresh, merge, and remove. It takes an afternoon on most sites, and it turns a vague brief into a specific list of work before anyone starts writing.
This walkthrough is for one situation. You have a blog that has been running for a while, you are about to pay someone to write for you, and nobody currently knows what is already on the site.
Skipping this step is how businesses end up paying for an article they already published in 2023.
Why do this before hiring rather than after?
Because the first thing a good writer will ask is what exists already, and the second thing is what has worked. If you cannot answer either, you will get a content plan built on guesses, and you will pay for it twice: once to write it and once to fix the overlap.
There is a commercial reason too. An inventory usually reveals that a chunk of the work you were about to commission has already been done badly rather than not at all. Fixing an existing page is cheaper than writing a new one and it compounds with whatever authority the URL has already earned.
It also changes the shape of the hire. If most of your gaps are refreshes, you want an editor. If they are genuine gaps, you want a writer. Those are different people and you should know which one you are looking for before you post the role.
Step one: how do you get the list of everything you have published?
Pull it from the source of truth, which is your CMS rather than your sitemap. Export every published item with its URL, title, publication date, author, and category. If your CMS cannot export, use its API, and if you have neither, crawl the site and accept that you will miss anything unlinked.
Add a column for the date the page was last meaningfully edited, not the date the record was touched. Those two are almost always different, and only one of them tells you whether the content is stale.
Watch for the pages that are live but unreachable from anywhere on the site. Those will not appear in a crawl and they will quietly skew your conclusions. I wrote a separate walkthrough of that problem in auditing orphan pages on a large blog.
Step two: which performance data should you attach?
Clicks and impressions per URL from Search Console, plus whatever your analytics tool gives you for engaged time and conversions. Search Console's performance report lets you group by dimensions including queries, pages, countries, devices, search appearance and dates, and the pages dimension is the one you want here.
Mind the default date range. Google's documentation notes that the default view shows click and impression data for your site in Google Search results for the past three months, so widen it deliberately if you want a fairer picture of an older archive.
Export limits and row caps change over time, and they differ between the interface and the API, so check Search Console's own documentation for the current numbers rather than trusting a figure from a blog post. On a large archive you will almost certainly need the API or the bulk export rather than the download button.
Step three: how do you score each page?
Three columns, each a simple judgement rather than a formula. Does this page still reflect what we do. Does it get any search impressions at all. Would we be happy for a buyer to land on it today. Yes, no, or unsure for each.
Resist the urge to build a weighted scoring model. I have tried, and the model always ends up encoding what I already believed while making the process slower. Three honest columns and a human reading each title gets you to the same answer in a fraction of the time.
Do the scoring in one sitting if you can, because consistency matters more than precision. A pass done over three weeks will apply a different standard on day one and day twenty, and you will not be able to tell which pages were graded which way.
Step four: what do you do with the four piles?
Keep means leave it alone, and most of your archive should land here. Refresh means the topic is right and the content is out of date. Merge means two or more pages are competing to answer the same question. Remove means it should not have been published or no longer represents you.
Handle merge before refresh, because merging changes what needs refreshing. Two thin pages on the same topic become one good page, one redirect, and one fewer thing to maintain forever. That decision is the highest leverage output of the whole inventory.
The refresh pile is the one your writer should start on, and it needs its own process so that updating a page does not quietly cost you the traffic it already had. I set that out in my content refresh workflow for old blog posts.
What does the writer actually receive?
The spreadsheet, filtered to their pile, plus one page of standing rules. Not the whole archive, and not a list of topics without the existing URLs attached, which is the most common way this goes wrong. The filtering is the point, because a writer handed everything will simply start at the top.
For each refresh item, include the URL, what is out of date about it, the queries it already receives impressions for, and what you want it to say instead. That is four fields and it removes almost all of the back and forth that usually eats the first month of a writing engagement.
For genuine gaps, include the question the page should answer in the reader's own phrasing, and which existing pages it should link to. Building that mapping is worth doing properly, and I went through it in building a keyword to page map for a site you inherited.
What mistakes should you avoid?
Three, and I have made all of them. Deleting aggressively in the first pass, because remove feels decisive and is the hardest action to reverse. Scoring on traffic alone, which buries genuinely useful pages that were never going to rank. And running the inventory without the person who will act on it.
The deletion one deserves emphasis. A page with no traffic may still be the page your sales team sends to every prospect, and that will not show up in any export you run. Ask before you cut.
The third mistake is the quietest. An inventory produced by one person and handed to another is a document. An inventory produced together is a plan, and only one of those gets used.
How often should you repeat it?
Once a year for most businesses, and before any hire or restructure regardless of when you last did it. More often than that and you are measuring noise, less often and the archive drifts far enough that the exercise becomes a project rather than an afternoon.
Keep the old versions of the sheet with dates on them. The comparison between this year's piles and last year's is more informative than either sheet alone, because it shows you whether the refresh work actually happened.
With more than 350 published articles on my own site, I run a lighter version of this quarterly, and the only column I care about between full passes is whether anything has become factually wrong since it was published.
What should you do next?
Export your published posts to a spreadsheet this week, even if you do nothing else with it. Just seeing the count and the oldest publication date changes how most people think about their next content decision. It is the cheapest hour of clarity available to you.
Then add the Search Console pages data, score three columns by hand, and count how many land in refresh. If that number is larger than your planned new article count, you have just changed your brief and saved yourself money.
Across more than seventy projects I have never once run this exercise and found the archive was in the state the client expected. If you want someone to run the inventory and hand you the four piles with the merges already worked out, reach out and let's chat.
Get found, cited and the back office automated
Let's make your site the source AI engines quote and wire up the systems behind it.
Read more blogs
Let's get your website found and cited by AI
Tell me what you're working on, whether AI search is skipping your product, your back office is buried in manual work, or you need a build that does both.