Technology

Why Does Your Sitemap Drift Away From Your Real Page Set?

Written by
Pravin Kumar
Published on
Oct 1, 2026

My sitemap says one thing and Google indexes another. Which one is wrong?

Usually the sitemap, and usually by neglect rather than error. A sitemap is a claim your site makes about which pages exist and when they changed. Nothing forces that claim to stay true, so on most sites it slowly stops matching reality while continuing to look perfectly valid.

This is worth caring about because the sitemap is one of the few direct statements you make to a crawler. Getting it wrong does not usually produce an error anybody sees. It produces a quiet mismatch, where you believe you have told search engines about your archive and the file is actually describing a site you had last spring.

I pulled my own sitemap while writing this, and I will use what I found as the example, because it is more honest than a hypothetical.

What is a sitemap actually for?

Discovery, not ranking. It helps a crawler find URLs it might otherwise reach slowly or not at all, which matters most on large sites, new sites, and sites with pages that are not well linked internally. It is a hint that a URL exists. It is not an instruction to index and it is not a quality signal.

Googlebot and every other crawler reads it that way. That distinction kills a lot of magical thinking. Adding a page to your sitemap does not make it rank, and removing a page does not remove it from the index. If you want a page out of search results, that is a different mechanism entirely, and conflating the two is one of the more common technical mistakes I get asked to unpick.

Where a sitemap genuinely earns its keep is on a CMS-driven archive that grows faster than anyone links to it. If you publish regularly and your older posts are reachable only through pagination, the sitemap is doing real work.

What does Google say it ignores?

Two things you may have spent time on. Its sitemap documentation states plainly that Google ignores the priority and changefreq values. Both of those tags still appear in generated sitemaps all over the web, and tuning them is effort spent on something the documentation says is not read.

I mention this because priority in particular invites a certain kind of wishful configuration. People set their key pages to 1.0 and feel they have done something. The honest version of that work is making those pages genuinely better linked and genuinely more useful, which is harder and is the thing that actually moves.

What Google does say it uses is the lastmod value, with a condition attached. Its wording is that Google uses lastmod if it is consistently and verifiably accurate, giving the example of comparing it to the last modification of the page. That condition is doing a lot of work in a short sentence.

Does lastmod matter, then?

Yes, and only if you are honest with it. Read that condition again. The value is used when it can be checked against the page and holds up. The implication is that a sitemap which stamps every URL with today's date, or which never updates at all, is teaching a crawler to stop trusting that field on your domain.

Google's guidance is also specific about what the date should mean, which is the last significant update to the page content, structured data or links, rather than a minor change like a copyright year ticking over. So a site-wide footer edit is not a reason for every lastmod to move.

This is the same principle as the visible date on a page. A date you treat carelessly is worse than no date, because it looks like information and is not. I have written about the page-level version of that problem, and the sitemap version is its mechanical twin.

What did my own sitemap actually say?

Something I was not entirely happy with. It returned successfully, about 199 kilobytes, listing 1,133 URLs in total of which 1,089 were blog posts. The newest last-modified date on any entry in the whole file was from the middle of September, on a site I add to regularly, which is precisely the drift this article is about.

That is not a disaster and nothing is broken. But it means the file was telling crawlers that nothing on this domain had changed in weeks, which was not true. The lesson I take from it is that generated sitemaps reflect whatever event the platform decided counts as a modification, and that may not be the event you care about.

I include this because it would be easy to write a confident piece about sitemap hygiene and quietly not check my own. The whole point of the exercise is that this happens to people who know better, which is why it is worth a scheduled look rather than a one-time fix.

What are the actual limits?

Google's documentation states that all formats limit a single sitemap to 50MB uncompressed or 50,000 URLs. Those are generous numbers and most sites will never approach them, but a large programmatic build or a big product catalog can, and the answer then is a sitemap index pointing at several files rather than one enormous one.

The number worth watching is not the limit, it is the count. If you know roughly how many pages your site should have and the sitemap says something different, you have found something. A gap of a few is noise. A gap of hundreds means either pages are missing from the file or the file is listing things that no longer exist.

Checking that count is a one-line job with curl and does not require a tool subscription. Counting the location entries in the file and comparing against what your CMS says you have published is the entire audit, and it takes a minute. Screaming Frog or a similar crawler will do it more thoroughly, but you do not need one to answer the first question.

Where does the drift come from?

Four places, in my experience. Pages that were deleted but are still listed. Pages that exist but were never added. URLs listed in a non-canonical form, which Google's guidance says to avoid by including only the canonical version you want in results. And dates that no longer describe anything real.

The non-canonical case is the sneakiest, because the page works when you click it. A trailing slash variant, an uppercase character, a stale URL from before a slug change, all resolve fine for a human and all quietly ask a crawler to treat a duplicate as the main version. Slug changes after publication are the usual culprit.

Deleted pages are the easiest to fix and the most often left. If a URL in your sitemap returns a not-found status, you are actively pointing crawlers at something that is gone, which is a small waste repeated at scale. That overlaps with ordinary link rot, which I covered in finding and fixing broken links.

Which pages should not be in there?

Anything you would not want as a search result. Thank-you pages, internal utility pages, duplicate filtered views, drafts that somehow became public, and anything you have deliberately kept out of the index. A sitemap listing a page you have told search engines not to index is your site contradicting itself.

That contradiction is worth hunting specifically, because it is invisible unless you look for it. Two systems, both configured by reasonable people at different times, producing opposite instructions about the same URL. Nothing errors. The crawler simply has to pick, and you have lost a decision you could have made yourself.

The same logic applies to crawler control more broadly. Your sitemap, your robots file and your page-level directives should be telling one consistent story, which is the point I was making in noindex against robots disallow.

What does ignoring this actually cost?

Mostly time, in the form of slower discovery of pages you want found and crawler attention spent on pages you do not. On a small site that cost is close to nothing and I would not lose sleep over it. On an archive of a thousand or more pages that publishes often, it compounds.

The larger cost is epistemic, and it is the one I care about. If your sitemap is unreliable, you have lost a measuring instrument. You can no longer use it to answer a simple question like how many pages this site has, and you will reach for Search Console or a crawler for something a file you own should have told you immediately.

That matters more now that several systems read your site besides Google. Bing has its own tooling in Bing Webmaster Tools, and the crawlers behind ChatGPT, Claude and Perplexity each vary in what they do, with the honest answer being that every vendor documents its own behavior and you should read theirs rather than assume mine. A sitemap that is simply accurate serves all of them without you needing to know.

What should you do next?

Fetch your own sitemap and count the URLs. Compare that number to what you think you have published. Then look at the newest last-modified date in the file and ask whether it matches when you last meaningfully changed something. Three checks, ten minutes, and you will know whether yours is describing your actual site.

If it is wrong, fix the count first and the dates second, because a missing page is a real loss and a stale date is a weakened signal. Then put the check on a recurring schedule rather than treating it as done. I found mine drifting while writing this piece, after 350 published articles, which tells you how much attention it quietly needs.

If you have a large Webflow archive and want someone to work out what your sitemap is really claiming and whether it matches your pages, reach out. It is a small job that tends to surface two or three other things worth knowing.

Get found, cited and the back office automated

Let's make your site the source AI engines quote and wire up the systems behind it.

Contact

Let's get your website found and cited by AI

Tell me what you're working on, whether AI search is skipping your product, your back office is buried in manual work, or you need a build that does both.

Got it, thanks. I read every message personally and reply within 1-2 business days.
Oops! Something went wrong while submitting the form.