AI

How Do AI Answer Engines Decide a Page Is Still Current?

Written by
Pravin Kumar
Published on
Oct 1, 2026

I updated the page. Why do AI engines still quote the old version?

Because an engine has to estimate how current a page is, and the date you changed is only one weak input into that estimate. If the words on the page still describe the old state of the world, a new timestamp does not change what gets quoted. Recency is inferred from content, not announced by metadata.

This is the single most common frustration I hear from people doing answer engine work, whether the surface is ChatGPT, Perplexity, Claude, Gemini or Google's own AI results. They did the update, the date moved, and three weeks later an assistant is still citing a claim they deleted. The gap is almost never a crawling problem. It is that the page did not actually change its mind about anything.

So the useful question is not how to signal freshness. It is what makes a page read as current to something that cannot see your changelog.

How does a search engine decide what date a page is?

Not from one place. Google's own documentation on publication dates states that it does not depend on a single date factor, because all factors can be prone to issues, and that its systems look at several factors to determine a best estimate of when a page was published or significantly updated. That phrasing is worth reading twice.

Google publishes that in its Search Central documentation, which is where I would send anyone who wants the wording rather than my summary. Two things follow from it. Your date is an estimate made by someone else, not a fact you control. And the estimate is of publication or significant update, which means a trivial edit is not what the system is trying to detect. It is looking for evidence that the page meaningfully changed.

The same documentation tells you to add a user-visible date to the page and feature it prominently, and to label dates clearly with text like "Publish" or "Last updated". That is a straightforward instruction that a surprising number of CMS templates ignore, either hiding the date entirely or showing one unlabeled number nobody can interpret.

What should the structured data say?

Google's guidance is to add a subtype of CreativeWork, such as Article, BlogPosting or VideoObject, and to specify the datePublished and dateModified fields. It also tells you not to specify future dates, or the date of the action described on the page, which rules out a mistake I see often on event and announcement pages.

That last point deserves emphasis because it is counterintuitive. If you publish on Monday about a conference happening in November, the date in your markup is Monday. The temptation to put the interesting date there is strong and it is wrong, and it produces pages that look like they are from the future.

Dates in that markup go in ISO 8601 format, and Googlebot is reading the page rather than your editorial calendar. Beyond that, keep the markup and the visible page agreeing with each other. A dateModified in your schema that contradicts the "Last updated" line a human can read is not a clever hedge. It is two conflicting claims on one page, and I have never seen that resolve in the site's favor.

Is a changed date the same as fresh content?

No, and treating them as equivalent is the core mistake. A date says when something happened to the file. Freshness, as a reader experiences it, is whether the page describes the current state of the thing it is about. Those come apart constantly, in both directions.

You can have a page edited yesterday that is badly stale, because the edit fixed a typo while the advice still assumes a tool that changed last year. You can also have a page untouched for two years that is perfectly current, because it explains a mechanism that has not moved. Neither is visible from the timestamp.

This is my own position rather than anyone's documented guidance, and I hold it firmly: rolling the date forward without changing the substance is a wasted edit at best. It spends the one signal you control on a page that does not deserve it, and it makes your own archive harder to audit, because you can no longer tell which pages you have genuinely reviewed.

Which signals actually say a page is current?

Specifics that could only be true now. A named version, a figure with a year attached, a reference to a thing that exists today, a sentence that acknowledges what changed and when. Those are what an engine can read as evidence of currency, because they are checkable against everything else it knows.

The inverse is also true, and it is why so many evergreen pages quietly rot. A page written entirely in timeless generalities has nothing in it that can be dated. That feels safe. In practice it means the page can never read as current, only as undatable, which is a weaker position when something has to choose between you and a page that states this year's specifics.

The most reliable move I know is to name the thing that changed. One sentence saying what is different now, with the date it changed, does more for how a page reads than any amount of metadata. It also forces you to confirm that something actually changed before you touch the date.

What makes an old page keep getting cited?

Being the clearest statement of something that is still true. Age is not a penalty when the content has not expired. A page from three years ago that explains a mechanism in plainer language than anything newer will keep getting pulled, because the alternative candidates are worse, not fresher.

What kills old pages is contradiction, not age. Once your old page says one thing and your newer page says another, you have handed a retrieval system a conflict, and the way it resolves that is not under your control. I wrote separately about how answer engines handle conflicting information, and the practical upshot is that two half-updated pages are worse than one honestly updated page.

The other killer is a broken premise. If a page opens by assuming something that is no longer the case, everything after it is suspect even where it is still correct. That is the one case where I would rather unpublish than update, because the whole frame is wrong rather than the details.

When should you update instead of writing a new page?

Update when the reader's question is unchanged and only your answer has moved. Write a new page when the question itself is new. That single distinction prevents most of the duplicate-content mess I get asked to untangle, and it is easier to apply than it sounds.

If someone would still search the same words to find this page, update it. If they would search different words, that is a different page, and forcing the new material into the old one produces something that serves neither query well. The URL should match the question, not the topic area.

Across 350 published articles I have ended up updating far less often than I expected, because most of what changes deserves its own page. The exception is anything that states a number, a limit or a price. Those I would rather update in place than leave a wrong figure live under an old URL.

How should you mark an update honestly?

Say what changed, in the page, in a sentence a human would write. Not "updated for 2026". Something like a short note at the top saying which section changed and why. That is useful to a reader, verifiable by anyone, and it gives a retrieval system real text to work with rather than a bare label.

I also keep the original claim visible when I have changed my mind, rather than silently deleting it. Partly that is honesty, and partly it is practical, because a page that says "I used to recommend this, here is why I no longer do" is far more quotable than one that simply presents a new opinion with no history.

What I avoid is the update banner that never comes down. If every page on your site says it was recently reviewed, the label carries no information and readers learn to ignore it, which is the same failure mode as a warning nobody acts on.

Where does this go wrong on a CMS-driven site?

In the template, usually. A Webflow CMS collection has whatever date fields you created, and the template shows whichever one the person who built it bound. I have opened sites where the visible date was the created date, the schema used a different field, and a bulk import had set both to the import day for every item at once.

Google Search Console will not flag that for you, and the Internet Archive's Wayback Machine is often the fastest way to find out what your dates used to say. That last one is worth checking specifically if you have ever migrated a blog. A migration that stamps 900 posts with the same date tells any system looking at your site that you published 900 pages in one day and have not updated since, which is a strange signal to send about an archive you spent years building.

The fix is unglamorous. Decide which field means published and which means updated, bind both deliberately in the template, label them in the markup, and spot-check a few items against what you actually know. I run this as part of a weekly citation audit rather than as a one-time task, because template changes reintroduce it.

What should you do next?

Open your three most important pages and read them as a stranger. Ask whether anything on each page could only be true now. If the answer is no, the page cannot read as current no matter what the timestamp says, and that is the thing to fix before touching any metadata.

Then check that your visible date and your structured data agree, and that neither is in the future. That takes ten minutes and removes an entire class of problem. The deeper work, which is keeping pages honest about what has changed, is the part that actually decides whether you get quoted, and it never stops being manual. Chunking matters here too, which I covered in why AI search quotes one paragraph.

If you have an archive that has drifted and you want help working out which pages to update, which to merge and which to retire, reach out. That triage is usually worth more than another round of new posts.

Get found, cited and the back office automated

Let's make your site the source AI engines quote and wire up the systems behind it.

Contact

Let's get your website found and cited by AI

Tell me what you're working on, whether AI search is skipping your product, your back office is buried in manual work, or you need a build that does both.

Got it, thanks. I read every message personally and reply within 1-2 business days.
Oops! Something went wrong while submitting the form.