Why has Google not come back to the page you updated last week?
Because recrawling is not a queue you join by editing something. Google decides how much of your site to crawl based on what your server can comfortably serve and how much it wants your pages, and an individual edit moves neither of those very much on its own.
This is one of the most common frustrations I hear from people running content-heavy sites. They rewrite a page, wait, check the cache, and conclude that something is broken. Usually nothing is broken. The page simply is not interesting enough, often enough, to jump the line.
Google documents how this works, and the mechanics are more useful to know than any trick for forcing a recrawl.
What actually decides how much Google crawls your site?
Google says a site's crawl budget is determined by two main elements, which are the crawl capacity limit and crawl demand. Capacity is about what your infrastructure can take. Demand is about how much Google wants your content. You need both to be healthy, and they fail in different ways.
Splitting the problem that way is the single most useful thing in Google's documentation on this, because the two halves have completely different fixes. A capacity problem is an engineering problem about server response. A demand problem is a content and structure problem about whether your pages are worth returning to.
Most people diagnose the wrong half. They assume Google is uninterested when the server is actually struggling, or they buy faster hosting when the real issue is that they have published four thousand pages nobody links to and nothing changes on.
What raises and lowers your crawl capacity limit?
Google defines the crawl capacity limit in terms of the total amount of time your server spends holding connections open for Google. It goes up if the site responds consistently and its response times remain stable or improve. It goes down if the site slows down, or responds with server errors or rate-limiting signals.
Google names the signals specifically. Server errors in the 5xx range and rate-limiting responses such as HTTP 429 both push the limit down. That is worth internalising, because a site under intermittent load can quietly train Google to ask for less, and the effect persists after the load event has passed.
The word doing the work in that definition is consistently. Stability matters more than raw speed here. A site that reliably answers in a moderate time is in a better position than one that is usually fast and occasionally times out, and the second pattern is far more common than people realise on sites with heavy third-party scripts or an overloaded database.
What makes Google want to crawl a page again?
Google names three factors you can influence on the demand side, which are perceived inventory, popularity and staleness. Perceived inventory covers duplicates and unwanted URLs that waste crawling time. Popularity means more popular URLs get crawled more often. Staleness is about recrawling frequency itself.
Perceived inventory is the one most teams can fix today and mostly ignore. If a large part of what Google can reach on your site is duplicate, parameterised or otherwise not worth having, you are spending your allocation on pages you did not want crawled. Cleaning that up is not glamorous work, and it is where I would start on almost any large site.
Popularity is the quiet reason internal linking matters beyond the reader. A page nobody links to, internally or externally, sits at the back of the queue. If you have a stack of pages in that position, an orphan page audit is a crawl intervention as much as a navigation one.
Does any of this apply to a small site?
Much less than you think. Google's own guidance is aimed at large sites with more than one million unique pages that change roughly weekly, at medium or larger sites with more than ten thousand unique pages that change daily, and at sites with a large proportion of URLs reported as discovered but currently not indexed.
If your site is a hundred pages, crawl budget is not your problem and optimising for it is a distraction. Your pages are not being recrawled slowly because of capacity. They are being recrawled at a rate that reflects how much demand there is, which is a different conversation about links, freshness and whether the page deserves attention.
I am fairly blunt about this with clients, because crawl budget has become a fashionable thing to worry about at every size. Read Google's thresholds, decide honestly which side of them you sit on, and spend your effort accordingly. The small-site version of this work is almost always writing better pages and linking to them properly.
How do you see your own crawl behaviour?
Search Console's Crawl stats report, which is the only honest view you have. Google says it shows total crawl requests, total download size and average response time, alongside host status and breakdowns by crawl responses, file type, crawl purpose and Googlebot type.
Average response time next to total crawl requests is the pairing I look at first, because together they tell you whether you have a capacity problem. If response time is climbing while requests fall, you are watching your capacity limit come down in real time, and that is an infrastructure conversation rather than a content one.
Host status is the second stop, and Google describes it as covering robots.txt fetching, DNS resolution and server connectivity. Those three break rarely and catastrophically. A robots.txt that intermittently fails to fetch is the kind of fault that produces bewildering crawl behaviour and never shows up in any content audit.
What is the difference between discovery and refresh crawling?
Google splits crawl purpose into exactly those two. Discovery means a URL that has never been crawled before, and refresh means a recrawl of a known page. The ratio between them tells you what Google is currently spending your site's crawl allocation on.
This breakdown answers a question most people try to answer by guessing. If you publish frequently and discovery is tiny, new pages are not being found, which is usually a linking or sitemap problem. If you publish rarely and discovery is large, something is generating URLs you did not intend, and that is inventory waste.
The file type and Googlebot type breakdowns are worth a look for the same reason. Google lists Googlebot types including smartphone, desktop, image, video, page resource load, AdsBot and StoreBot. Discovering that a large share of your crawl requests are image or resource loads reframes a slow-recrawl complaint entirely.
What should you change if recrawls are genuinely too slow?
Fix capacity signals first, because they are unambiguous. Eliminate server errors and rate limiting under normal traffic, and make response times consistent rather than occasionally excellent. Then reduce the inventory Google can reach but you do not want crawled, because those two moves address opposite halves of the same budget.
On the inventory side, the practical work is deciding what should exist in your sitemap and what should not, and making those two agree. Mismatches here cause more confusion than almost anything else, which is why I wrote separately about why your sitemap and your index coverage disagree. On Webflow specifically, controlling what gets into your sitemap is the lever you have.
What I would not do is treat a manual recrawl request as a strategy. It is a fine tool for a handful of important pages after a meaningful change. It is not a substitute for a site Google wants to come back to, and if you find yourself using it routinely, the underlying problem is demand.
What should you do next?
Open the Crawl stats report and look at three things. Average response time across the window the report covers, the split between discovery and refresh, and whether host status shows anything other than a clean record for robots.txt, DNS and server connectivity.
Then compare your page count against Google's stated thresholds and decide whether this is your problem at all. Being below them is good news, and it means your effort belongs somewhere else entirely. Be honest rather than flattered by the idea that you have a crawl budget issue.
If you are above them, work the two halves separately and in order. Capacity first because it is measurable, then inventory, then the linking that drives demand. If you are staring at a Crawl stats report that does not make sense next to what you are publishing, reach out and send me the screenshots.
Get found, cited and the back office automated
Let's make your site the source AI engines quote and wire up the systems behind it.
Read more blogs
Let's get your website found and cited by AI
Tell me what you're working on, whether AI search is skipping your product, your back office is buried in manual work, or you need a build that does both.