Industry News

Cloudflare Now Sorts AI Crawlers by Purpose: What It Means for Your Site

Written by
Pravin Kumar
Published on
Jul 24, 2026

Should you care that Cloudflare just changed how AI crawlers reach your site?

Yes, if you want to show up in AI answers. On July 1, 2026, Cloudflare gave every customer, including the free tier, new controls that sort AI crawlers into three groups by purpose. That choice now shapes whether tools like ChatGPT and Google can read and cite your pages.

I read the announcement the day it landed because it touches the exact work I do every day, which is getting sites cited by AI answer engines. Most coverage framed it as a publisher story about ads and money. That misses the part that matters for founders and marketers who want visibility.

So I want to walk through what actually changed, in plain language, and what I would check on your own site this week.

What exactly did Cloudflare announce in July 2026?

Cloudflare rolled out new AI traffic options that let site owners allow or block automated bots based on why the bot is visiting. It splits AI crawlers into Search, Agent, and Training. The controls reached all plans on July 1, 2026, per Cloudflare's own blog, and free accounts got them too.

Cloudflare sits in front of a large share of the web as a content delivery network and security layer. When it changes a default, that default ripples across millions of sites at once. So a quiet settings update is not really quiet. It becomes a de facto standard for how bots and websites talk to each other.

The old model was blunt. You either let a crawler in or you blocked it, usually through robots.txt or a firewall rule. The new model asks a smarter question. It asks what the bot plans to do with your content once it has it.

What is the difference between Search, Agent, and Training crawlers?

Search crawlers index your content so an engine can answer questions about it later and point people back to you. Agent crawlers act in real time for a person, like ChatGPT fetching a page you asked about. Training crawlers pull your content to bake into a model. Same page, three very different trades.

Cloudflare frames Search as the category where you should expect referral traffic or other fair compensation in return. That is the trade most website owners actually want. You give an engine your content, and it sends readers or credit back to you.

Agent traffic is different. Cloudflare describes it as automated behavior acting on a person's behalf to get something done right now. Think of ChatGPT-User fetching a link, or Gemini and Claude driving a browser to complete a task. A human is often waiting on the other end, so this traffic can carry real intent.

Training is the category that makes people nervous. Your data gets permanently absorbed into the model itself, and there is usually no link back, no visit, and no credit. Once it is in the weights, it is in the weights.

Why does splitting crawlers by purpose matter for your traffic?

Because a single on-or-off switch forced a bad choice. Block everything and you vanish from AI search. Allow everything and you feed model training for free. Sorting by purpose lets you welcome the crawlers that can send readers back while limiting the ones that only take. That is a much fairer deal.

I care about this because the whole point of answer engine optimization is getting cited, not getting scraped into silence. If you block AI blindly, you can accidentally cut yourself out of ChatGPT, Perplexity, and Google's AI results. I wrote more about that tension in my post on which AI crawlers you should allow or block in robots.txt, and this Cloudflare move makes those decisions sharper.

The nuance is that Search and Training can come from the same company. Google crawls for classic Search, and it also has separate crawlers tied to model training. Treating them as one thing was always a mistake. Now you can treat them as what they are.

What happens by default now, and could it hurt your AI visibility?

For new domains onboarding to Cloudflare that display ads, Training and Agent crawlers are blocked by default, while Search stays allowed. That default protects publishers, but it can surprise anyone who never opened these settings. A default you did not choose is still a choice your site is making for you.

This is the part I would not ignore. Defaults are powerful because most people never change them. If your site runs ads and sits behind Cloudflare, you may now be blocking Agent traffic without realizing it. That could mean a person asking ChatGPT to pull up your page gets nothing back.

Whether that is good or bad depends on your goal. A publisher protecting premium articles may love it. A founder who wants to be the answer an AI hands to a buyer may hate it. The tooling is neutral. Your strategy is not, and only you can set it.

Does this only matter if you are a Cloudflare customer?

Directly, yes, these switches live in the Cloudflare dashboard. Indirectly, no. When the largest network provider normalizes purpose-based crawler rules, other platforms and AI companies tend to follow. The vocabulary of Search, Agent, and Training is likely to spread well beyond one vendor's control panel.

Even if your site runs on Webflow or anywhere else, the underlying shift affects you. AI companies now have a clear, public framework for how their bots should identify themselves and behave. That makes it easier for any site owner to reason about crawler intent, no matter who hosts them.

It also raises the bar on honesty. Crawlers that misrepresent their purpose to slip past a Search-only rule will face more scrutiny. That pressure is good for everyone who plays fair.

How does this connect to robots.txt and llms.txt?

It layers on top of them. robots.txt is a polite request that well-behaved bots choose to honor. Cloudflare's controls are enforced at the network edge, so they can actually stop traffic. llms.txt is a separate idea for guiding AI to your best content. Together they form a stack, not a single lever.

I think of it as three tiers. robots.txt states your intent, llms.txt points AI to what matters, and a network layer like Cloudflare enforces the rules when a bot ignores the polite version. If you want the deeper distinction between blocking and hiding, I broke it down in my piece on noindex versus robots.txt disallow for AI crawlers.

The takeaway is that no single file is a magic switch anymore. You need your polite signals and your enforced rules to agree. When they conflict, you get the worst of both, like inviting Search in through robots.txt while a firewall quietly slams the door.

What should you actually do about this on your own site?

Audit before you touch anything. Confirm whether you are behind Cloudflare, check which AI categories are allowed, and decide, per category, whether you want the trade. For most founders I work with, the answer is allow Search, allow Agent, and think hard about Training. Then align robots.txt to match.

Start by naming your goal in one sentence. If the goal is to be cited and to earn referral visits, you almost certainly want Search crawlers welcomed. Blocking them to feel safe is like unplugging your phone so no one can sell you anything. You also stop the calls you wanted.

Then decide how you feel about training. There is no universally right answer here. Some clients want their expertise reflected in models because it builds authority. Others guard proprietary content. Both are valid. What matters is that you decide on purpose rather than inherit a default. And remember that being cited still depends on people clicking, which I covered in whether people still click through when Google shows an AI answer.

What should you do next?

Pick one hour this week to review your crawler settings and your robots.txt together, category by category, and write down the trade you are choosing for each. If they disagree, fix them so they tell one story. That single review will teach you more than any think piece about AI and the open web.

This space keeps shifting, and the sites that win are the ones that make deliberate choices instead of drifting on defaults. If you want a second set of eyes on your crawler strategy, or you are not sure whether your Webflow site is quietly blocking the AI traffic you want, reach out through pravinkumar.co. I am always happy to talk through it with you.

Get found, cited and the back office automated

Let's make your site the source AI engines quote and wire up the systems behind it.

Contact

Let's get your website found and cited by AI

Tell me what you're working on, whether AI search is skipping your product, your back office is buried in manual work, or you need a build that does both.

Got it, thanks. I read every message personally and reply within 1-2 business days.
Oops! Something went wrong while submitting the form.