Breaking
Google, Cloudflare, Microsoft Test AI Payment ModelsLazy Loading for Ads: The 2026 Marketer's PlaybookDisplay Ads Now Add Layers Without Replacing Old OnesCloudflare's AI Training Block Now Spares GooglebotAnonymizing Prompts Cuts GPT-4o Mini Retrieval Score 60%Google, Cloudflare, Microsoft Test AI Payment ModelsLazy Loading for Ads: The 2026 Marketer's PlaybookDisplay Ads Now Add Layers Without Replacing Old OnesCloudflare's AI Training Block Now Spares GooglebotAnonymizing Prompts Cuts GPT-4o Mini Retrieval Score 60%
SEO

Cloudflare’s AI Training Block Now Spares Googlebot

Cloudflare's new Disallow AI Training setting lets sites block AI training without losing Google, Bing, or Apple search visibility. Here's what to check before toggling.

Cloudflare's AI training block now spares Googlebot

Cloudflare has untangled one of the messiest choices in technical SEO: block AI training and risk losing search bots, or stay visible in search and feed AI models. On September 15, the company rolled out a new setting called Disallow AI Training on every plan, including the free tier.

Why this matters: Googlebot, Bingbot, and Applebot have been doing double duty. They index pages for search and gather content that can be used to train AI models. If a site blocked one of them to protect content, search visibility could disappear too. Cloudflare says fewer than 1% of sites on its network block search bots, while 17% have enabled some type of training block.

How the new setting works

When you enable Disallow AI Training, Cloudflare’s Bot Preference Sync adds two rules to robots.txt: one disallowing Google-Extended and one disallowing Applebot-Extended. Google and Apple use those tokens only for AI training. Googlebot, Bingbot, and Applebot remain allowed.

Amazon, Anthropic, Meta, and OpenAI already run separate crawlers for training and search, so Cloudflare blocks their training crawlers outright. Your pages should still be readable by ChatGPT search and similar tools.

Cloudflare now labels bot operators “Accountable” when they offer four things:

  • A robots.txt or similar opt-out for AI training
  • A way to opt out of AI summaries
  • URL-level visibility into which pages were made available for training
  • Assurance that opting out of training won’t affect search rankings

Google says URL-level reporting will arrive in the coming weeks. Apple’s is planned for next year. Microsoft is furthest behind: Bingbot won’t read the robots.txt training rule until early 2027. For now, the only way to keep Bing from training on a page is to add a NOARCHIVE meta tag.

Early data: search crawlers became more welcome

SeenSure tested 1,046 websites against eight AI crawlers on September 14 and again on September 16. Of the 746 Cloudflare-hosted sites, training crawler refusals rose after the change. GPTBot refusals climbed from 18.9% to 22.0%, and ClaudeBot refusals from 19.9% to 22.8%.

But refusals of AI search crawlers dropped by roughly 13 points each. OAI-SearchBot, which feeds ChatGPT search, was refused by 16.9% of Cloudflare sites before the change and only 3.0% after. The 300 non-Cloudflare sites stayed flat. As SEO consultant Aleyda Solis put it on X: “Search moved the most of any category.”

What to do before you toggle

The setting only controls training. It does not decide whether you appear in AI answers, and Cloudflare says AI summary controls are next on its roadmap. That leaves three separate decisions: search indexing, AI training, and AI answers.

Cloudflare’s default leaves training set to Allow for sites that don’t run ads. But for publishers that monetize content with ads, it recommends Disallow AI Training.

Remember: a robots.txt rule is a request, not a block. Cloudflare can detect and block crawlers that ignore it, but for Google, Apple, and Microsoft it is relying on commitments. Cloudflare says it will track progress publicly on Cloudflare Radar.

Before flipping the switch, check your AI crawler settings in Semrush Site Audit under Blocked from AI Search. It reads robots.txt and will show Google-Extended as blocked once the setting is on. It won’t detect firewall-level blocks, so check server logs to confirm search crawlers still get 200 status codes. Then record your AI visibility numbers in the AI Visibility Toolkit and compare them a few weeks later.

Source: Semrush Blog

Leave a Reply