Breaking
Instagram Tools 2026: Scheduling, AI and Analytics ConvergeHow to Check If Your Page Shows Up in AI SearchWhat Is Agentic SEO? A Repeatable AI WorkflowWhy Most Creator Ambassador Programs Underdeliver2026 Social Algorithms: Ranking Signals That MatterInstagram Tools 2026: Scheduling, AI and Analytics ConvergeHow to Check If Your Page Shows Up in AI SearchWhat Is Agentic SEO? A Repeatable AI WorkflowWhy Most Creator Ambassador Programs Underdeliver2026 Social Algorithms: Ranking Signals That Matter
SEO

Cloudflare’s New AI Training Setting Keeps Googlebot Crawling

Cloudflare’s Disallow AI Training setting blocks AI crawlers while preserving Googlebot, Applebot, and Bingbot search access. Here’s what to check.

Cloudflare Keeps Googlebot While Blocking AI Training

Cloudflare has shipped a new Disallow AI Training control that separates AI training opt-outs from traditional search crawling. Instead of forcing sites to choose between visibility and privacy, the setting blocks most training crawlers while still letting Googlebot, Applebot, and Bingbot index content for search.

The change reverses a more aggressive July plan. Back then, Cloudflare said sites that blocked AI training would also block those three mixed-use crawlers because they feed both search and model training. Now, selecting Block shuts down Googlebot, Applebot, and Bingbot entirely, including search crawling. The new setting avoids that trade-off.

What changed on September 15

Disallow AI Training lives under Cloudflare’s Training control, alongside Search and Agent settings. When enabled, mixed-use crawlers remain allowed for search only if Cloudflare considers the operator “Accountable.” Other training crawlers are blocked.

Existing Training selections of Block or Block on pages with ads will migrate automatically. For most sites, no manual action is needed. Sites that previously used blocking under the older Block AI Bots toggle will see different presets: Allow for Search, Disallow AI Training for Training, and Block on pages with ads for Agent. Cloudflare also says Block AI Bots and the Managed Robots.txt feature will be deprecated. To keep mixed-use crawlers off entirely, you now need to select Block.

Who gets “Accountable” status

Cloudflare created the Accountable designation after conversations with crawler operators. To qualify, a company has to meet or commit to four requirements:

  • A robots.txt or similar opt-out for AI training.
  • A way to opt out of AI summaries, now with the operator and later through Cloudflare.
  • URL-level visibility into pages used for training plus search performance metrics.
  • An assurance that training opt-outs won’t hurt traditional search results.

Apple, Google, and Microsoft meet the requirements, with some features live now and others tied to deadlines. Amazon, Anthropic, Meta, and OpenAI are also listed as Accountable because they run separate search and training crawlers. Their training crawlers remain blocked under Disallow AI Training.

What the setting means per platform

Google: Cloudflare applies a Disallow rule for Google-Extended, the token that opts content out of Gemini model training. Google’s documentation says this does not affect Search inclusion or ranking. AI Overviews, AI Mode, and Discover generative features are controlled separately in Search Console.

Apple: The setting uses Applebot-Extended, which Apple says does not crawl pages or influence search ranking. To keep content out of AI-generated answers in Siri and Search, you still need the nosnippet meta tag.

Bing: Microsoft does not yet support a robots.txt no-training preference through Cloudflare. Bing’s current training opt-out is the NOARCHIVE meta tag. Content marked NOARCHIVE won’t be used to train Microsoft’s generative AI models, and it won’t be linked in Chat or Copilot.

Why marketers should audit this now

The biggest risk is choosing Block instead of Disallow AI Training. That stops AI training crawlers but also removes Google, Apple, and Bing from your pages, crushing search visibility. For performance marketers and SEO teams, the practical framework is simple: separate search access from training access, then verify each platform’s opt-out.

Check robots.txt rules, confirm Google Search Console AI feature settings, and add NOARCHIVE if Bing exclusion matters for your brand. If you want to monitor how content is used, Google plans URL-level transparency tools for Google-Extended in the coming weeks. Apple’s equivalent is expected next year, and Microsoft’s robots.txt no-training support is targeted for early 2027.

Source: Search Engine Journal

Leave a Reply