How-To Guide

How to Prepare for Cloudflare's AI Bot Defaults

A step-by-step guide to Cloudflare's Search, Training and Agent controls: audit which pages show ads, set each category on purpose, avoid blocking Googlebot by accident, and check the result in your logs.

By Ramanath, CTO & Co-Founder at Presenc AI · Last updated: September 21, 2026

What Changed

Since July 1, 2026, Cloudflare controls AI traffic with three separate settings: Search (indexing content to answer questions later), Training (using content to train or fine-tune models) and Agent (acting in real time for a person). From September 15, 2026, domains newly onboarded to Cloudflare block Training and Agent by default on pages that display ads, and allow Search everywhere. Cloudflare applies the most restrictive rule that fits, so multi-purpose crawlers such as Googlebot, Applebot and BingBot are blocked wherever Training is blocked. The details are in Cloudflare's new AI bot defaults. This guide covers what to do about them.

Step 1: Find Out Which Defaults You Have

Open each zone in the Cloudflare dashboard and check the AI traffic controls under Security. Write down the Search, Training and Agent settings for each. Zones onboarded before September 15 keep whatever you set. Zones added later start on the new defaults. Don't assume every zone in an account matches, because agencies and multi-brand companies often have zones from different years.

Step 2: List Which Pages Show Ads

The new defaults only affect ad-monetized pages. List the templates that carry ad slots, such as blog posts, recipe pages or free tools, and separate them from product, pricing and documentation pages that don't. Brand sites often have fewer ad pages than they think, and publishers more.

Step 3: Decide Each Category on Purpose

CategoryAllow ifBlock if
SearchYou want to be cited in AI answers and search results. This is almost always the right choiceRarely. Blocking it removes you from AI answers
TrainingYou want the brand in future models' built-in knowledge, which helps with local and offline modelsThe content is your product, as for publishers, datasets and paid research
AgentAgents can complete a sale, booking or sign-up on the pageThe page earns only from ad views and agents add cost without value

For most brand sites the answer is to allow all three on commercial pages. Blocking agents on a product page turns away buyers. See optimizing for AI shopping agents.

Step 4: Check the Googlebot Trade-Off

Before blocking Training on any ad page, remember that crawlers which both index and train are caught by the more restrictive rule. If blocking Training would stop Googlebot or BingBot on pages you rely on for organic traffic, either allow Training on those pages or accept losing search on them. Check again as crawler operators split their bots by purpose.

Step 5: Line Up robots.txt and Content Signals

Cloudflare's network-level controls sit on top of robots.txt, and the two should agree. A robots.txt that allows GPTBot while Cloudflare blocks Training sends mixed signals and makes debugging harder. Cloudflare is also testing a Content Signals "use" field with three levels, immediate, reference and full, to state how content may be reused. It expresses a preference and does not block. See optimizing robots.txt for AI.

Step 6: Decide Whether to Charge

If you publish content that AI companies want, blocking is not your only choice. Cloudflare's Pay Per Crawl marketplace, now moving toward a per-use model, lets you charge instead. The setup is in setting up ai.txt and Pay Per Crawl, and pricing is in pricing content for AI crawlers.

Step 7: Check It in Your Logs

After any change, look at bot traffic for two to four weeks. You should see Training crawlers drop on blocked pages and Search crawlers stay steady. If Googlebot or OAI-SearchBot fall on pages you meant to keep open, a multi-purpose rule is catching them. Also watch your brand's mention rate in AI answers. That is how you'll know whether the change cost you any visibility.

Frequently Asked Questions

The automatic defaults apply to domains onboarded from September 15, 2026. Existing zones keep their current settings, but every customer, Free plan included, can set the Search, Training and Agent controls. Some reports say Free-plan zones were also switched, so check each zone's Security settings rather than assuming.
It can, on the pages where you block it. Cloudflare applies the most restrictive rule that fits, and its announcement says multi-purpose crawlers such as Googlebot, Applebot and BingBot are blocked where Training is blocked. Crawlers that declare a search-only purpose are still allowed.
Usually not on commercial pages. Agents that can read your product, pricing and sign-up pages can recommend you and complete purchases. Blocking agents mainly makes sense on pages that earn only from ad impressions.
Look at crawler traffic by bot for two to four weeks after the change, and track your brand's mention rate in AI answers over the same period. A drop in search crawlers or in mentions means a rule is broader than you meant.

Track Your AI Visibility

See how your brand appears across ChatGPT, Claude, Perplexity, and other AI platforms. Start monitoring today.