The Cloudflare Problem Nobody's Reading About the Right Way
2026-08-04 — marketing implementation consulting
Someone on Reddit tried to block AI training bots. Googlebot started returning 403s. They panicked.
Worth paying attention to.
Here's what actually happened, stripped of the noise. When the user set AI Training to Block, both Googlebot and Bingbot started receiving HTTP 403 responses when trying to fetch the sitemap. They turned it off. Bots came back. This person wasn't making it up—they could see in the Cloudflare dashboard that Googlebot and Bingbot were shown as blocked automatically. Google's John Mueller got involved.
Now, the tricky part.
Cloudflare isn't broken. The logic is actually intentional, just poorly named. Cloudflare's new AI crawler defaults treat mixed-use bots like Googlebot under the most restrictive rule that applies to them, and because Googlebot crawls for both Search and AI training in a single bot, sites that block Training will also block Googlebot on those pages unless they explicitly opt out. Googlebot does two jobs at once. It indexes your site for Google Search. It also collects content for AI training—powering Google Gemini, AI Overviews, and AI Mode. One bot. Two behaviors. That's the wrinkle.
Most consultants get this wrong, including us sometimes.
The deadline matters. Cloudflare announced on July 1, 2026, that starting September 15, new default rules will treat crawlers that serve more than one purpose, search indexing and AI training combined, under whichever rule is most restrictive. So if your settings already block Training, September 15 is when Cloudflare will start enforcing that logic globally. But here's what people keep glossing over—new defaults apply only to domains newly onboarding to Cloudflare. Existing sites don't get migrated to the new logic unless they're on the Free tier and haven't touched their settings.
Actually, that's not quite right.
Let me be more precise. Multi-purpose crawlers are evaluated by all of their behaviors, and if you block Training, per Cloudflare you can also end up blocking Googlebot. The question is timing. Are you seeing this now, or only after September 15? The Reddit post suggests it's happening already, which points either to Bot Fight Mode being overly aggressive or to something else running in their setup. Cloudflare's actual statement doesn't explain this gap.
What you should do.
Check your Cloudflare dashboard before the deadline. Cloudflare's documentation lets you set your preference under Security Settings > Block AI bots. Cloudflare's own announcement breaks down the three categories: Search, Agent, and Training. If you want Google's content for search results but not for its AI products, the solution isn't to block all AI. You'd specifically block Training, then manually allow Search—or better, contact Cloudflare support before September 15 to understand your exact setup.
The other option: Google offers a bot called Google-Extended that lets site owners opt out of having content used for training and AI products such as Gemini Apps and Vertex AI, without affecting inclusion in Google Search. You could use robots.txt to block Google-Extended instead of fighting with Cloudflare's logic.
This isn't a Cloudflare bug. It's a design that treats search and AI training as inseparable when they technically aren't—they're just bundled. That bundle is what some coverage frames as Cloudflare pushing back on Google. Google has the option to split its crawlers. It hasn't.
So now you're the one caught between two incompatible rules.