If your site is on Cloudflare, AI crawlers can be blocked in three different places: the AI bot policies (Search, Agent, Training), AI Crawl Control, where each crawler has an Allow or Block action, and the managed robots.txt, which adds Disallow: / groups for known AI training bots. The first two block the request at Cloudflare's edge, so your robots.txt never gets a say. Since 15 September 2026, Cloudflare's defaults for new domains block Training and Agent bots on pages that display ads, which matters for any AdSense site. Check all three before assuming robots.txt is the problem.
The three switches
| Feature | What it does | Where |
|---|---|---|
| AI bot policies | Block or allow by behaviour: Search, Agent, Training. Each has "Block (on all pages)", "Block on pages with ads" or "Allow (do not block)" | Security Settings > Configure AI bot policies |
| AI Crawl Control | A table of AI crawlers that visit your site, each with an Allow or Block action. Blocking creates a WAF custom rule | AI Crawl Control > Crawlers |
| Managed robots.txt | Prepends Disallow: / groups for AI training crawlers to your robots.txt, plus a Content-signal line |
Security Settings > Bot traffic > "Set your preference to block training in robots.txt" |
Sources: AI bot policies, Manage AI crawlers, managed robots.txt. Dashboard labels change from time to time; the names above are the ones in Cloudflare's docs at the time of writing.
AI bot policies and the new defaults
Cloudflare classifies AI bots by behaviour (AI bots):
- Search: collects or indexes your content so it can answer questions about it later.
- Agent: acts in real time on a person's behalf, "such as chat fetch bots and browser-use agents".
- Training: crawls content to train or fine-tune a model, including mixed-purpose crawlers used for both training and search.
Cloudflare's docs state: "On September 15, 2026, Cloudflare will set updated defaults for new domains: bots classified as Training or as Agent will be blocked on pages that display ads, and Search will remain allowed." "Pages that display ads" is detected automatically. If you added an AdSense site to Cloudflare after that date and never opened this screen, user-triggered fetches from AI assistants may be blocked on every page that shows ads, while pages without ads look fine.
Note the last point under Training: mixed-purpose crawlers that combine search and training are blocked by every option that blocks training. If you want to be found by a crawler that does both, you can't block training here and still allow it.
The older single toggle, "Block AI bots", is marked as deprecated on the same date.
AI Crawl Control
AI Crawl Control lists the AI crawlers requesting your pages, their request counts, robots.txt violations and an Action column. Cloudflare says that when you block a crawler there, "the system creates or updates a WAF custom rule on your zone to enforce that block". On the free plan it identifies crawlers by user agent string; Bot Management plans add Cloudflare's detection IDs (Get started).
Managed robots.txt
When this is on and your origin already serves a robots.txt with a 200 status, Cloudflare puts its managed block in front of your file. Cloudflare's own example of the managed content:
User-Agent: *
Content-signal: search=yes, ai-train=no, use=reference
Allow: /
User-agent: Amazonbot
Disallow: /
User-agent: Applebot-Extended
Disallow: /
User-agent: Bytespider
Disallow: /
User-agent: CCBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: Google-Extended
Disallow: /
User-agent: GPTBot
Disallow: /
User-agent: meta-externalagent
Disallow: /
That list contains training tokens, not OAI-SearchBot, Claude-SearchBot or PerplexityBot. Two practical consequences:
- You now have two
User-agent: *groups (Cloudflare's and yours). Under RFC 9309 groups that match the same crawler are combined, so your ownDisallowlines still apply. - If you edit robots.txt on your server and the change "doesn't show", look at the live file at
https://yourdomain.com/robots.txt; the managed part is added at Cloudflare, not on your server.
Cloudflare itself notes that "robots.txt compliance is voluntary" and that AI Crawl Control is how you enforce a block.
Other Cloudflare features that catch AI bots
AI-specific settings aren't the only ones. The same features that block Google can block AI crawlers:
- Bot Fight Mode challenges traffic that matches known bot patterns and, per Cloudflare, "cannot be customized, adjusted, or reconfigured via WAF custom rules" (Bot Fight Mode). Challenged requests show "Bot Fight Mode" in the Service field of Security Events.
- Custom WAF rules that challenge by country, ASN or empty referrer.
- Rate limiting with low thresholds.
- "I'm Under Attack" mode, which shows a challenge to everyone.
The Cloudflare and AdSense guide walks through these settings in detail.
Confirm what Cloudflare did
- Open Security > Analytics > Events and filter by user agent containing
GPTBot,OAI-SearchBot,ChatGPT-User,ClaudeBot,Claude-SearchBotorPerplexityBot. Each event names the action and the service or rule responsible. - Open AI Crawl Control > Crawlers and look at the Action column and the request trend for each crawler.
- Fetch the live robots.txt and read the managed block at the top.
You can also test from a terminal:
curl -sI -A "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; OAI-SearchBot/1.4; +https://openai.com/searchbot" https://example.com/
A 403 points to a block rule. But Cloudflare can tell your request doesn't come from OpenAI's IP ranges, so a 200 or a 403 here is not the whole truth. Security Events is.
Let AI search through, keep training blocked
- In Configure AI bot policies, set Search to Allow. Decide on Agent: "Allow" lets user-triggered fetches from ChatGPT, Claude or Perplexity read the page a user asked about. Set Training as you like.
- In AI Crawl Control, set the search crawlers you want (OAI-SearchBot, Claude-SearchBot, PerplexityBot) to Allow.
- If you keep the managed robots.txt, confirm the final file doesn't disallow those search tokens.
- If you wrote your own WAF rules, exempt verified bots rather than matching user agent text. Perplexity's documentation describes a WAF rule that combines its user agent with its published IP ranges; OpenAI publishes ranges for each of its crawlers too (OpenAI crawlers).
None of this touches AdSense. The AdSense crawler, Mediapartners-Google, is not in any of Cloudflare's AI categories in the docs above; if Google itself is being challenged, that's the Cloudflare and AdSense problem, not an AI setting.
Check it from outside
The free AI crawler checker requests your page with each AI bot's user agent and reads your live robots.txt, including Cloudflare's managed block, so you can see which bots get a challenge, a 403 or the page. Our requests are simulated: they carry the bot's user agent but come from our server, and Cloudflare can tell. Use the result to know where to look, then confirm in Security Events. For the AdSense side, the Googlebot access checker and a free scan check Google's crawlers.
FAQ
Did Cloudflare turn this on without asking me?
The 15 September 2026 defaults apply to new domains, according to Cloudflare's documentation. Older domains keep the settings they had, but check the screen anyway; someone on your team may have switched it on.
Does "Block on pages with ads" affect AdSense?
Cloudflare describes it as blocking the chosen AI behaviours only on pages where it detects ads. It's about AI bots. Nothing in the documentation says it applies to Google's ad crawlers.
I allowed GPTBot in robots.txt. Why is it still blocked?
robots.txt can only ask; it can't override a firewall. If Cloudflare blocks the request at the edge, the bot never reaches robots.txt or your page. Change the setting in Cloudflare.
Should I block AI bots at Cloudflare or in robots.txt?
robots.txt states your preference to the bots that honour it; a Cloudflare block enforces it. Many sites use robots.txt for training opt-outs and keep firewall blocks for crawlers that ignore it. The should I block AI training bots guide covers the trade-off.
Spotted something out of date or wrong? Tell us and we'll correct it.
Read this guide in Turkish →Check it on your own site. Free, no sign-up.
Free tools for this
Free scan
Check your own site
Free scan: readiness score and every issue, usually in a few minutes.