Skip to content
Approvalens

Reading room · 5 min read

Should I Block AI Training Bots? Trade-offs for Publishers

Blocking GPTBot, ClaudeBot, Google-Extended and other training bots: what it changes, what it doesn't, how it relates to AI search, and AdSense.

By the Approvalens team

Fixes these report findings

  • AI training crawlers in robots.txt

You can block AI training bots without disappearing from AI search, because the big operators split the two: OpenAI (GPTBot vs OAI-SearchBot), Anthropic (ClaudeBot vs Claude-SearchBot), Google (Google-Extended vs Googlebot), Apple (Applebot-Extended vs Applebot) and Amazon (Amazonbot vs Amzn-SearchBot) each document separate controls. So the real question is narrower: are you comfortable with your content being used to train models? There's no documented ranking or traffic reward for allowing training, and no documented effect on AdSense either way. Block training if you'd rather not contribute; keep search bots open if you want to be cited.

Training vs search, company by company

Company Training control Search / answer control Source
OpenAI GPTBot OAI-SearchBot (ChatGPT search) OpenAI
Anthropic ClaudeBot Claude-SearchBot, Claude-User Anthropic
Google Google-Extended (Gemini training and grounding) Googlebot (Search, AI Overviews, AI Mode) Google
Apple Applebot-Extended Applebot; nosnippet for AI answers Apple
Amazon Amazonbot ("may be used to train Amazon AI models") Amzn-SearchBot ("does not crawl content for generative AI model training") Amazon
Perplexity none needed: PerplexityBot "is not used to crawl content for AI foundation models" PerplexityBot Perplexity
Common Crawl CCBot (open web archive) n/a Common Crawl

A typical "search yes, training no" block:

User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: Amazonbot
User-agent: CCBot
Disallow: /

The allow AI crawlers in robots.txt guide has the complete file.

Reasons to block

  • You don't want your work in training data. That's the documented purpose of these tokens, and the main reason most publishers use them.
  • Server load. Crawlers that don't send visitors back cost bandwidth. robots.txt blocks from compliant bots reduce that.
  • Licensing. Some publishers block by default and license content separately. Cloudflare's AI Crawl Control even lets you return a 402 Payment Required instead of a plain block (Manage AI crawlers).

Reasons to allow

  • Your goal is reach, not exclusivity. If you publish facts, documentation or how-to content you want widely repeated, training can spread it, though with no guarantee of credit.
  • Simplicity. Fewer rules mean fewer chances of blocking the wrong bot.

What you shouldn't count on: none of the companies above says allowing its training crawler improves your visibility in its search or assistant. Google states that Google-Extended "is not used as a ranking signal" in Search, and Apple says Applebot-Extended rules "are not considered in ranking for Search".

What blocking doesn't do

  • It isn't retroactive. Anthropic says blocking ClaudeBot excludes the site's "future materials"; Squarespace notes its AI crawler setting "doesn't retroactively remove content previously scraped" (Squarespace).
  • It's a request. Cloudflare's docs: "robots.txt compliance is voluntary" (managed robots.txt). The companies above say they honour it; others may not. Enforcing a block takes a firewall rule.
  • It doesn't cover user-triggered fetches. OpenAI and Perplexity say robots.txt may not apply to ChatGPT-User and Perplexity-User.
  • Mixed-purpose crawlers are all-or-nothing. Cloudflare's AI bot policies block crawlers used for both training and search under any setting that blocks training (AI bot policies).

The mistakes that cost visibility

Most damage comes from blocking more than intended:

Mistake What you lose
A copied list that includes OAI-SearchBot, Claude-SearchBot or PerplexityBot AI search citations
Blocking Googlebot to stop Gemini Google Search, including AI Overviews
Blocking Applebot instead of Applebot-Extended Spotlight, Siri and Safari results
Cloudflare "Block (on all pages)" for Search All AI search crawlers Cloudflare classifies as Search
Ticking every box in a host's blocklist, such as Netlify's User Agent Blocker, which offers both AI and SEO/Search crawler options Possibly search engines too; check what each option covers

What about AdSense?

Blocking or allowing AI training bots has no documented effect on AdSense. AdSense crawls with Mediapartners-Google, which only follows robots.txt groups that name it, and none of the training tokens above is that crawler. Two things to watch:

  • Don't over-block. A "block all bots" rule that hits Googlebot or Mediapartners-Google does hurt; the robots.txt and AdSense guide explains which groups matter.
  • Cloudflare's "pages with ads" option. Cloudflare can block AI Training or Agent bots only on pages where it detects ads, and that's the default for new domains since 15 September 2026. It targets AI bots, not Google's ad crawlers, but it's worth knowing it exists if you're puzzled why an AI assistant can't read your article pages. See Cloudflare blocking AI bots.

A sensible default

For most content sites applying for or running AdSense:

  1. Allow search crawlers: Googlebot, Mediapartners-Google, OAI-SearchBot, Claude-SearchBot, PerplexityBot, Applebot.
  2. Decide on training once, then list the training tokens explicitly.
  3. Check your CDN and plugins don't override that choice.
  4. Recheck after installing any security or SEO plugin.

Check your choices for free

The free AI crawler checker groups AI bots into search and training and shows which ones your robots.txt blocks, so you can see whether your file says what you meant. Page requests use simulated user agents from our server, so firewall results are a strong hint rather than proof. To make sure the AdSense and Google crawlers stay open, use the robots.txt tester or run a free scan.

FAQ

Will blocking AI training bots lower my Google rankings?

Google says Google-Extended isn't used as a ranking signal and doesn't affect inclusion in Search. Other companies' bots aren't part of Google Search.

If I block GPTBot, can ChatGPT still cite me?

Yes, if OAI-SearchBot is allowed. OpenAI documents them as independent.

Is blocking training bots enough to protect my content?

No technical measure fully protects public content. robots.txt covers compliant crawlers; a firewall can enforce blocks; neither removes what was collected before.

Do I need to block CCBot?

Common Crawl documents how to block it in robots.txt. Whether you should depends on whether you want your pages in its openly available archive.

Spotted something out of date or wrong? Tell us and we'll correct it.

Read this guide in Turkish →

Check it on your own site. Free, no sign-up.

Free tools for this

Free scan

Check your own site

Free scan: readiness score and every issue, usually in a few minutes.

Free scan · score and every problem found · no sign-up

All guides →