Skip to content
Approvalens

Reading room · 6 min read

How to Allow AI Crawlers in robots.txt (ChatGPT, Claude, Perplexity)

The exact robots.txt tokens for OpenAI, Anthropic, Perplexity, Google and Apple, how to allow AI search but block training, and mistakes that hide you.

By the Approvalens team

Fixes these report findings

  • AI search crawlers allowed in robots.txt
  • AI training crawlers in robots.txt
  • Sitemap pages open to AI crawlers

To let AI assistants find and cite your pages, your robots.txt must not disallow their search crawlers: OAI-SearchBot (ChatGPT search), Claude-SearchBot (Claude), PerplexityBot (Perplexity) and, for Google's AI features, plain Googlebot. Training crawlers such as GPTBot, ClaudeBot, CCBot and the Google-Extended and Applebot-Extended tokens are separate switches. You can block them and still be visible in AI search, because every company below documents search and training as independent settings.

The tokens that matter

Each operator publishes the exact token to use in a User-agent: line. These are copied from their own documentation, checked on the date at the top of this page.

Token Operator What it controls Source
OAI-SearchBot OpenAI Whether your pages can be shown in ChatGPT search answers OpenAI crawlers
GPTBot OpenAI Whether content may be used to train OpenAI's foundation models same
ChatGPT-User OpenAI Fetches made when a ChatGPT user asks; "robots.txt rules may not apply" same
Claude-SearchBot Anthropic Indexing for Claude's search results Anthropic crawlers
ClaudeBot Anthropic Collection of content that may be used for model training same
Claude-User Anthropic Fetches made when a Claude user asks a question same
PerplexityBot Perplexity Surfacing and linking sites in Perplexity search; "not used to crawl content for AI foundation models" Perplexity crawlers
Perplexity-User Perplexity User-requested fetches; "generally ignores robots.txt rules" same
Googlebot Google Google Search, including AI Overviews and AI Mode AI features and your website
Google-Extended Google Use of content for Gemini training and grounding; not a separate crawler Google's common crawlers
Applebot-Extended Apple Use of Applebot data to train Apple's foundation models; does not crawl About Applebot
CCBot Common Crawl Common Crawl's open web archive CCBot

Tokens are matched case-insensitively (RFC 9309), so gptbot and GPTBot are the same group. What does matter is spelling: Claude-Searchbot works, Claude Search Bot does not.

Allow AI search, block AI training

This is the setup most content sites want: be cited in AI answers, keep content out of training sets.

# AI search and assistants: allowed
User-agent: OAI-SearchBot
User-agent: Claude-SearchBot
User-agent: PerplexityBot
Allow: /

# AI training: not allowed
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: CCBot
User-agent: Google-Extended
User-agent: Applebot-Extended
Disallow: /

# Everyone else
User-agent: *
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php

Sitemap: https://example.com/sitemap.xml

A group can start with several User-agent: lines; the rules below apply to all of them (RFC 9309, section 2.1). OpenAI's own wording confirms the split works: "a webmaster can allow OAI-SearchBot in order to appear in search results while disallowing GPTBot". Perplexity and Amazon describe their user agents the same way, as independent settings.

The Allow: / lines in the search group are not strictly needed if nothing else blocks those bots, but they document intent and survive someone adding a broad block later.

Allow everything

If you're happy for your content to be used for both search and training, the shortest correct robots.txt names no AI bot at all:

User-agent: *
Disallow: /wp-admin/

Sitemap: https://example.com/sitemap.xml

Every crawler without its own group falls back to *. Problems start when the * group is restrictive, because AI bots inherit it. A staging leftover like User-agent: * + Disallow: / hides you from every AI search crawler at once. The robots.txt precedence guide explains the fallback rules in detail.

Mistake Effect Fix
Copying a "block all AI bots" list that includes OAI-SearchBot or PerplexityBot You disappear from ChatGPT search or Perplexity answers. OpenAI: sites opted out of OAI-SearchBot "will not be shown in ChatGPT search answers, though can still appear as navigational links" Keep search tokens out of the block list
Expecting Disallow for GPTBot to hide you from ChatGPT search It doesn't; GPTBot is training only Use OAI-SearchBot for search
Blocking Googlebot to keep content out of Gemini Removes you from Google Search, including AI Overviews Use Google-Extended for Gemini training and grounding
User-agent: * with Disallow: / Every AI bot without its own group is blocked Remove it from production
Disallowing /sitemap.xml or the folder it lives in Crawlers can't discover your URLs from it Let the sitemap path through for search bots
robots.txt returning a 5xx error Under RFC 9309 a crawler "MUST assume complete disallow" Serve a static file that always returns 200
Blocking a bot by IP in the firewall instead of robots.txt Anthropic warns IP blocks "may not work correctly" as an opt-out because they stop it reading robots.txt Use robots.txt for preferences, firewall rules for abuse

robots.txt is also cached. OpenAI says it "can take ~24 hours" for its search systems to pick up a change; Perplexity says "up to 24 hours"; Amazon may use a copy up to 30 days old (Amazonbot). Don't judge a change after an hour.

robots.txt is not the whole story

A perfect robots.txt still fails if the request never reaches your server. Cloudflare, Vercel, Netlify and many security plugins can block AI user agents in their firewall before robots.txt is consulted. If your robots.txt is open but AI tools still can't read your pages, check the Cloudflare AI bot settings and other firewall and plugin rules. Also check that your main text is in the HTML and not only rendered by JavaScript; see JavaScript content and AI crawlers.

AI bots and AdSense are separate

None of these tokens is the AdSense crawler. AdSense uses Mediapartners-Google, which only obeys groups that name it directly. Blocking or allowing GPTBot, ClaudeBot or Google-Extended doesn't change what AdSense can crawl. The robots.txt and AdSense guide covers the AdSense side, and the should I block AI training bots guide covers the trade-offs.

Check your robots.txt for AI bots

The free AI crawler checker reads your robots.txt the way each AI bot would, shows which search and training bots are allowed or blocked, and requests your page with their user agents. Those requests come from our server, not from OpenAI's or Anthropic's IP ranges, so a firewall that verifies bots by IP may treat the real crawler differently; treat a pass as a strong hint, not proof. For the AdSense side of the same file, use the robots.txt tester or run a free scan.

FAQ

Do I need to list every AI bot to allow it?

No. A bot with no group of its own follows User-agent: *. Named groups are only needed when you want a bot treated differently from everyone else.

Does disallowing GPTBot remove my site from ChatGPT?

Not from ChatGPT search. OpenAI documents GPTBot as the training crawler and OAI-SearchBot as the search crawler, and says each setting is independent.

Can robots.txt stop ChatGPT-User, Claude-User or Perplexity-User?

Partly at most. OpenAI says robots.txt "may not apply" to user-initiated ChatGPT-User fetches and Perplexity says Perplexity-User "generally ignores robots.txt rules". Anthropic says disabling Claude-User prevents its system from retrieving your content for user queries.

Is there one line that allows AI search and blocks all training?

No. Training opt-outs are per company, so you list each token. Cloudflare's managed robots.txt also adds a Content-signal: search=yes, ai-train=no line, but that is Cloudflare's own signal, not a standard every crawler reads.

Spotted something out of date or wrong? Tell us and we'll correct it.

Read this guide in Turkish →

Check it on your own site. Free, no sign-up.

Free tools for this

Free scan

Check your own site

Free scan: readiness score and every issue, usually in a few minutes.

Free scan · score and every problem found · no sign-up

All guides →