Anthropic, the company behind Claude, uses three bots. ClaudeBot collects web content that may be used to train its models. Claude-SearchBot indexes content to improve Claude's search results. Claude-User fetches pages when a Claude user asks a question. All three respect robots.txt, and each can be allowed or blocked on its own, so you can keep your content out of training while staying visible when people use Claude to search.
The three bots
From Anthropic's help article Does Anthropic crawl data from the web, and how can site owners block the crawler?:
| Bot | What it does | What happens if you block it |
|---|---|---|
ClaudeBot |
"Collecting web content that could potentially contribute to their training" | It "signals that the site's future materials should be excluded from our AI model training datasets" |
Claude-SearchBot |
"Navigates the web to improve search result quality for users" | "Prevents our system from indexing your content for search optimization, which may reduce your site's visibility and accuracy in user search results" |
Claude-User |
Accesses websites when "individuals ask questions to Claude" | "Prevents our system from retrieving your content in response to a user query, which may reduce your site's visibility for user-directed web search" |
Note "future materials": blocking ClaudeBot is about what's collected from now on.
What Anthropic commits to
The same article lists these principles:
- "Anthropic's Bots respect 'do not crawl' signals by honoring industry standard directives in robots.txt."
- "Anthropic's Bots respect anti-circumvention technologies (e.g., we will not attempt to bypass CAPTCHAs for the sites we crawl.)"
- Anthropic supports the non-standard
Crawl-delaydirective to slow crawling down. - Opt-outs apply per host: add the rule "for every subdomain that you wish to opt out from".
One consequence of the CAPTCHA commitment: if your CDN shows a challenge page to Anthropic's bots, they stop there. A challenge is a block, whether you meant it as one or not.
robots.txt examples
Visible in Claude, no training
User-agent: Claude-SearchBot
User-agent: Claude-User
Allow: /
User-agent: ClaudeBot
Disallow: /
Slow ClaudeBot down instead of blocking it
User-agent: ClaudeBot
Crawl-delay: 1
Crawl-delay isn't part of the robots.txt standard (RFC 9309) and support varies: Apple and Amazon say their bots don't follow it (Applebot, Amazonbot), while Anthropic says it does.
Block everything from Anthropic
User-agent: ClaudeBot
User-agent: Claude-SearchBot
User-agent: Claude-User
Disallow: /
Put the file on every subdomain you want covered: blog.example.com/robots.txt is a separate file from example.com/robots.txt.
Don't block by IP to opt out
It sounds stricter, but Anthropic warns: "Alternate methods like blocking IP address(es) from which Anthropic Bots operates may not work correctly or persistently guarantee an opt-out, as doing so impedes our ability to read your robots.txt file." If ClaudeBot can't read your robots.txt, it can't see that you opted out.
Anthropic does publish its IP addresses at claude.com/crawling/bots.json: "If a crawler has a source IP address on this list, it indicates that the crawler is coming from Anthropic." Use that list to verify a request that claims to be ClaudeBot, or to allow Anthropic's bots through a firewall, not to opt out.
Old names in copied block lists
Some "block AI bots" lists, including the list Squarespace adds when you tick its AI crawler box (Squarespace), contain anthropic-ai alongside ClaudeBot. Anthropic's current page names only the three bots above. Extra lines for other names don't hurt, but they don't replace a ClaudeBot group either, and none of these lists is a substitute for checking that Claude-SearchBot and Claude-User aren't blocked if you want to appear in Claude.
Firewalls and CDNs
Because Anthropic's bots don't bypass challenges, any bot challenge on your CDN acts as a block. Check:
- Cloudflare AI bot policies (Search, Agent, Training), AI Crawl Control and Bot Fight Mode: see Cloudflare blocking AI bots.
- Hosting and WordPress firewalls: see firewall and plugin rules.
Look in your logs for ClaudeBot, Claude-SearchBot and Claude-User and their status codes. A 403 or a challenge response means the request was stopped before robots.txt mattered.
Claude's bots and AdSense
They're unrelated to AdSense. Blocking or allowing ClaudeBot doesn't change what Google's AdSense crawler can see. The robots.txt and AdSense guide covers that crawler, and a free Approvalens scan checks AdSense readiness.
Check your setup
The free AI crawler checker shows whether your robots.txt allows ClaudeBot, Claude-SearchBot and Claude-User, and requests your page with their user agents to catch firewall blocks and challenge pages. Those requests are simulated from our server rather than coming from Anthropic's IP addresses, so a firewall that checks IPs may answer the real bots differently. The robots.txt tester shows the same file from Google's side.
FAQ
Does blocking ClaudeBot remove my site from Claude's answers?
Anthropic describes ClaudeBot as the training bot. Search and user-requested fetches use Claude-SearchBot and Claude-User, which you control separately.
Is Claude-User subject to robots.txt?
Anthropic says its bots honour robots.txt and that disabling Claude-User prevents its system from retrieving your content for user queries. That's different from OpenAI and Perplexity, which say robots.txt may not apply to their user-triggered fetchers.
How do I contact Anthropic about crawling?
The help article gives a contact address for crawler reports and asks you to write from an email address at the domain concerned, so they can verify it.
Does Crawl-delay work for Claude-SearchBot too?
Anthropic's example uses ClaudeBot. The article says Crawl-delay is supported "to limit crawling activity" without listing exceptions; add a group for each bot you want to slow down.
Spotted something out of date or wrong? Tell us and we'll correct it.
Read this guide in Turkish →Check it on your own site. Free, no sign-up.
Free tools for this
Free scan
Check your own site
Free scan: readiness score and every issue, usually in a few minutes.