Why AI crawlers are different from Googlebot
Each AI company runs several crawlers, and they do different jobs. OpenAI's GPTBot collects pages to train models, OAI-SearchBot finds pages to show in ChatGPT search, and ChatGPT-User fetches a page when someone asks ChatGPT to. Anthropic and Perplexity split their crawlers the same way. Blocking one does not block the others.
That means a single line in robots.txt can keep your site out of AI training and still let assistants cite you, or, by accident, do the opposite. The check shows the verdict for each crawler, so you can see which of the two you have.
How to read the result
- Allowed: robots.txt lets the crawler fetch your homepage and our simulated request got the page.
- Blocked by robots.txt: a rule in your file tells the crawler not to fetch the homepage. We quote the rule, its line and its group. This result is exact.
- Blocked at server: robots.txt says yes, but your server or firewall refused our request or showed a challenge page. This result is simulated; confirm it in your firewall log.
- Unclear: the request was rate-limited, failed or returned a much shorter page than a browser gets.
Blocking training is a choice, blocking search is usually a mistake
Many sites block training crawlers on purpose, and that is a legitimate decision. We report it as information, not as a problem. Search and assistant crawlers are different: if they are blocked, your pages can't appear with a link in ChatGPT, Claude or Perplexity answers.
None of this affects Google Search or AdSense. Google-Extended, for example, only controls whether Google may use your pages for Gemini; Google says it does not affect your inclusion or ranking in Search.