Skip to content
Approvalens

Tool 01 · Free · no sign-up

Can ChatGPT see your site?

We check whether GPTBot, ClaudeBot, PerplexityBot and 11 other AI crawlers may read your homepage: the exact robots.txt rule for each one, what your server answers when they knock, noai tags, JavaScript, llms.txt and your sitemap. Free, no sign-up, about 20 seconds.

Free · no sign-up · homepage only · about 20 seconds

What we test

Six things decide whether an AI assistant can read your site

  1. 1

    robots.txt, per crawler

    We read your robots.txt the way RFC 9309 describes and find the group and the exact rule that applies to each AI crawler. Exact.

  2. 2

    Server access, including Cloudflare

    We request your homepage with each crawler's user agent and spot refusals, challenge pages and rate limits. Simulated.

  3. 3

    noai directives

    Robots meta tags and X-Robots-Tag headers with noai, noimageai, noindex or nosnippet.

  4. 4

    JavaScript dependency

    We compare the words in the raw HTML with the words after the page renders. Most AI crawlers don't run JavaScript.

  5. 5

    llms.txt

    Whether you have one and, if so, whether it follows the proposed format, problem by problem.

  6. 6

    Sitemap

    We sample pages from your sitemap and check which AI crawlers robots.txt keeps out of them.

Who we test

14 AI crawlers and robots.txt names

Each links to its vendor's own documentation.

Why AI crawlers are different from Googlebot

Each AI company runs several crawlers, and they do different jobs. OpenAI's GPTBot collects pages to train models, OAI-SearchBot finds pages to show in ChatGPT search, and ChatGPT-User fetches a page when someone asks ChatGPT to. Anthropic and Perplexity split their crawlers the same way. Blocking one does not block the others.

That means a single line in robots.txt can keep your site out of AI training and still let assistants cite you, or, by accident, do the opposite. The check shows the verdict for each crawler, so you can see which of the two you have.

How to read the result

  • Allowed: robots.txt lets the crawler fetch your homepage and our simulated request got the page.
  • Blocked by robots.txt: a rule in your file tells the crawler not to fetch the homepage. We quote the rule, its line and its group. This result is exact.
  • Blocked at server: robots.txt says yes, but your server or firewall refused our request or showed a challenge page. This result is simulated; confirm it in your firewall log.
  • Unclear: the request was rate-limited, failed or returned a much shorter page than a browser gets.

Blocking training is a choice, blocking search is usually a mistake

Many sites block training crawlers on purpose, and that is a legitimate decision. We report it as information, not as a problem. Search and assistant crawlers are different: if they are blocked, your pages can't appear with a link in ChatGPT, Claude or Perplexity answers.

None of this affects Google Search or AdSense. Google-Extended, for example, only controls whether Google may use your pages for Gemini; Google says it does not affect your inclusion or ranking in Search.

Pricing

Start free. Go deeper when it matters.

Free check

$0homepage

  • Your homepage against 14 AI crawlers
  • robots.txt rules and simulated server access
  • noai, JavaScript, llms.txt and sitemap
  • A shareable result link
Most complete

Full AI visibility report

$9one-time, per site

  • Your whole site crawled, up to 20,000 pages
  • Every page against every AI crawler we test
  • Every AI-visibility problem with its evidence and fix steps
  • An llms.txt built from your own pages
  • Unlimited rescans for one month

One-time payment via Lemon Squeezy · no account needed · open for 30 days

Monitoring

AdSense & AI site health monitoring

$4.99per month

  • Uptime: your homepage loaded every 5 minutes, an email when it goes down and when it is back
  • TLS certificate: a warning 14 and 7 days before it expires, and at once if it becomes invalid
  • ads.txt: an alert if it disappears or loses your google.com publisher line
  • Google access: robots.txt rules plus simulated homepage requests as Googlebot and Mediapartners-Google (AdSense)
  • 14 AI crawlers: robots.txt rules, simulated homepage requests and llms.txt, rechecked daily
  • Alerts only when something changes, with before and after, plus a weekly summary
What “simulated” means

Crawler requests come from our server with the crawler's user agent. A firewall that verifies crawlers by IP address can treat the real Googlebot or GPTBot differently, so a server-side change has to show on two daily checks in a row before we email you. robots.txt, ads.txt, llms.txt and TLS changes are emailed the same day.

Every full report already includes one month of this monitoring for its site.

Questions

Frequently asked questions

Q1

Does blocking GPTBot remove my site from ChatGPT?

Not from ChatGPT search. OpenAI documents three agents: GPTBot collects content for training, OAI-SearchBot finds pages to show in ChatGPT search results, and ChatGPT-User fetches a page when a user asks. Blocking GPTBot keeps your pages out of training; to appear in ChatGPT search, OAI-SearchBot must be allowed.

Source:developers.openai.com

Q2

What is the difference between ClaudeBot, Claude-SearchBot and Claude-User?

Anthropic says ClaudeBot collects public web content that may be used to train its models, Claude-SearchBot crawls to improve the quality of search results, and Claude-User fetches pages when a Claude user asks a question. Anthropic states that its bots honour robots.txt, and each one can be blocked separately.

Source:support.claude.com

Q3

Does Perplexity follow robots.txt?

Perplexity documents two agents. PerplexityBot finds and links sites in Perplexity's search results and follows robots.txt. Perplexity-User fetches a page because a user asked for it, and Perplexity says it generally ignores robots.txt for those requests.

Source:docs.perplexity.ai

Q4

Does blocking Google-Extended hurt my Google ranking?

No. Google describes Google-Extended as a standalone robots.txt product token that controls whether your content may be used for its Gemini models; it is not a separate crawler and it does not affect inclusion or ranking in Google Search.

Source:developers.google.com

Q5

Why would Cloudflare block AI crawlers on my site?

Cloudflare offers a setting that blocks known AI crawlers across a site, and it can be switched on from the dashboard in one step. If it is on, AI crawlers get refused or challenged even when your robots.txt allows them. Our server test shows this, but only as a simulation, because Cloudflare can recognise real crawlers by their network.

Source:developers.cloudflare.com

Q6

What is llms.txt, and do I need one?

llms.txt is a proposal for a Markdown file at /llms.txt that gives AI tools a short map of your site: a title, a summary and links to your key pages. It is not a standard and the major AI crawlers do not say they use it, so we report a missing file as information. If you have one, we check it against the proposed format.

Source:llmstxt.org

Q7

Do noai and noimageai meta tags work?

They are not part of any standard, and Google's list of supported robots meta rules does not include them. They state your wish but do not block a crawler. To keep a crawler out, use robots.txt, and to keep pages out of search entirely, use noindex.

Source:developers.google.com

Q8

How exact is this check?

The robots.txt part is exact: we read your file the way RFC 9309 describes, with the most specific group and the longest matching rule. The server part is simulated: we send each crawler's user agent from our own servers, and a firewall that verifies bots by IP address may treat the real crawler differently.

Source:rfc-editor.org

Reading room

Guides that use this tool

Where this check comes up, and how to fix what it finds.

Other free tools