Skip to content
Approvalens

Reading room · 4 min read

Can ChatGPT See My Website? A 5-Minute Check

Whether ChatGPT can find and read your site depends on OAI-SearchBot, ChatGPT-User, your firewall and your HTML. How to check each one and fix what blocks it.

By the Approvalens team

Fixes these report findings

  • AI crawlers get your pages from the server
  • AI search crawlers allowed in robots.txt
  • Content is in the HTML AI crawlers fetch

ChatGPT reaches websites in two ways: OAI-SearchBot crawls pages so they can be shown in ChatGPT search answers, and ChatGPT-User fetches a page when someone asks ChatGPT about it. For ChatGPT to see your site, OAI-SearchBot must be allowed in robots.txt, your firewall or CDN must let OpenAI's requests through, and your text must be in the HTML rather than added later by JavaScript. GPTBot, the one most block lists mention, is only about training and doesn't decide whether you appear in ChatGPT search.

How ChatGPT gets to your pages

OpenAI documents three crawlers (OpenAI crawlers):

User agent Used for robots.txt
OAI-SearchBot Surfacing websites in ChatGPT's search features Respected. Opted-out sites "will not be shown in ChatGPT search answers, though can still appear as navigational links"
ChatGPT-User Visiting a page when a user asks ChatGPT or a custom GPT something "Because these actions are initiated by a user, robots.txt rules may not apply"
GPTBot Crawling content that may be used to train OpenAI's foundation models Respected. Disallowing it means content "should not be used in training"

OpenAI is explicit that the settings are independent: "a webmaster can allow OAI-SearchBot in order to appear in search results while disallowing GPTBot". The GPTBot vs OAI-SearchBot vs ChatGPT-User guide goes deeper into each.

The 5-minute check

1. robots.txt

Open https://yourdomain.com/robots.txt and look for:

  • a group naming OAI-SearchBot with Disallow: /;
  • a User-agent: * group with Disallow: /, which OAI-SearchBot inherits if it has no group of its own;
  • a long "block AI bots" list pasted from somewhere, which often includes search crawlers along with training ones.

If OAI-SearchBot is blocked, remove it from the list or give it its own group with Allow: /. The allow AI crawlers in robots.txt guide has a ready-made file.

2. Firewall and CDN

robots.txt is only read if the request gets through. OpenAI recommends "allowing requests from our published IP ranges", which it lists for each crawler: searchbot.json, chatgpt-user.json and gptbot.json.

Places that commonly block OpenAI's requests:

  • Cloudflare: AI bot policies, AI Crawl Control, Bot Fight Mode, custom WAF rules. New Cloudflare domains since 15 September 2026 block Agent bots on pages with ads by default, and Cloudflare counts chat fetch bots as Agents. See Cloudflare blocking AI bots.
  • Vercel, Netlify and similar hosts: optional AI bot rules in the firewall or extensions.
  • WordPress security plugins and server rules: user-agent blocklists, rate limits and country blocks. See firewall and plugin rules.

The quickest look is your server or CDN logs: search for OAI-SearchBot and ChatGPT-User and check the status codes. 200 is good; 403, 429 or 503 means something is turning them away.

3. Your HTML

Fetch a page and look for a sentence from the middle of the article:

curl -s -A "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; OAI-SearchBot/1.4; +https://openai.com/searchbot" https://example.com/your-article/ | grep -c "a sentence from your article"

If the count is 0 but the sentence shows in your browser, the text is added by JavaScript. OpenAI doesn't document whether its crawlers run JavaScript, so don't rely on it. See JavaScript content and AI crawlers.

This request comes from your computer with OpenAI's user agent string. It's a simulation: a firewall that verifies OpenAI's IP ranges will treat it differently from the real bot, in either direction.

4. Give it time

OpenAI says "it can take ~24 hours from a site's robots.txt update for our systems to adjust". Allowing OAI-SearchBot today doesn't mean you'll be cited tomorrow, and nothing guarantees ChatGPT will choose your page for a given question.

Ask ChatGPT directly, carefully

Asking ChatGPT "what's on example.com?" tells you something, but not much. A good answer can come from other sites that mention you, and a bad answer can happen even when nothing is blocked. Server logs and the checks above are more reliable than one conversation.

ChatGPT and AdSense

Being visible to ChatGPT has nothing to do with AdSense approval. AdSense uses its own crawler, Mediapartners-Google, and none of OpenAI's settings affect it. If you're working on both, keep the rules separate in robots.txt; the robots.txt and AdSense guide covers the AdSense side, and a free Approvalens scan checks AdSense readiness, with the full report covering up to 20,000 pages.

Run the check for free

The free AI crawler checker does steps 1 to 3 in one go for OpenAI's crawlers and the other major AI bots: it reads your robots.txt per bot, requests your page with each bot's user agent and flags pages whose text needs JavaScript. Like the curl test above, its requests are simulated from our server, so treat a block as a strong lead and confirm it in your firewall logs. The robots.txt tester checks the same file for Google's crawlers.

FAQ

I blocked GPTBot. Will ChatGPT stop showing my site?

No. GPTBot is training only. Search visibility is controlled by OAI-SearchBot.

Can I stop ChatGPT from reading a page a user pastes?

OpenAI says robots.txt rules "may not apply" to ChatGPT-User because the fetch is user-initiated. Blocking it in your firewall works for requests that reach it, but also means ChatGPT can't read the page for people who ask about you.

Does ChatGPT use Bing or Google to find my site?

OpenAI's crawler documentation describes OAI-SearchBot for its search features and doesn't describe any other search provider, so we don't make claims about one here.

How do I know if OpenAI's crawlers visit me?

Search your server or CDN logs for OAI-SearchBot, ChatGPT-User and GPTBot. To make sure a request really came from OpenAI and not from someone faking the name, compare its IP address with the ranges OpenAI publishes for that crawler.

Spotted something out of date or wrong? Tell us and we'll correct it.

Read this guide in Turkish →

Check it on your own site. Free, no sign-up.

Free tools for this

Free scan

Check your own site

Free scan: readiness score and every issue, usually in a few minutes.

Free scan · score and every problem found · no sign-up

All guides →