OpenAI's bots do different jobs. GPTBot collects content that may be used to train OpenAI's models. OAI-SearchBot decides whether your site can appear in ChatGPT search answers. ChatGPT-User fetches a page when a ChatGPT user asks for it, and robots.txt "may not apply" to it. OAI-AdsBot only visits pages submitted as ads on ChatGPT. Most content sites that want ChatGPT traffic but not training should allow OAI-SearchBot and disallow GPTBot. OpenAI says the settings are independent.
The four OpenAI user agents
Everything in this table comes from OpenAI's crawler documentation.
| GPTBot | OAI-SearchBot | ChatGPT-User | OAI-AdsBot | |
|---|---|---|---|---|
| Purpose | Crawl content "that may be used in training our generative AI foundation models" | "Surface websites in search results in ChatGPT's search features" | Visit a page when "users ask ChatGPT or a CustomGPT a question"; also GPT Actions | "Validate the safety of web pages submitted as ads on ChatGPT" |
| Crawls automatically? | Yes | Yes | No: "not used for crawling the web in an automatic fashion" | Only pages submitted as ads |
| robots.txt | Disallowing it "indicates a site's content should not be used in training" | Opted-out sites "will not be shown in ChatGPT search answers, though can still appear as navigational links" | "robots.txt rules may not apply" | Not stated |
| Used for training? | Yes, that's its job | Not stated as a training crawler | Not stated | "Not used to train generative AI foundation models" |
| IP ranges | gptbot.json | searchbot.json | chatgpt-user.json | adsbot.json |
Two more details from the same page:
- Shared crawls. "If your site has allowed both bots, we may use the results from just one crawl for both use cases to avoid duplicative crawling." Seeing only one of them in your logs doesn't mean the other is blocked.
- ChatGPT-User doesn't control search. It "is not used to determine whether content may appear in Search". Use OAI-SearchBot for that.
What the user agent strings look like
OpenAI gives example strings and notes that version numbers change:
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.4; +https://openai.com/gptbot
Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36; compatible; OAI-SearchBot/1.4; +https://openai.com/searchbot
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; ChatGPT-User/1.0; +https://openai.com/bot
When fetching robots.txt, GPTBot and OAI-SearchBot may add a robots.txt marker to the string. In robots.txt you only write the token (GPTBot, OAI-SearchBot), never the whole string. Filters that match on the full string break whenever OpenAI bumps a version number; match on the token, or better, on the published IP ranges.
Three setups
Search yes, training no
The most common choice for publishers.
User-agent: OAI-SearchBot
Allow: /
User-agent: GPTBot
Disallow: /
Everything allowed
No OpenAI groups needed, as long as your User-agent: * group isn't restrictive. Add explicit groups if you want to document the choice:
User-agent: OAI-SearchBot
User-agent: GPTBot
Allow: /
Nothing from OpenAI
User-agent: OAI-SearchBot
User-agent: GPTBot
Disallow: /
This removes you from ChatGPT search answers (navigational links can still appear) and opts out of training. It doesn't reliably stop ChatGPT-User, since robots.txt may not apply to user-initiated fetches. If you need that, block ChatGPT-User's IP ranges in your firewall, knowing ChatGPT then can't read your pages for anyone who asks about them.
Which one should you allow?
| If you want… | OAI-SearchBot | GPTBot |
|---|---|---|
| ChatGPT search traffic, no training | Allow | Disallow |
| Maximum reach in ChatGPT | Allow | Allow |
| Nothing used by OpenAI | Disallow | Disallow |
OpenAI doesn't say that allowing GPTBot improves your ChatGPT search visibility, so there's no documented visibility reason to allow training. The decision is about whether you're comfortable with your content in training data. The should I block AI training bots guide lays out the trade-offs.
Firewalls and verification
Anyone can send a request that says GPTBot. To allow the real crawlers through a firewall, match OpenAI's published IP ranges rather than the user agent string. To block fakes, do the reverse: drop requests that claim an OpenAI user agent from outside those ranges. If you use Cloudflare, check its AI settings first; the Cloudflare blocking AI bots guide lists them.
OpenAI bots and AdSense
None of these is related to Google AdSense. OAI-AdsBot is about ads on ChatGPT, not ads on your site. AdSense crawls with Mediapartners-Google, and the robots.txt and AdSense guide covers it.
Check which OpenAI bots you allow
The free AI crawler checker shows, bot by bot, whether your robots.txt allows GPTBot, OAI-SearchBot and ChatGPT-User, and requests your page with each user agent to catch firewall blocks. These requests are simulated: they use OpenAI's user agent strings but come from our server, not OpenAI's IP ranges, so a firewall that verifies IPs may answer them differently. For a line-by-line read of your robots.txt, use the robots.txt tester.
FAQ
Is GPTBot the same as ChatGPT?
No. GPTBot crawls for training. ChatGPT's search uses OAI-SearchBot, and live page visits on a user's behalf use ChatGPT-User.
Can I allow GPTBot on some folders only?
Yes. robots.txt paths work per group, for example User-agent: GPTBot, Allow: /blog/, Disallow: /. The longest matching path wins.
How quickly do changes apply?
For search, OpenAI says it "can take ~24 hours" after a robots.txt update.
My robots.txt names a different OpenAI bot. Does it matter?
OpenAI's current documentation lists the four agents above. A group for a name that isn't on that list does nothing, so check the spelling against OpenAI's page rather than a copied list.
Spotted something out of date or wrong? Tell us and we'll correct it.
Read this guide in Turkish →Check it on your own site. Free, no sign-up.
Free tools for this
Free scan
Check your own site
Free scan: readiness score and every issue, usually in a few minutes.