On this page (12)
- §1The rules, from RFC 9309
- §2Surprise 1: a named group cancels your * rules
- §3Surprise 2: * blocks every AI bot you didn't name
- §4Surprise 3: duplicate groups merge
- §5Surprise 4: some bots borrow another bot's rules
- §6Surprise 5: blocking the sitemap path
- §7Surprise 6: spelling and paths
- §8Surprise 7: one file per host
- §9A quick way to read any file
- §10Precedence and AdSense
- §11Test it
- §12FAQ
A crawler obeys only the robots.txt group that names it. If there's a User-agent: GPTBot group, GPTBot follows that group and ignores everything under User-agent: *; it doesn't add the two together. Only when no group names a bot does it fall back to *. Groups that name the same bot are merged, order in the file doesn't matter, and inside a group the longest matching path wins. Most "why is this AI bot allowed (or blocked)?" puzzles come from forgetting the first rule.
The rules, from RFC 9309
robots.txt is standardised as RFC 9309. For choosing a group:
- Match the product token, case-insensitively. Crawlers "MUST use case-insensitive matching to find the group that matches the product token and then obey the rules of the group."
- Merge duplicates. "If there is more than one group matching the user-agent, the matching groups' rules MUST be combined into one group."
- Fall back to
*only if nothing matches. "If no matching group exists, crawlers MUST obey the group with a user-agent line with the '*' value, if present." - No group, no rules. If nothing matches and there's no
*group, "no rules apply".
For the paths inside the chosen group:
- Longest match wins. "The most specific match is the match that has the most octets."
- Ties go to Allow. "If an 'allow' rule and a 'disallow' rule are equivalent, then the 'allow' rule SHOULD be used."
- No match means allowed. And "the /robots.txt URI is implicitly allowed".
A group can start with several User-agent: lines that share the same rules.
Surprise 1: a named group cancels your * rules
User-agent: *
Disallow: /wp-admin/
Disallow: /checkout/
Disallow: /search/
User-agent: GPTBot
Crawl-delay: 10
You meant "everyone stays out of admin, checkout and search; GPTBot also slows down". What GPTBot actually sees is a group with no Disallow lines at all, so it may fetch everything, including /checkout/. Fix it by repeating the paths:
User-agent: GPTBot
Crawl-delay: 10
Disallow: /wp-admin/
Disallow: /checkout/
Disallow: /search/
(Crawl-delay isn't in RFC 9309. OpenAI's crawler page doesn't mention it; Anthropic says it supports it for its bots.)
Surprise 2: * blocks every AI bot you didn't name
User-agent: *
Disallow: /
User-agent: Googlebot
Allow: /
This lets Googlebot in and blocks OAI-SearchBot, Claude-SearchBot, PerplexityBot and every other crawler without its own group. If you want AI search bots in, name them:
User-agent: Googlebot
User-agent: OAI-SearchBot
User-agent: Claude-SearchBot
User-agent: PerplexityBot
Allow: /
User-agent: *
Disallow: /
Surprise 3: duplicate groups merge
User-agent: ClaudeBot
Disallow: /drafts/
User-agent: ClaudeBot
Disallow: /members/
ClaudeBot follows both: /drafts/ and /members/ are blocked. This matters when a plugin or CDN adds its own block on top of yours. Cloudflare's managed robots.txt, for example, prepends a User-agent: * group and named groups for AI training bots to your existing file (managed robots.txt). Your own * group and Cloudflare's are combined into one.
Surprise 4: some bots borrow another bot's rules
Not every crawler stops at *. Two documented exceptions:
- Applebot: "If robots instructions don't mention Applebot but mention Googlebot, the Apple robot will follow Googlebot instructions" (About Applebot).
- Amzn-SearchBot: "If robots.txt files don't mention Amzn-SearchBot but allow other search bots, Amzn-SearchBot will crawl in accordance with the robots.txt directives given to other search bots" (Amazonbot).
And one that never falls back: Google's AdSense crawler, Mediapartners-Google, ignores the * group entirely (robots.txt and AdSense).
Surprise 5: blocking the sitemap path
User-agent: *
Disallow: /*.xml
Meant to hide feeds, this also blocks /sitemap.xml for every bot that follows *, so search crawlers can't use it to discover your URLs. (Wildcards like * in paths are part of RFC 9309.) Give your sitemap an explicit Allow, which wins because it's longer:
User-agent: *
Allow: /sitemap.xml
Disallow: /*.xml
Sitemap: https://example.com/sitemap.xml
The Sitemap: line isn't tied to any group and applies to the whole file.
Surprise 6: spelling and paths
User-agent: GPT-BotorUser-agent: Claude Search Botmatch nothing. Copy tokens from the operator's own page; the allow AI crawlers in robots.txt guide lists them with sources.- Paths are case-sensitive:
Disallow: /Blog/doesn't block/blog/. - Paths are prefixes:
Disallow: /aiblocks/ai-guidetoo.
Surprise 7: one file per host
robots.txt applies to the host it's served from. blog.example.com needs its own file. Anthropic, for instance, asks you to add opt-out rules "for every subdomain that you wish to opt out from", and Amazon says its bots honour "robots rules exposed under each host".
A quick way to read any file
For each bot you care about, ask:
- Is there a group with its exact token? If yes, use only that group (merged with any other groups naming it).
- If not, does the bot document a fallback (Applebot → Googlebot)? Use that.
- Otherwise use
*. No*group means everything is allowed. - Within the chosen group, find the longest matching path for the URL.
Precedence and AdSense
The same rules decide what Googlebot and Mediapartners-Google can fetch, so a mistake made for AI bots can spill over. After any change, check the Google side with the robots.txt tester or a free scan.
Test it
The free AI crawler checker applies these rules for each major AI bot and shows which group it ended up in and whether your homepage and sitemap are allowed. It then requests the page with simulated user agents from our server; a firewall that checks IP ranges may treat the real bots differently, so a pass there is a hint, not a guarantee.
FAQ
Does the order of groups in robots.txt matter?
No. A crawler picks the group by matching its token, not by position, and merges groups that name it.
If I name GPTBot, do I need to repeat my * rules?
Yes, any rule you want GPTBot to follow must be in its group.
Can one group list several bots?
Yes. Several User-agent: lines in a row share the rules that follow.
Does Allow work everywhere?
It's part of RFC 9309. Crawlers that follow the standard support it.
Spotted something out of date or wrong? Tell us and we'll correct it.
Read this guide in Turkish →Check it on your own site. Free, no sign-up.
Free tools for this
Free scan
Check your own site
Free scan: readiness score and every issue, usually in a few minutes.