Skip to content
Approvalens

Reading room · 5 min read

robots.txt Group Precedence: Named AI Bots vs User-agent: *

Why a GPTBot group makes GPTBot ignore your * rules, how duplicate groups merge, Applebot's Googlebot fallback and other robots.txt surprises for AI bots.

By the Approvalens team

Fixes these report findings

  • Sitemap pages open to AI crawlers
  • AI search crawlers allowed in robots.txt
  • AI training crawlers in robots.txt

A crawler obeys only the robots.txt group that names it. If there's a User-agent: GPTBot group, GPTBot follows that group and ignores everything under User-agent: *; it doesn't add the two together. Only when no group names a bot does it fall back to *. Groups that name the same bot are merged, order in the file doesn't matter, and inside a group the longest matching path wins. Most "why is this AI bot allowed (or blocked)?" puzzles come from forgetting the first rule.

The rules, from RFC 9309

robots.txt is standardised as RFC 9309. For choosing a group:

  1. Match the product token, case-insensitively. Crawlers "MUST use case-insensitive matching to find the group that matches the product token and then obey the rules of the group."
  2. Merge duplicates. "If there is more than one group matching the user-agent, the matching groups' rules MUST be combined into one group."
  3. Fall back to * only if nothing matches. "If no matching group exists, crawlers MUST obey the group with a user-agent line with the '*' value, if present."
  4. No group, no rules. If nothing matches and there's no * group, "no rules apply".

For the paths inside the chosen group:

  1. Longest match wins. "The most specific match is the match that has the most octets."
  2. Ties go to Allow. "If an 'allow' rule and a 'disallow' rule are equivalent, then the 'allow' rule SHOULD be used."
  3. No match means allowed. And "the /robots.txt URI is implicitly allowed".

A group can start with several User-agent: lines that share the same rules.

Surprise 1: a named group cancels your * rules

User-agent: *
Disallow: /wp-admin/
Disallow: /checkout/
Disallow: /search/

User-agent: GPTBot
Crawl-delay: 10

You meant "everyone stays out of admin, checkout and search; GPTBot also slows down". What GPTBot actually sees is a group with no Disallow lines at all, so it may fetch everything, including /checkout/. Fix it by repeating the paths:

User-agent: GPTBot
Crawl-delay: 10
Disallow: /wp-admin/
Disallow: /checkout/
Disallow: /search/

(Crawl-delay isn't in RFC 9309. OpenAI's crawler page doesn't mention it; Anthropic says it supports it for its bots.)

Surprise 2: * blocks every AI bot you didn't name

User-agent: *
Disallow: /

User-agent: Googlebot
Allow: /

This lets Googlebot in and blocks OAI-SearchBot, Claude-SearchBot, PerplexityBot and every other crawler without its own group. If you want AI search bots in, name them:

User-agent: Googlebot
User-agent: OAI-SearchBot
User-agent: Claude-SearchBot
User-agent: PerplexityBot
Allow: /

User-agent: *
Disallow: /

Surprise 3: duplicate groups merge

User-agent: ClaudeBot
Disallow: /drafts/

User-agent: ClaudeBot
Disallow: /members/

ClaudeBot follows both: /drafts/ and /members/ are blocked. This matters when a plugin or CDN adds its own block on top of yours. Cloudflare's managed robots.txt, for example, prepends a User-agent: * group and named groups for AI training bots to your existing file (managed robots.txt). Your own * group and Cloudflare's are combined into one.

Surprise 4: some bots borrow another bot's rules

Not every crawler stops at *. Two documented exceptions:

  • Applebot: "If robots instructions don't mention Applebot but mention Googlebot, the Apple robot will follow Googlebot instructions" (About Applebot).
  • Amzn-SearchBot: "If robots.txt files don't mention Amzn-SearchBot but allow other search bots, Amzn-SearchBot will crawl in accordance with the robots.txt directives given to other search bots" (Amazonbot).

And one that never falls back: Google's AdSense crawler, Mediapartners-Google, ignores the * group entirely (robots.txt and AdSense).

Surprise 5: blocking the sitemap path

User-agent: *
Disallow: /*.xml

Meant to hide feeds, this also blocks /sitemap.xml for every bot that follows *, so search crawlers can't use it to discover your URLs. (Wildcards like * in paths are part of RFC 9309.) Give your sitemap an explicit Allow, which wins because it's longer:

User-agent: *
Allow: /sitemap.xml
Disallow: /*.xml

Sitemap: https://example.com/sitemap.xml

The Sitemap: line isn't tied to any group and applies to the whole file.

Surprise 6: spelling and paths

  • User-agent: GPT-Bot or User-agent: Claude Search Bot match nothing. Copy tokens from the operator's own page; the allow AI crawlers in robots.txt guide lists them with sources.
  • Paths are case-sensitive: Disallow: /Blog/ doesn't block /blog/.
  • Paths are prefixes: Disallow: /ai blocks /ai-guide too.

Surprise 7: one file per host

robots.txt applies to the host it's served from. blog.example.com needs its own file. Anthropic, for instance, asks you to add opt-out rules "for every subdomain that you wish to opt out from", and Amazon says its bots honour "robots rules exposed under each host".

A quick way to read any file

For each bot you care about, ask:

  1. Is there a group with its exact token? If yes, use only that group (merged with any other groups naming it).
  2. If not, does the bot document a fallback (Applebot → Googlebot)? Use that.
  3. Otherwise use *. No * group means everything is allowed.
  4. Within the chosen group, find the longest matching path for the URL.

Precedence and AdSense

The same rules decide what Googlebot and Mediapartners-Google can fetch, so a mistake made for AI bots can spill over. After any change, check the Google side with the robots.txt tester or a free scan.

Test it

The free AI crawler checker applies these rules for each major AI bot and shows which group it ended up in and whether your homepage and sitemap are allowed. It then requests the page with simulated user agents from our server; a firewall that checks IP ranges may treat the real bots differently, so a pass there is a hint, not a guarantee.

FAQ

Does the order of groups in robots.txt matter?

No. A crawler picks the group by matching its token, not by position, and merges groups that name it.

If I name GPTBot, do I need to repeat my * rules?

Yes, any rule you want GPTBot to follow must be in its group.

Can one group list several bots?

Yes. Several User-agent: lines in a row share the rules that follow.

Does Allow work everywhere?

It's part of RFC 9309. Crawlers that follow the standard support it.

Spotted something out of date or wrong? Tell us and we'll correct it.

Read this guide in Turkish →

Check it on your own site. Free, no sign-up.

Free tools for this

Free scan

Check your own site

Free scan: readiness score and every issue, usually in a few minutes.

Free scan · score and every problem found · no sign-up

All guides →