What the generator writes
A short file with up to three groups: one that names Mediapartners-Google with an empty Disallow, so the AdSense crawler is plainly open; one for the AI crawlers you chose to keep out; and the * group for everyone else, with the paths you don't want crawled and your sitemap.
Google's own documentation says Mediapartners-Google ignores the * group, so the first group isn't strictly needed. It's there because it makes your intent readable to anyone who opens the file later, including you.
How the tester decides
It reads the file the way Google documents it: a crawler obeys only the most specific group that names it; inside that group the longest matching rule wins; on a tie, Allow wins; * matches any run of characters and $ marks the end of the URL. Mediapartners-Google and AdsBot-Google skip the * group. AI crawlers are matched by their exact token, so a group for gpt doesn't apply to GPTBot.
For each crawler you see the verdict and the line that decided it. That's usually enough to spot the classic mistake: an Allow rule sitting in a group the crawler never reads.
Blocking AI crawlers: training vs search
Training crawlers (GPTBot, ClaudeBot, CCBot and the Google-Extended and Applebot-Extended tokens) collect text to train models. Search and assistant agents (OAI-SearchBot, ChatGPT-User, Claude-SearchBot, PerplexityBot) fetch pages to answer a question and often link back to you. Blocking the second group removes your site from those answers.
Google-Extended doesn't affect Google Search or AdSense. It's a separate token that controls whether content is used for Gemini models; blocking it leaves Googlebot and the AdSense crawler alone.
Common robots.txt mistakes on ad-supported sites
- Disallow: / under User-agent: * on a live site, often left over from staging
- Blocking /wp-content/ or theme folders, which hides CSS and JavaScript Google needs to render pages
- A sitemap line with a relative path instead of a full URL
- Rules without a leading slash, which don't match the paths you think
- Several groups for the same crawler written as if only the last one counts (they merge)