Skip to content
Approvalens

Free · no sign-up

robots.txt Generator and Tester

Pick your platform, choose which AI crawlers to keep out and add your sitemap. The file is written as you go, and the Test tab shows, for 16 crawlers, whether a path is allowed and which line decided it. Everything runs in your browser.

Build or test a robots.txt

What do you want to do?

Your platform

Your platform

One per line, starting with /. Search results and admin pages are typical.

One per line. WordPress needs /wp-admin/admin-ajax.php open.

Block AI training crawlers

They collect text to train models. Blocking them doesn't affect Google Search or AdSense.

Block AI search and assistant agents

They fetch pages to answer questions in ChatGPT, Claude and Perplexity, often with a link to you.

Full address, one per line, e.g. https://example.com/sitemap.xml.

robots.txt
# AdSense crawler: everything open (it ignores the * group anyway)
User-agent: Mediapartners-Google
Disallow:

User-agent: *
Disallow: /wp-admin/
Disallow: /?s=
Disallow: /search/
Allow: /wp-admin/admin-ajax.php

Save it as robots.txt at the root of your site.

What the generator writes

A short file with up to three groups: one that names Mediapartners-Google with an empty Disallow, so the AdSense crawler is plainly open; one for the AI crawlers you chose to keep out; and the * group for everyone else, with the paths you don't want crawled and your sitemap.

Google's own documentation says Mediapartners-Google ignores the * group, so the first group isn't strictly needed. It's there because it makes your intent readable to anyone who opens the file later, including you.

How the tester decides

It reads the file the way Google documents it: a crawler obeys only the most specific group that names it; inside that group the longest matching rule wins; on a tie, Allow wins; * matches any run of characters and $ marks the end of the URL. Mediapartners-Google and AdsBot-Google skip the * group. AI crawlers are matched by their exact token, so a group for gpt doesn't apply to GPTBot.

For each crawler you see the verdict and the line that decided it. That's usually enough to spot the classic mistake: an Allow rule sitting in a group the crawler never reads.

Blocking AI crawlers: training vs search

Training crawlers (GPTBot, ClaudeBot, CCBot and the Google-Extended and Applebot-Extended tokens) collect text to train models. Search and assistant agents (OAI-SearchBot, ChatGPT-User, Claude-SearchBot, PerplexityBot) fetch pages to answer a question and often link back to you. Blocking the second group removes your site from those answers.

Google-Extended doesn't affect Google Search or AdSense. It's a separate token that controls whether content is used for Gemini models; blocking it leaves Googlebot and the AdSense crawler alone.

Common robots.txt mistakes on ad-supported sites

  • Disallow: / under User-agent: * on a live site, often left over from staging
  • Blocking /wp-content/ or theme folders, which hides CSS and JavaScript Google needs to render pages
  • A sitemap line with a relative path instead of a full URL
  • Rules without a leading slash, which don't match the paths you think
  • Several groups for the same crawler written as if only the last one counts (they merge)

Questions

Frequently asked questions

Will this robots.txt get my site approved for AdSense?

No file can do that. What it can do is make sure the AdSense crawler and Googlebot aren't shut out, which is one of the reasons Google lists when a site isn't ready to show ads.

Should I block AI crawlers?

It depends on what you want from them. Blocking training crawlers keeps your text out of future training sets that honour robots.txt. Blocking search agents also removes you from AI answers that could send you visitors. Our guide on blocking AI training bots walks through the trade-off.

Does robots.txt hide a page from Google?

It stops crawling, not indexing. A blocked URL can still appear in search results without a description if other pages link to it. Use a noindex tag on a crawlable page to keep it out of the index.

Where does the file go?

At the root of the host: https://example.com/robots.txt. Each subdomain and each protocol-host combination has its own file. On Blogger, use Settings, Crawlers and indexing, Custom robots.txt.

Why test it here instead of in Search Console?

Search Console's robots.txt report shows which versions of the file Google fetched and any parse errors. It doesn't test one path for AdsBot, Mediapartners-Google or AI crawlers, which is what the Test tab does.

Guides

Guides that use this tool

Where this check comes up, and how to fix what it finds.

Next step

One check here, the whole site in the report

This is one of 258 checks. Scan the whole site for the full picture; the first 50 pages are free.

First 50 pages free · no sign-up · no card