Approvalens

robots.txt Tester for the AdSense Crawler and Googlebot

Check whether Mediapartners-Google, Googlebot, AdsBot-Google and Googlebot-Image may fetch any path on your site, using the matching rules Google documents. The verdict for each crawler appears above.

What this robots.txt tester checks

Google runs several crawlers, and each one reads robots.txt on its own. The tester fetches your live /robots.txt, picks the group that applies to each crawler and tells you whether the path you entered is allowed or blocked, and which line decided it.

  • Mediapartners-Google: the AdSense crawler, which reads your pages to choose relevant ads
  • Googlebot: Google Search's main crawler. Google also asks that robots.txt does not block your ads.txt
  • AdsBot-Google: checks landing page quality for Google Ads
  • Googlebot-Image: crawls images for Google Images

How Google decides: the matching rules

Google follows the robots.txt standard (RFC 9309) plus a few documented details. The tester applies them in this order:

  • One group per crawler. The group with the most specific matching user-agent wins. Googlebot-Image uses its own group if there is one, then the Googlebot group, then *. Rules in other groups are ignored.
  • Mediapartners-Google and AdsBot-Google ignore the * group. Only a group that names them can block them.
  • The longest match wins. Of the rules that match the URL, the one with the longest path decides.
  • Ties go to Allow. If an Allow and a Disallow rule are the same length, the less restrictive Allow wins.
  • * matches any run of characters and $ marks the end of the URL. Paths are case-sensitive.
  • A missing robots.txt (404) means everything is allowed. A 5xx server error makes Google treat the whole site as blocked for a while.

Common robots.txt mistakes that hurt AdSense

Each of these turns up regularly on sites that get rejected or lose ad coverage.

  • Blocking the AdSense crawler by name. User-agent: Mediapartners-Google followed by Disallow: / stops Google from serving ads on your pages. Remove the group, or leave the rule empty: Disallow: with nothing after the colon.
  • Blocking everything for everyone. User-agent: * with Disallow: / does not stop Mediapartners-Google, but it blocks Googlebot, can make Google ignore your ads.txt and stops your content from being evaluated. Staging sites often go live with this line still in place.
  • Prefix rules that catch too much. Disallow: /ads blocks /ads.txt and /ads-guide as well. Write Disallow: /ads/ if you meant a folder.
  • Blocking CSS, JavaScript or images. Googlebot needs them to render the page. Disallow: /wp-content/ breaks your theme for crawlers and hides your images.
  • Wrong location. robots.txt only works at the root of each host: https://example.com/robots.txt. A file in a subfolder is ignored, and shop.example.com needs its own file.

Platform tips: WordPress, Blogger and CDNs

WordPress serves a virtual robots.txt unless a real file exists in the site root, and a real file always wins. SEO plugins such as Yoast SEO and Rank Math include an editor for it. Also open Settings, Reading: the Search engine visibility box asks search engines not to index your site, so untick it before you apply.

Blogger's default robots.txt already contains a Mediapartners-Google group with an empty Disallow, which lets the AdSense crawler in. If you turned on Custom robots.txt under Settings, Crawlers and indexing, make sure that group is still there.

Cloudflare and other CDNs can add their own lines to robots.txt, for example for AI crawlers. Test the live file, not the copy in your theme or repository.

robots.txt and AdSense approval

AdSense help is direct about it: if robots.txt disallows the AdSense crawler, Google can't serve ads on those pages. Google also says it may disable ads on content it cannot evaluate, including content blocked by robots.txt. Blocking Googlebot is not an automatic rejection, because the AdSense crawler is separate from Search, but it makes your site harder to review.

Google caches robots.txt for up to 24 hours, so a fix can take a day to register. robots.txt is also only one layer: a firewall can still turn crawlers away. Our Googlebot access checker tests that layer, and the full Approvalens scan runs both checks with more than 100 others. Sources: https://support.google.com/adsense/answer/10532 and https://developers.google.com/search/docs/crawling-indexing/google-special-case-crawlers

FAQ

Does User-agent: * Disallow: / block AdSense?

Not the AdSense crawler itself. Google documents that Mediapartners-Google ignores the * group. It does block Googlebot, which affects ads.txt crawling and how Google evaluates your content, so it can still cause problems.

How do I allow the AdSense crawler in robots.txt?

Add a group that names it with an empty rule: User-agent: Mediapartners-Google on one line and Disallow: on the next. An empty Disallow allows everything.

Is the Search Console robots.txt tester gone?

Yes. Google retired the old tester in 2023 and replaced it with a robots.txt report that shows which versions of the file Google fetched. The report does not test single paths per crawler, which is what this tool does.

What is the difference between Mediapartners-Google and Googlebot?

Mediapartners-Google reads pages to decide which ads to show, and Googlebot crawls for Google Search. AdSense says its crawler is different from the Google crawler, and your site does not have to be indexed in Search to be reviewed.

Why is my path blocked when I have an Allow rule?

The longest matching rule wins, so Allow: /blog loses to Disallow: /blog/drafts for every URL under /blog/drafts. Make the Allow rule more specific, and check that it sits in the group that actually applies to that crawler.

Free AdSense readiness tools