Approvalens

robots.txt and the AdSense Crawler: Rules That Matter

Approvalens · Updated October 4, 2026 · 5 min read

The AdSense crawler, Mediapartners-Google, only obeys robots.txt groups that name it directly; it ignores the User-agent: * group. So the one line that blocks AdSense is a Disallow: / under User-agent: Mediapartners-Google. That does not make * rules harmless: they still block Googlebot, which can stop ads.txt from being read and makes your content harder for Google to evaluate. The safest setup is an explicit open group for Mediapartners-Google and a light * group.

The crawlers that matter for AdSense

User agent What it does Reads the * group?
Mediapartners-Google Crawls pages to decide which ads to show No. Google: "The global user agent (*) is ignored." (special-case crawlers)
Google-Display-Ads-Bot Used with Mediapartners-Google when AdSense checks a newly added site (Add a new site) Do not block it
Googlebot Google Search crawler Yes
Google-Extended A robots.txt token that controls use of content for Gemini model training Not a separate crawler; does not affect AdSense

AdSense's help is direct about the consequence: "If you've modified your site's robots.txt file to disallow the AdSense crawler from indexing the pages of your site, then we can't serve Google ads on the site" (crawler access). And on the * group: "If you're serving ads on pages that are being roboted out with the line User-agent: *, then the AdSense crawler will still crawl these pages" (About the AdSense crawler).

How robots.txt groups are read

robots.txt is standardised in RFC 9309. Three rules explain most surprises:

  1. A crawler uses the most specific group that matches its name. If there is a User-agent: Googlebot group, Googlebot follows that group and ignores *. If there is no named group, it falls back to * (except Mediapartners-Google, which never falls back).
  2. Within a group, the longest matching path wins. Allow: /wp-admin/admin-ajax.php beats Disallow: /wp-admin/ for that one URL.
  3. Paths are prefixes. Disallow: /ads blocks /ads, /ads/, /ads-guide and /ads.txt.

A blank Disallow: means "nothing is disallowed", which is the cleanest way to say "this crawler may fetch everything".

Safe examples

A typical WordPress site

User-agent: Mediapartners-Google
Disallow:

User-agent: *
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php
Disallow: /?s=
Disallow: /search/

Sitemap: https://example.com/sitemap_index.xml

The Mediapartners-Google group is technically redundant here, since the bot would not read * anyway, but it documents your intent and protects you if someone later adds a broader block.

Blocking AI training crawlers without touching AdSense

User-agent: Mediapartners-Google
Disallow:

User-agent: GPTBot
Disallow: /

User-agent: CCBot
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: *
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php

Sitemap: https://example.com/sitemap.xml

Google-Extended only controls whether content is used for Gemini training. It does not block Googlebot or the AdSense crawler. Check any "block AI bots" plugin or CDN feature to make sure it only adds groups like these and does not add a blanket Disallow: / for everything.

A staging site that should never be crawled

Keep staging behind a password or IP allowlist rather than relying on robots.txt. Then make sure the staging robots.txt does not get deployed to production. A production site serving User-agent: * Disallow: / is one of the most common self-inflicted problems we see.

Mistakes that block AdSense or hurt the review

Mistake Effect Fix
User-agent: Mediapartners-Google + Disallow: / AdSense cannot crawl; no ads, "Site down or unavailable" during review Change to Disallow: (empty)
User-agent: * + Disallow: / left from staging Googlebot blocked; Google says it may disable ads on "content whose robots.txt file blocks Google's crawling" (ad placement policies) Remove the line
Disallow: /ads to hide an ads folder Also blocks /ads.txt Use Disallow: /ads/ with a trailing slash
Blocking CSS and JS folders Google cannot render pages properly Allow theme and plugin assets
robots.txt returns 5xx or times out Google may treat the whole site as off-limits until it can fetch the file Serve a static file that always returns 200
robots.txt is an HTML page Unreadable rules Serve plain text at /robots.txt
Sitemap line points to the wrong domain Crawlers ignore it or crawl a staging host Use the live, canonical sitemap URL

Why blocking Googlebot still matters for AdSense

The AdSense crawler is separate from Googlebot, and being indexed in Search is not an AdSense requirement. But Googlebot-related blocks still cause trouble:

  • ads.txt. "The ads.txt file for a domain may be ignored by crawlers if the robots.txt file on a domain disallows … the crawling of the URL path on which an ads.txt file is posted" (ads.txt crawling). See the ads.txt guide.
  • Content evaluation. Google may disable ads on content it cannot evaluate, including content behind a robots.txt block.
  • Noindex everywhere. A sitewide noindex meta tag or X-Robots-Tag: noindex header is not robots.txt, but it sends the same "do not look" signal. Remove it from the pages you want reviewed.

Test your robots.txt

curl -s https://example.com/robots.txt
curl -sI https://example.com/robots.txt

Check the status (200), the content type (text/plain) and the groups. In Search Console, the robots.txt report shows which version Google fetched and any parse errors. If a firewall challenges bots, robots.txt itself can be unreachable for Google even when it loads for you; the Cloudflare guide covers that case, and site down or unavailable covers the broader reachability checks. Our methodology describes how we evaluate each group the way the crawler would.

Check your robots.txt for free

A free scan parses your robots.txt per RFC 9309, evaluates it for Mediapartners-Google, Googlebot and ads.txt, and flags anything that blocks them. Run a free scan.

FAQ

Does User-agent: * with Disallow: / block AdSense?

Not the AdSense crawler, which ignores the * group. It does block Googlebot, which can affect ads.txt crawling and content evaluation, so do not leave it on a live site.

Do I need a Mediapartners-Google group at all?

No, if nothing blocks it. Adding User-agent: Mediapartners-Google with an empty Disallow: is a harmless way to make your intent explicit.

Will blocking Google-Extended hurt my AdSense earnings?

Google-Extended is a separate token for Gemini training use. It does not control Googlebot or Mediapartners-Google.

How fast does Google pick up robots.txt changes?

Google generally caches robots.txt for up to about a day. Changes are usually seen within that window.

Check your own site

Free scan: readiness score and every issue, usually in a few minutes.

Free scan · score and top 3 issues · no sign-up

More guides