Skip to content
Approvalens

Reading room · 5 min read

AI Bots Blocked by Your Host, Firewall or Plugin? Where to Look

robots.txt allows AI crawlers but they still get 403s? Where to find blocks in WordPress plugins, nginx, Apache, Vercel and Netlify, and how to fix them.

By the Approvalens team

Fixes these report findings

  • AI crawlers get your pages from the server

When AI crawlers get a 403, 429 or challenge page even though your robots.txt allows them, the block is in front of your content: a CDN or hosting firewall, a server rule, or a WordPress plugin. robots.txt only expresses a preference; these layers decide whether the request is answered at all. Check them in order from the outside in (CDN, host firewall, web server, CMS plugins), look for user-agent lists that name AI bots, and allow the bots you want by their published IP ranges rather than by user agent text.

How to tell it's a server-side block

Symptom Likely layer
robots.txt allows the bot, but logs show 403 for its requests Firewall, server rule or plugin
Requests never appear in your server logs CDN or host firewall in front of the server
A "Just a moment", "Checking your browser" or similar page Bot challenge (CDN or host)
429 Too Many Requests Rate limiting
200, but with a short block message in the body Plugin or edge function returning a custom page

Start with your CDN's and host's logs or security events, then the web server's access log, then the CMS.

Cloudflare

Cloudflare has its own AI bot policies, AI Crawl Control and managed robots.txt, plus Bot Fight Mode and WAF rules. It's common enough to have its own guide: Cloudflare blocking AI bots.

Vercel

Vercel's firewall has an AI Bots Managed Ruleset. Per Vercel's docs:

  • It's available on all plans and "inactive by default", shown as Allow in the dashboard.
  • You set it under the project's Firewall > Rules > Bot Management, choosing Log (record only) or Deny ("blocks all traffic identified as coming from AI bots").
  • The separate Bot Protection Managed Ruleset can serve "a JavaScript challenge to traffic that is unlikely to be a browser" when set to Challenge.

If AI crawlers are blocked on a Vercel site, check whether either ruleset is set to Deny or Challenge. To let specific traffic through a managed ruleset, Vercel documents a bypass action in a WAF custom rule.

Netlify

Netlify offers a User Agent Blocker extension that "can block web requests from a preset list of common AI crawlers, SEO/Search crawlers that you choose from", using an Edge Function (Netlify docs). It's set per project under Extensions > User Agent Blocker > Block User Agents.

Two things to check: which options are ticked (the list includes search crawlers, not only AI training bots), and whether the extension is active on the project in question.

WordPress plugins

SEO and security plugins can add AI bot rules to robots.txt or block requests outright.

  • Yoast SEO Premium: Settings > Advanced > Crawl optimization > Block unwanted bots has toggles for Google AdsBot, Google-Extended, GPTBot and CCBot. These write robots.txt rules (Yoast). Yoast's own help warns that blocking Google AdsBot "can have serious negative consequences if you are running Google Ads".
  • The SEO Framework: SEO Settings > Robots Settings > Robots.txt has curated "AI" and "SEO" blocklists. The AI list includes Amazonbot, Applebot-Extended, CCBot, ClaudeBot, GPTBot, Google-Extended, GoogleOther, Meta-ExternalAgent and FacebookBot (TSF). It doesn't include OAI-SearchBot or PerplexityBot.
  • Security and firewall plugins: many have user-agent blocklists, "bad bot" lists, country blocking and rate limiting. Look for any list containing GPTBot, ClaudeBot, PerplexityBot, OAI-SearchBot, Bytespider, CCBot or a pattern like bot|crawl|spider.

On a hosted platform, check the platform's own setting too, for example Squarespace's "Block known artificial intelligence crawlers" box (Squarespace).

nginx

Search the configuration for user-agent rules:

grep -rniE "http_user_agent|gptbot|claudebot|perplexity|oai-searchbot|ccbot|bytespider" /etc/nginx/

A typical block looks like this:

if ($http_user_agent ~* "(GPTBot|ClaudeBot|PerplexityBot|CCBot)") {
    return 403;
}

Remove the bots you want to allow from the pattern, or the whole block if robots.txt already says what you want. Also check limit_req zones (rate limits answer with 503 by default, or the status set by limit_req_status) and any deny rules for IP ranges. Reload nginx after editing (nginx -t first).

Apache and .htaccess

Search the virtual host files and every .htaccess:

grep -rniE "HTTP_USER_AGENT|BrowserMatch|SetEnvIf|gptbot|claudebot|perplexity" /etc/apache2/ /var/www/

Typical patterns:

RewriteCond %{HTTP_USER_AGENT} (GPTBot|ClaudeBot|PerplexityBot) [NC]
RewriteRule .* - [F,L]

or

SetEnvIfNoCase User-Agent "GPTBot" bad_bot
Require not env bad_bot

[F] returns 403 Forbidden. Edit the pattern to drop the bots you want to allow. If a plugin wrote the rule into .htaccess, it may come back while the plugin setting is still on, so change it in the plugin too.

Allow by IP, not by user agent

Matching on the user agent alone lets anyone in who types GPTBot into a request. The operators publish their IP ranges for this purpose:

Operator Ranges
OpenAI gptbot.json, searchbot.json, chatgpt-user.json (docs)
Anthropic bots.json (docs)
Perplexity perplexitybot.json, perplexity-user.json (docs)
Apple applebot.json, or reverse DNS in *.applebot.apple.com (docs)
Common Crawl ccbot.json, reverse DNS in crawl.commoncrawl.org (docs)
Amazon Amazonbot, Amzn-SearchBot, Amzn-User (docs)

Perplexity recommends refreshing these automatically, since the ranges change.

One caution from Anthropic: blocking its IPs to opt out "may not work correctly", because the bot then can't read your robots.txt to see the opt-out. Use robots.txt to say no, and firewall rules for abuse.

Don't break Google while you're at it

A broad pattern like bot|crawler|spider also catches Googlebot and the AdSense crawler. If the same rule is blocking Google, that's an AdSense problem too; the site down or unavailable guide covers what that looks like in a review. Check with the Googlebot access checker after any change.

Check from outside

The free AI crawler checker requests your page with each major AI bot's user agent and reports the status code and any challenge page, next to what your robots.txt says, so a mismatch ("robots.txt allows, server says 403") stands out. These requests are simulated: they come from our server, not the operators' IP ranges, so an IP-based allowlist or blocklist can answer the real bots differently. Confirm in your logs. A free Approvalens scan runs the equivalent checks for Google's crawlers.

FAQ

Why does my site load for me but block AI bots?

Rules that match on user agent or on "automated" traffic don't touch your browser. Test with the bot's user agent, and read your server logs for the real bot's requests.

I removed the rule but bots are still blocked. Why?

Check for another layer: a CDN in front, a host-level firewall, a plugin that rewrites .htaccess, or a cached robots.txt. Operators also take time to retry; OpenAI and Perplexity mention up to about 24 hours for robots.txt changes.

Is blocking by user agent useless?

It stops honest bots that announce themselves, which also respect robots.txt. It doesn't stop anyone faking a browser. IP-based rules are more precise for both allowing and blocking.

Can a rate limit look like a block?

Yes. A crawler that keeps getting 429 or 503 may slow down or stop. Exempt verified crawlers or raise limits on content pages.

Spotted something out of date or wrong? Tell us and we'll correct it.

Read this guide in Turkish →

Check it on your own site. Free, no sign-up.

Free tools for this

Free scan

Check your own site

Free scan: readiness score and every issue, usually in a few minutes.

Free scan · score and every problem found · no sign-up

All guides →