When AI crawlers get a 403, 429 or challenge page even though your robots.txt allows them, the block is in front of your content: a CDN or hosting firewall, a server rule, or a WordPress plugin. robots.txt only expresses a preference; these layers decide whether the request is answered at all. Check them in order from the outside in (CDN, host firewall, web server, CMS plugins), look for user-agent lists that name AI bots, and allow the bots you want by their published IP ranges rather than by user agent text.
How to tell it's a server-side block
| Symptom | Likely layer |
|---|---|
| robots.txt allows the bot, but logs show 403 for its requests | Firewall, server rule or plugin |
| Requests never appear in your server logs | CDN or host firewall in front of the server |
| A "Just a moment", "Checking your browser" or similar page | Bot challenge (CDN or host) |
| 429 Too Many Requests | Rate limiting |
| 200, but with a short block message in the body | Plugin or edge function returning a custom page |
Start with your CDN's and host's logs or security events, then the web server's access log, then the CMS.
Cloudflare
Cloudflare has its own AI bot policies, AI Crawl Control and managed robots.txt, plus Bot Fight Mode and WAF rules. It's common enough to have its own guide: Cloudflare blocking AI bots.
Vercel
Vercel's firewall has an AI Bots Managed Ruleset. Per Vercel's docs:
- It's available on all plans and "inactive by default", shown as Allow in the dashboard.
- You set it under the project's Firewall > Rules > Bot Management, choosing Log (record only) or Deny ("blocks all traffic identified as coming from AI bots").
- The separate Bot Protection Managed Ruleset can serve "a JavaScript challenge to traffic that is unlikely to be a browser" when set to Challenge.
If AI crawlers are blocked on a Vercel site, check whether either ruleset is set to Deny or Challenge. To let specific traffic through a managed ruleset, Vercel documents a bypass action in a WAF custom rule.
Netlify
Netlify offers a User Agent Blocker extension that "can block web requests from a preset list of common AI crawlers, SEO/Search crawlers that you choose from", using an Edge Function (Netlify docs). It's set per project under Extensions > User Agent Blocker > Block User Agents.
Two things to check: which options are ticked (the list includes search crawlers, not only AI training bots), and whether the extension is active on the project in question.
WordPress plugins
SEO and security plugins can add AI bot rules to robots.txt or block requests outright.
- Yoast SEO Premium: Settings > Advanced > Crawl optimization > Block unwanted bots has toggles for Google AdsBot, Google-Extended, GPTBot and CCBot. These write robots.txt rules (Yoast). Yoast's own help warns that blocking Google AdsBot "can have serious negative consequences if you are running Google Ads".
- The SEO Framework: SEO Settings > Robots Settings > Robots.txt has curated "AI" and "SEO" blocklists. The AI list includes Amazonbot, Applebot-Extended, CCBot, ClaudeBot, GPTBot, Google-Extended, GoogleOther, Meta-ExternalAgent and FacebookBot (TSF). It doesn't include OAI-SearchBot or PerplexityBot.
- Security and firewall plugins: many have user-agent blocklists, "bad bot" lists, country blocking and rate limiting. Look for any list containing
GPTBot,ClaudeBot,PerplexityBot,OAI-SearchBot,Bytespider,CCBotor a pattern likebot|crawl|spider.
On a hosted platform, check the platform's own setting too, for example Squarespace's "Block known artificial intelligence crawlers" box (Squarespace).
nginx
Search the configuration for user-agent rules:
grep -rniE "http_user_agent|gptbot|claudebot|perplexity|oai-searchbot|ccbot|bytespider" /etc/nginx/
A typical block looks like this:
if ($http_user_agent ~* "(GPTBot|ClaudeBot|PerplexityBot|CCBot)") {
return 403;
}
Remove the bots you want to allow from the pattern, or the whole block if robots.txt already says what you want. Also check limit_req zones (rate limits answer with 503 by default, or the status set by limit_req_status) and any deny rules for IP ranges. Reload nginx after editing (nginx -t first).
Apache and .htaccess
Search the virtual host files and every .htaccess:
grep -rniE "HTTP_USER_AGENT|BrowserMatch|SetEnvIf|gptbot|claudebot|perplexity" /etc/apache2/ /var/www/
Typical patterns:
RewriteCond %{HTTP_USER_AGENT} (GPTBot|ClaudeBot|PerplexityBot) [NC]
RewriteRule .* - [F,L]
or
SetEnvIfNoCase User-Agent "GPTBot" bad_bot
Require not env bad_bot
[F] returns 403 Forbidden. Edit the pattern to drop the bots you want to allow. If a plugin wrote the rule into .htaccess, it may come back while the plugin setting is still on, so change it in the plugin too.
Allow by IP, not by user agent
Matching on the user agent alone lets anyone in who types GPTBot into a request. The operators publish their IP ranges for this purpose:
| Operator | Ranges |
|---|---|
| OpenAI | gptbot.json, searchbot.json, chatgpt-user.json (docs) |
| Anthropic | bots.json (docs) |
| Perplexity | perplexitybot.json, perplexity-user.json (docs) |
| Apple | applebot.json, or reverse DNS in *.applebot.apple.com (docs) |
| Common Crawl | ccbot.json, reverse DNS in crawl.commoncrawl.org (docs) |
| Amazon | Amazonbot, Amzn-SearchBot, Amzn-User (docs) |
Perplexity recommends refreshing these automatically, since the ranges change.
One caution from Anthropic: blocking its IPs to opt out "may not work correctly", because the bot then can't read your robots.txt to see the opt-out. Use robots.txt to say no, and firewall rules for abuse.
Don't break Google while you're at it
A broad pattern like bot|crawler|spider also catches Googlebot and the AdSense crawler. If the same rule is blocking Google, that's an AdSense problem too; the site down or unavailable guide covers what that looks like in a review. Check with the Googlebot access checker after any change.
Check from outside
The free AI crawler checker requests your page with each major AI bot's user agent and reports the status code and any challenge page, next to what your robots.txt says, so a mismatch ("robots.txt allows, server says 403") stands out. These requests are simulated: they come from our server, not the operators' IP ranges, so an IP-based allowlist or blocklist can answer the real bots differently. Confirm in your logs. A free Approvalens scan runs the equivalent checks for Google's crawlers.
FAQ
Why does my site load for me but block AI bots?
Rules that match on user agent or on "automated" traffic don't touch your browser. Test with the bot's user agent, and read your server logs for the real bot's requests.
I removed the rule but bots are still blocked. Why?
Check for another layer: a CDN in front, a host-level firewall, a plugin that rewrites .htaccess, or a cached robots.txt. Operators also take time to retry; OpenAI and Perplexity mention up to about 24 hours for robots.txt changes.
Is blocking by user agent useless?
It stops honest bots that announce themselves, which also respect robots.txt. It doesn't stop anyone faking a browser. IP-based rules are more precise for both allowing and blocking.
Can a rate limit look like a block?
Yes. A crawler that keeps getting 429 or 503 may slow down or stop. Exempt verified crawlers or raise limits on content pages.
Spotted something out of date or wrong? Tell us and we'll correct it.
Read this guide in Turkish →Check it on your own site. Free, no sign-up.
Free tools for this
Free scan
Check your own site
Free scan: readiness score and every issue, usually in a few minutes.