Skip to content
Approvalens

Reading room · 4 min read

PerplexityBot Not Crawling Your Site? Causes and Fixes

Why Perplexity may not cite your pages: robots.txt, WAF and CDN blocks, IP allowlists and JavaScript-only content, and how to check each one.

By the Approvalens team

Fixes these report findings

  • AI crawlers get your pages from the server
  • AI search crawlers allowed in robots.txt

If Perplexity never cites your site, the usual causes are a robots.txt rule that blocks PerplexityBot, a firewall or CDN that rejects its requests, or pages whose text only appears after JavaScript runs. Perplexity says PerplexityBot "is designed to surface and link websites in search results on Perplexity" and is "not used to crawl content for AI foundation models", and it recommends allowing the bot in robots.txt and allowing requests from its published IP ranges. Fix those three things, wait at least 24 hours, and then check your logs.

Perplexity's two user agents

From Perplexity's crawler documentation:

User agent Purpose robots.txt IP ranges
PerplexityBot Surfacing and linking websites in Perplexity search results. Not used for AI foundation model training Respected; Perplexity recommends allowing it perplexitybot.json
Perplexity-User Visiting a page when a user asks Perplexity a question, to answer accurately and link to the page "Since a user requested the fetch, this fetcher generally ignores robots.txt rules" perplexity-user.json

Full user agent strings, as published:

Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; PerplexityBot/1.0; +https://perplexity.ai/perplexitybot)
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Perplexity-User/1.0; +https://perplexity.ai/perplexity-user)

Perplexity notes that "each setting works independently, and it may take up to 24 hours for our systems to reflect changes."

Cause 1: robots.txt

Look for any of these in https://yourdomain.com/robots.txt:

  • User-agent: PerplexityBot followed by Disallow: /;
  • User-agent: * with Disallow: / and no separate PerplexityBot group;
  • a pasted list of AI bots that includes PerplexityBot.

Because PerplexityBot isn't a training crawler, blocking it to "stop AI training" removes you from Perplexity's search results without achieving what you wanted. To allow it explicitly:

User-agent: PerplexityBot
Allow: /

The allow AI crawlers in robots.txt guide shows a full file that allows AI search bots and blocks training bots.

Cause 2: firewall, WAF or CDN

Perplexity's own docs have a section for this: "If you're using a Web Application Firewall (WAF) to protect your site, you may need to explicitly whitelist Perplexity's bots". Its Cloudflare instructions describe a custom rule that matches the user agent PerplexityBot or Perplexity-User and an IP source address in its published ranges, set to let the traffic through. Its AWS WAF instructions do the same with IP sets and string match conditions.

It also asks you to keep those ranges current: "Set up automated processes to periodically fetch and update your WAF rules with the latest IP ranges from these endpoints."

Common blockers to check:

  • Cloudflare AI bot policies and AI Crawl Control: a block on Search, or on PerplexityBot individually. See Cloudflare blocking AI bots.
  • Bot protection and challenges: Bot Fight Mode, JavaScript challenges, "I'm Under Attack" mode.
  • Country or ASN rules that only allow visitors from your own country.
  • Hosting firewalls and WordPress security plugins with user-agent blocklists. See firewall and plugin rules.

Cause 3: content that needs JavaScript

Perplexity doesn't document whether PerplexityBot renders JavaScript. If your article text is loaded by scripts after the page loads, assume a crawler may see an empty page. Check with View Source or curl and look for a sentence from the middle of an article. See JavaScript content and AI crawlers.

Cause 4: nothing is wrong yet

Being crawlable doesn't mean Perplexity will cite you for a given question. Changes need up to 24 hours to apply, and Perplexity chooses which sources to show. If your logs show PerplexityBot fetching pages with 200 responses, access isn't the problem.

Read your logs

Search your server or CDN logs for PerplexityBot and Perplexity-User:

What you see Likely meaning
No requests at all Blocked before your server (CDN, WAF), or not crawled yet
Requests to /robots.txt only robots.txt disallows the bot
403, 429 or 503 responses Firewall, rate limit or challenge
200 responses for pages Access works

Check that requests claiming to be PerplexityBot come from the published ranges; Perplexity recommends combining user agent and IP checks.

Perplexity and AdSense

Unrelated. PerplexityBot has nothing to do with AdSense, and allowing or blocking it doesn't affect approval. The robots.txt and AdSense guide covers Google's ad crawler, and a free Approvalens scan checks the AdSense side.

Check it for free

The free AI crawler checker reads your robots.txt as PerplexityBot and the other AI bots would and requests your page with their user agents, so you can spot a robots.txt block, a 403 or a challenge page. Our requests are simulated: they carry Perplexity's user agent but come from our server, not Perplexity's IP ranges, so a WAF rule that checks IPs can respond differently to the real bot. The robots.txt tester checks the same file for Google.

FAQ

Does blocking PerplexityBot stop Perplexity from training on my content?

Perplexity says PerplexityBot is not used to crawl content for AI foundation models. Blocking it removes you from Perplexity search results.

Can I stop Perplexity-User with robots.txt?

Perplexity says it "generally ignores robots.txt rules" because a user requested the fetch. A firewall block works, but then Perplexity can't read your page for people asking about it.

I allowed PerplexityBot yesterday. Why am I still not cited?

Perplexity says changes may take up to 24 hours, and being reachable doesn't guarantee a citation. Check your logs for successful fetches first.

Should I allowlist Perplexity by user agent only?

No. Anyone can send that user agent. Perplexity recommends matching both the user agent and its published IP ranges.

Spotted something out of date or wrong? Tell us and we'll correct it.

Read this guide in Turkish →

Check it on your own site. Free, no sign-up.

Free tools for this

Free scan

Check your own site

Free scan: readiness score and every issue, usually in a few minutes.

Free scan · score and every problem found · no sign-up

All guides →