PerplexityBot and Perplexity-User: what they do on your site

·7 min read

PerplexityBot is Perplexity's search crawler. It reads pages so Perplexity can show and link them in its answers, and Perplexity says it does not crawl for AI model training. Perplexity-User is a separate agent that opens a page when someone's question needs it, and Perplexity says that one generally ignores robots.txt because a person asked for the fetch. A rule for PerplexityBot does nothing to Perplexity-User, and blocking PerplexityBot mostly costs you citations in Perplexity.

Most sites should allow both. If you want out of AI training, PerplexityBot is the wrong bot to block. The harder question is trust. In 2025 Cloudflare accused Perplexity of reaching blocked sites under a disguised identity, and Perplexity said Cloudflare had its facts wrong. Both accounts are below, with checks you can run on your own logs.

What PerplexityBot does

PerplexityBot crawls the web to build the index Perplexity searches when it writes an answer. Perplexity's crawler documentation says the bot exists to surface and link sites in Perplexity's results and is not used to crawl content for AI foundation models. Its robots.txt help article, last modified September 14, 2026, adds that Perplexity builds no foundation models, so allowing the bot does not put your pages into pre-training. Perplexity recommends allowing it.

PerplexityBot is not Perplexity's only source. The same article says Perplexity also uses third-party crawlers for its index and has updated its agreements so they respect robots.txt, particularly on news publisher sites. It doesn't name them, so you have no user agent to write a rule for.

What Perplexity-User does

Perplexity-User loads a page in real time when a Perplexity user asks something the page can answer, and Perplexity may link the page in its reply. Perplexity says the agent does not crawl the web or collect training data.

The docs say that because a user requested the fetch, Perplexity-User generally ignores robots.txt rules. The help center adds that users could once ask Perplexity to summarize a URL that robots.txt blocked, and that the feature has been disabled. Neither page says how the two statements fit together, so treat a Perplexity-User rule as a stated preference and use your firewall if you need it enforced.

PerplexityBot Perplexity-User
What starts a visit Perplexity building its search index A user's question
robots.txt Follows it, per Perplexity Generally ignores it, per Perplexity
Used to train models No, per Perplexity No, per Perplexity
What blocking it costs Perplexity has none of your text to quote Perplexity can't open the page when someone asks about it

The PerplexityBot user agent and IP lists

Perplexity's docs give these user agent strings:

Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; PerplexityBot/1.0; +https://perplexity.ai/perplexitybot)
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Perplexity-User/1.0; +https://perplexity.ai/perplexity-user)

Search your logs for the token, PerplexityBot or Perplexity-User, without regard to case. robots.txt matches the token the same way, so User-agent: perplexitybot works.

Anyone can send that string, so the IP lists are the real check. We fetched both files on October 10, 2026:

Agent IP list Date in the file Entries
PerplexityBot perplexitybot.json February 7, 2025 8 IPv4 ranges, 18 addresses
Perplexity-User perplexity-user.json October 17, 2025 4 IPv4 ranges, 14 addresses

Both lists are short and IPv4 only. Perplexity's docs say the files are updated regularly and should be your source of truth, so fetch them fresh instead of pasting today's ranges into a rule. Perplexity documents no reverse DNS hostname for either agent, which leaves the IP list as the only verification it offers.

Find and verify Perplexity requests in your logs

These commands assume the Nginx or Apache combined log format, with the client IP in the first field. zgrep reads rotated .gz files as well as the live log.

# Requests from each agent
zgrep -ih "perplexitybot" /var/log/nginx/access.log* | wc -l
zgrep -ih "perplexity-user" /var/log/nginx/access.log* | wc -l

# Pages Perplexity-User opened for people's questions
zgrep -ih "perplexity-user" /var/log/nginx/access.log* | awk '{print $7}' | sort | uniq -c | sort -rn | head -20

# Status codes PerplexityBot received
zgrep -ih "perplexitybot" /var/log/nginx/access.log* | awk '{print $9}' | sort | uniq -c | sort -rn

Then test every IP that claimed to be PerplexityBot against the published list:

curl -sL https://www.perplexity.com/perplexitybot.json -o perplexitybot.json
zgrep -ih "perplexitybot" /var/log/nginx/access.log* | awk '{print $1}' | sort -u > claimed.txt
python3 - <<'EOF'
import ipaddress, json
prefixes = json.load(open("perplexitybot.json"))["prefixes"]
nets = [ipaddress.ip_network(p.get("ipv4Prefix") or p.get("ipv6Prefix")) for p in prefixes]
for line in open("claimed.txt"):
    ip = ipaddress.ip_address(line.strip())
    print(ip, "listed" if any(ip in net for net in nets) else "NOT LISTED")
EOF

Swap in perplexity-user.json and "perplexity-user" for the other agent. NOT LISTED means a request used Perplexity's name from an address Perplexity doesn't publish. If every line says NOT LISTED, check your proxy first. Behind Cloudflare or a load balancer the first field may be the proxy, and the client IP sits in CF-Connecting-IP or X-Forwarded-For. Our AI Bot Log Analyzer runs the same check on a whole log and counts the requests that used a crawler's name from another IP.

Does the Perplexity crawler respect robots.txt?

Perplexity says PerplexityBot does. Its help article says the bot will not index the full or partial text of a site that disallows it, though Perplexity may still index the domain, the headline and a brief factual summary. A Disallow keeps your words out of Perplexity's index. It doesn't make Perplexity forget the page exists. Perplexity-User remains the documented exception.

Cloudflare's stealth crawler report and Perplexity's reply

On August 4, 2025, Cloudflare published a report saying Perplexity used undeclared crawlers to reach sites that had blocked it, and Perplexity published its rebuttal the same day.

What Cloudflare reported. Customers told Cloudflare that Perplexity still reached their content after they disallowed it in robots.txt and blocked both of its agents with firewall rules. Cloudflare set up new domains that no search engine had indexed, disallowed all crawlers, and asked Perplexity about them. According to the Cloudflare post, Perplexity answered with details of the content. Once the declared agents were blocked, requests arrived with a generic Chrome-on-macOS user agent, from IPs outside Perplexity's lists that rotated across networks. Cloudflare counted 20 to 25 million declared Perplexity-User requests a day and 3 to 6 million from the undeclared crawler, which it saw on tens of thousands of domains. It removed Perplexity from its verified bots and added the crawler's signatures to its managed rule for AI crawlers.

What Perplexity replied. In Agents or bots? Making sense of AI on the open web, Perplexity said Cloudflare had confused user-driven fetching with crawling and that the 20 to 25 million requests came from people asking questions. It said the 3 to 6 million came from BrowserBase, a third-party cloud browser service that Perplexity uses only occasionally for specialized tasks, at fewer than 45,000 requests a day. It compared its agents to Google's user-triggered fetchers, said they neither store nor train on what they fetch, and said Cloudflare obscured its methodology.

Nobody outside the two companies can check the raw traffic, so I won't pick a winner. The lesson for your site is narrower. robots.txt is a request, and Perplexity says one of its two agents generally doesn't follow it. If you need Perplexity out, enforce that at your CDN or firewall, which our guide to stopping AI scraping covers.

How to block Perplexity in robots.txt

Give PerplexityBot its own group with Disallow: /. Perplexity says robots.txt changes can take up to 24 hours to reach its systems.

Block PerplexityBot only

User-agent: PerplexityBot
Disallow: /

Perplexity-User has no group here, so it falls back to your User-agent: * rules, to the extent it reads them.

Block both Perplexity agents

User-agent: PerplexityBot
User-agent: Perplexity-User
Disallow: /

This stops PerplexityBot and records your preference for Perplexity-User. Enforcing it on Perplexity-User takes a firewall rule.

Keep PerplexityBot out of some paths

User-agent: *
Disallow: /account/
Disallow: /checkout/

User-agent: PerplexityBot
Disallow: /account/
Disallow: /checkout/
Disallow: /members/

Copy your * rules into the PerplexityBot group. Under RFC 9309, a crawler that finds a group naming it ignores the * group, so a PerplexityBot group listing only /members/ would open /account/ and /checkout/ to it.

What blocking PerplexityBot costs you

Blocking PerplexityBot costs you citations in Perplexity and buys no training opt-out, because Perplexity says the bot collects no training data. A blocked page can still sit in the index as a headline and a summary with none of your text to quote, so the pages Perplexity can read win the answer instead. How Perplexity picks among those is covered in our guide to Perplexity SEO.

By Cloudflare's count, Perplexity also sends back more traffic per crawl than Anthropic does. Its crawl-to-refer data, published August 29, 2025, counted about 195 Perplexity crawls for every visit Perplexity referred in July 2025, up from about 55 in January. Anthropic's July figure was about 38,000.

Blocking makes sense when your text is the product, as with paywalled reporting. Block both agents and add the firewall rule. If you only want out of AI training, block training crawlers such as GPTBot, ClaudeBot and CCBot instead, which our AI crawler list sorts by job. Check your CDN as well, because a bot-protection setting aimed at AI bots can answer PerplexityBot with a 403 while robots.txt allows it.

Check your robots.txt for Perplexity's crawlers

Our robots.txt checker fetches your live file and shows whether your rules let 14 AI user agents, PerplexityBot and Perplexity-User among them, fetch the URL you enter. For each one it names the rule that decided it and says when the bot fell back to your * group, which is how the missing-rules mistake above shows up. You can also test any path against any user agent. It flags syntax errors, a user agent named in more than one group, Allow and Disallow rules with the same pattern, a missing Sitemap line, and Content-Signal lines that opt out of AI input or search.

It doesn't pretend to be Perplexity, so the Perplexity-User row shows what your file asks of that agent, and your logs show whether it listens. It does report when your server answers its own request with a 401, 403 or 429, a sign that a firewall rule may be turning crawlers away too. A run costs 10 credits.

Keep reading