Perplexity's current published position

Perplexity's help documentation states that PerplexityBot will not index the full or partial text of a site that disallows the crawler through robots.txt. The company also says it does not build foundation models and does not use PerplexityBot content for foundation-model pretraining. That documented purpose is relevant when a publisher decides whether the crawler belongs in a search-discovery bucket or a training-access bucket.

The same documentation notes that a blocked page can still leave limited domain, headline, or brief factual signals discoverable. That is an important reminder that robots.txt governs cooperative crawling of the page; it is not a universal deletion mechanism for every reference to a URL that may exist elsewhere.

Access policy still needs edge verification

As with other bots, a permissive robots.txt policy does not guarantee that PerplexityBot receives a useful response. Bot management can return a challenge, application middleware can block the request, and an origin can return errors. If search visibility matters, the site should test the actual response path and monitor status codes rather than considering the robots file the final proof.

For publishers that intentionally block PerplexityBot, first-party logs provide the inverse test: after the rule is deployed, does the observed full-page crawl behavior fall away? A policy that is not reflected in request outcomes deserves investigation.

Measure machine and human value separately

PerplexityBot request volume is a machine-access metric. Human visits from Perplexity are a referral metric. The two may correlate, but they are not interchangeable. A publisher trying to value access should compare bot requests with referral sessions, engaged time, signups, conversions, or other site outcomes instead of assigning a dollar figure to crawler hits.

That same separation makes licensing decisions more rational. If crawler activity is heavy but human referral value is negligible, the publisher has evidence to consider a more restrictive policy. If referrals convert unusually well, maintaining discoverability may be strategically valuable even when request volume itself is modest.

The durable operating rule

Document the intended policy, verify the deployed robots response, inspect actual edge behavior, and track downstream referrals independently. Revisit the policy when provider documentation changes. The machine web is moving quickly enough that a robots rule copied from an old blog post should not be treated as permanent truth.

Related technical evidence

PerplexityBot in the AI Bot Index · robots.txt for AI crawlers · AI crawler policy matrix