AI BOT INDEX
Know which machine identity you are actually controlling.
Search crawlers, model-development crawlers, user-requested fetchers, advertising validators, and product-control tokens should not be collapsed into one generic “AI bot” category. This index records the provider's documented purpose and the site-owner control surface.
| Provider | Identity | Purpose | Control | What it means |
|---|---|---|---|---|
| OpenAI | OAI-SearchBot OAI-SearchBot | Search discovery | robots.txt | OpenAI documents OAI-SearchBot as the crawler site owners should allow when they want public pages eligible for ChatGPT search discovery. Allowing access does not guarantee ranking, citation, referral traffic, or inclusion in a particular answer. Source |
| OpenAI | GPTBot GPTBot | Model development / training control | robots.txt | GPTBot is separate from OpenAI search discovery. Publishers can use a distinct robots.txt policy for GPTBot when they want search visibility without granting the same access for potential model-development use. Source |
| OpenAI | OAI-AdsBot OAI-AdsBot | Advertising landing-page validation | robots.txt and edge access | OpenAI documents OAI-AdsBot for validating ad landing pages. It is operationally different from search discovery and training crawlers, which is why site operators should not collapse every OpenAI request into a single policy bucket. Source |
| Anthropic | ClaudeBot ClaudeBot | Model development | robots.txt | Anthropic describes ClaudeBot as a crawler used to collect public web content that may contribute to model development. Anthropic says the bot honors robots.txt directives and supports a non-standard Crawl-delay directive. Source |
| Anthropic | Claude-SearchBot Claude-SearchBot | Search indexing and result quality | robots.txt | Anthropic separates its search crawler from its model-development crawler. Blocking Claude-SearchBot can reduce the system's ability to index or surface the site in search-oriented Claude experiences. Source |
| Anthropic | Claude-User Claude-User | User-requested retrieval | robots.txt | Claude-User is used when a person asks Claude to access web content. Treating it separately from ClaudeBot lets a publisher distinguish user-requested retrieval from model-development crawling. Source |
| Perplexity | PerplexityBot PerplexityBot | Search indexing | robots.txt | Perplexity says PerplexityBot respects robots.txt and is used to index public pages for search-like retrieval. Perplexity also says blocking a page may still leave limited domain, headline, or brief factual information discoverable through other signals. Source |
| Google-Extended Google-Extended | AI training and grounding control token | robots.txt product token | Google-Extended is not a distinct HTTP crawler user agent. Google documents it as a robots.txt product token controlling whether content Google already crawls may be used for certain Gemini training and grounding purposes. It does not control inclusion in Google Search. Source |