AIWebSignalsobserve the machine web

AI BOT INDEX

Know which machine identity you are actually controlling.

Search crawlers, model-development crawlers, user-requested fetchers, advertising validators, and product-control tokens should not be collapsed into one generic “AI bot” category. This index records the provider's documented purpose and the site-owner control surface.

Evidence boundary: a documented user-agent token can support classification, but the string itself is not cryptographic proof of provider identity. Google-Extended is listed because it matters to policy even though Google explicitly says it is not a separate HTTP request user agent.
ProviderIdentityPurposeControlWhat it means
OpenAIOAI-SearchBot
OAI-SearchBot
Search discoveryrobots.txtOpenAI documents OAI-SearchBot as the crawler site owners should allow when they want public pages eligible for ChatGPT search discovery. Allowing access does not guarantee ranking, citation, referral traffic, or inclusion in a particular answer. Source
OpenAIGPTBot
GPTBot
Model development / training controlrobots.txtGPTBot is separate from OpenAI search discovery. Publishers can use a distinct robots.txt policy for GPTBot when they want search visibility without granting the same access for potential model-development use. Source
OpenAIOAI-AdsBot
OAI-AdsBot
Advertising landing-page validationrobots.txt and edge accessOpenAI documents OAI-AdsBot for validating ad landing pages. It is operationally different from search discovery and training crawlers, which is why site operators should not collapse every OpenAI request into a single policy bucket. Source
AnthropicClaudeBot
ClaudeBot
Model developmentrobots.txtAnthropic describes ClaudeBot as a crawler used to collect public web content that may contribute to model development. Anthropic says the bot honors robots.txt directives and supports a non-standard Crawl-delay directive. Source
AnthropicClaude-SearchBot
Claude-SearchBot
Search indexing and result qualityrobots.txtAnthropic separates its search crawler from its model-development crawler. Blocking Claude-SearchBot can reduce the system's ability to index or surface the site in search-oriented Claude experiences. Source
AnthropicClaude-User
Claude-User
User-requested retrievalrobots.txtClaude-User is used when a person asks Claude to access web content. Treating it separately from ClaudeBot lets a publisher distinguish user-requested retrieval from model-development crawling. Source
PerplexityPerplexityBot
PerplexityBot
Search indexingrobots.txtPerplexity says PerplexityBot respects robots.txt and is used to index public pages for search-like retrieval. Perplexity also says blocking a page may still leave limited domain, headline, or brief factual information discoverable through other signals. Source
GoogleGoogle-Extended
Google-Extended
AI training and grounding control tokenrobots.txt product tokenGoogle-Extended is not a distinct HTTP crawler user agent. Google documents it as a robots.txt product token controlling whether content Google already crawls may be used for certain Gemini training and grounding purposes. It does not control inclusion in Google Search. Source