Original AIWebSignals evidence · fixed five-provider documentation cohort

AI Bot Purpose Controls at the Edge: A Five-Provider Taxonomy Study

The edge is moving beyond a single “AI crawler” bucket. The useful operator question is increasingly which purpose a machine request serves: background training, search/indexing, or real-time user-directed fetching and agent activity.

Aggregate result

  • 2/5 providers publish an explicit three-way AI-purpose taxonomy closely separating training, search and fetcher/agent activity.
  • 1/5 publishes a distinct AI crawler/fetcher split plus search-engine crawler classification, but its crawler class combines training and LLM indexing.
  • 2/5 document broader purpose-aware bot controls or identity categories without the same three-preset policy taxonomy in this bounded review.

This is not a provider ranking. It is a documentation taxonomy study of the five managed edge/CDN connector providers already represented in AIWebSignals.

Why this changed now

Cloudflare documents Search, Agent and Training presets and says new defaults take effect September 15, 2026: Training and Agent are blocked on ad-displaying pages while Search remains allowed for new domains. Akamai announced on September 3 that its Bot Directory has split the former AI Bots category into AI training crawlers, AI search crawlers, and AI fetchers and agents.

Those changes reinforce a product boundary AIWebSignals already treats as important: a crawler or agent identity is not enough to infer purpose, and a public robots policy is not the same thing as an edge-provider behavior policy.

Five-provider cohort

Observed at . Fixed five-provider cohort defined by AIWebSignals' existing managed edge/CDN connector paths: Cloudflare, Fastly, Vercel, AWS CloudFront/AWS WAF, and Akamai.

ProviderClassificationDocumented shapeEvidence
CloudflareExplicit three-way AI purpose policySearch / Agent / TrainingCloudflare documents separate AI behavior presets for Search, Agent, and Training. Its September 15, 2026 defaults block Training and Agent on ad-displaying pages while leaving Search allowed for new domains. First-party source
AkamaiExplicit three-way AI purpose policyAI training crawlers / AI search crawlers / AI fetchers and agentsAkamai announced on September 3, 2026 that its Bot Directory split the former AI Bots category into AI training crawlers, AI search crawlers, and AI fetchers and agents so customers can apply purpose-specific policies. First-party source
FastlyCrawler/fetcher purpose splitAI Crawler / AI Fetcher / Search Engine CrawlerFastly exposes distinct bot categories for AI crawlers, AI fetchers, and search-engine crawlers. Fastly's AI Crawler definition covers training or LLM indexes, so this is not the same three-way taxonomy as Cloudflare or Akamai. First-party source
VercelBroader purpose-aware controlsManaged AI-bot rules cover training, search, and user-generated fetchesVercel documents an AI Bots managed ruleset spanning bots used for training, search, or user-generated fetches and exposes verified-bot identity/category data. The bounded review did not establish an equivalent three-preset Search/Agent/Training policy surface. First-party source
AWSBroader purpose-aware controlsBot Control categories, verified search crawlers, browser automation and AI labelsAWS WAF Bot Control documents category and verification labels plus use cases that distinguish verified search crawlers from other bots and automated browsers. The bounded review did not establish an equivalent three-preset Search/Agent/Training policy surface. First-party source

What operators should keep separate

  1. Robots policy: what a cooperative crawler is declared to be allowed to request.
  2. Edge behavior policy: what a CDN/WAF/bot-management layer allows, challenges, rate-limits or blocks based on its own classifications.
  3. Observed first-party traffic: requests that actually reached a first-party logging surface.
  4. Verified identity: cryptographic or provider-native evidence that strengthens who sent a request.
  5. Referral/citation and economic outcomes: downstream evidence that must not be inferred from crawl access alone.

Methodology and limitations

Documentation-only classification from public first-party provider material. Categories describe documented control or taxonomy surfaces; they do not prove a feature is enabled for a particular customer or that any named crawler reached a site.

This study compares purpose granularity, not provider quality. Exact taxonomies differ, so similar-looking labels are not treated as interchangeable verification or traffic evidence.

The review used public first-party documentation only. It did not access customer consoles, impersonate crawlers, bypass access controls, or measure provider network traffic. Provider taxonomies can change, and a documented category does not prove that a specific request was correctly categorized or that the corresponding feature is enabled on a specific property.

Related AIWebSignals evidence

Measure policy and traffic as different evidence layers.

Use the public readiness scan for externally observable controls, then connect first-party traffic when you need to know what actually arrived.

Run a readiness scan