Measuring the Machine Web Without Pretending We Can Read an Agent's Mind
A practical evidence model for separating observed machine requests, inferred journeys, referrals, policy effects, and revenue.
AIWEBSIGNALS RESEARCH
We study what website operators can actually observe as AI systems crawl, retrieve, refer, interact with, and eventually pay for web resources. Every brief separates measured evidence from inference and links back to the provider or protocol documentation behind the claim.
This panel loads the current Activity Radar sweep so the publication has a live observation layer alongside long-form analysis.
Provider policies change, protocols evolve, and crawler behavior is easy to overstate. These pieces are written to remain useful by making the evidence boundary explicit.
A practical evidence model for separating observed machine requests, inferred journeys, referrals, policy effects, and revenue.
How to reason about OAI-SearchBot, GPTBot, OAI-AdsBot, robots.txt, and edge access without collapsing different purposes into one rule.
A site-owner guide to separating model-development crawling, search indexing, and user-requested retrieval in Anthropic traffic policy.
A current evidence-based guide to PerplexityBot access, indexing behavior, blocked-page limitations, and measurement.
Google-Extended is a robots.txt product token, not a separate HTTP user agent. Here's what it controls and what it does not change in Google Search.
A purpose-first framework for deciding which machine traffic to allow, observe, block, rate-limit, or eventually charge.
An evidence hierarchy for AI bot identification using declared tokens, provider documentation, verified network signals, request behavior, and uncertainty labels.
How to measure machine requests, AI-origin human visits, conversions, and revenue as separate stages of one commercial funnel.
A practical guide to 402 payment requirements, verification, settlement, replay safety, and evidence-based machine-access revenue.
Why semantic HTML, ARIA, stable controls, crawlable content, and predictable forms help both people and AI agents interact with a site.
A practical reading of Google's current guidance for AI Overviews and AI Mode, including indexing, snippets, unique content, and measurement.
Ten reproducible robots.txt observations show how often OAI-SearchBot, OAI-AdsBot, and GPTBot are explicitly separated versus governed by wildcard policy.
Ten reproducible robots.txt observations show how often ClaudeBot, Claude-SearchBot, and Claude-User are explicitly named versus governed by wildcard policy.
Ten reproducible robots.txt observations measure explicit versus inherited Google-Extended policy while preserving Google's no-separate-HTTP-user-agent distinction.
Four of five managed edge providers in the AIWebSignals connector cohort explicitly document cryptographic Web Bot Auth support; one remains unknown from bounded public documentation review.
Two provider paths can carry verification in request logs, one needs runtime instrumentation, and two require separate security streams for the strongest documented identity evidence.
Two providers now publish explicit three-way training/search/fetcher-agent taxonomies; the remaining cohort uses different purpose-aware control shapes.
The research library is complemented by seven implementation guides covering crawler access, robots.txt, llms.txt, structured data, sitemaps, and metadata.
FOUNDATIONAL READING
Before crawler rules, structured data, agent journeys, or machine payments, understand the basic difference between the web people experience and the networked resources machines actually request.