Technical guide · original methodology
AI referral traffic and crawler traffic are different datasets
How to avoid treating automated requests as customers or attributing an outcome to a crawler without evidence.
Do not mix the denominators
A request log counts requests, a session report counts sessions under that system’s rules, and an account ledger counts accounts. A crawler can make many requests without bringing a human visitor. A human can arrive from an AI answer without a same-day crawler request appearing in the site’s logs. The observatory keeps these measures separate.
Describe identities narrowly
Provider documentation can explain the declared purpose of a crawler token. It cannot reveal what happened to a particular fetched page. OpenAI, for example, distinguishes search-oriented crawling from model-development crawling. A token label is therefore useful classification metadata, but it is not proof of a specific user’s journey or a private downstream use.
Build an evidence chain
For public research, report only approved observations with a time window and coverage statement. For an authorized commercial evaluation, connect an attributed arrival to a real account or paid activation through the product’s existing first-party system. Keep the original acquisition source; internal research-to-pricing links should not overwrite it with invented campaign parameters.
Leave missing links missing
Referrer removal, shared devices, consent choices, blocked analytics, and incomplete logging can leave gaps. Record coverage limits rather than manufacturing a deterministic identity from IP address or timing. In particular, do not join unrelated people across websites simply because those sites share an operator.
An actionable scorecard
Use separate rows for public scan starts, completed reports, connected-site activations, paid activations, verified automated requests, and attributed human referrals. If one increases while another does not, that is a finding to investigate. It is not a reason to rename the first metric until the story looks more successful.
Sources and scope
OpenAI: documented crawler purposes · AIWebSignals: current public capability contract
Provider documents support the referenced technical facts. The interpretation, examples and proposed checks are AIWebSignals methodology, not provider endorsement or measured industry results.
Review the assessment method and its limits · Request a correction privately · Read the JSON version