{"slug":"ai-referrals-vs-crawlers","title":"AI referral traffic and crawler traffic are different datasets","description":"How to avoid treating automated requests as customers or attributing an outcome to a crawler without evidence.","sources":[{"name":"OpenAI: documented crawler purposes","url":"https://developers.openai.com/api/docs/bots"},{"name":"AIWebSignals: current public capability contract","url":"https://aiwebsignals.com/api/capabilities"}],"sections":[["Do not mix the denominators","A request log counts requests, a session report counts sessions under that system’s rules, and an account ledger counts accounts. A crawler can make many requests without bringing a human visitor. A human can arrive from an AI answer without a same-day crawler request appearing in the site’s logs. The observatory keeps these measures separate."],["Describe identities narrowly","Provider documentation can explain the declared purpose of a crawler token. It cannot reveal what happened to a particular fetched page. OpenAI, for example, distinguishes search-oriented crawling from model-development crawling. A token label is therefore useful classification metadata, but it is not proof of a specific user’s journey or a private downstream use."],["Build an evidence chain","For public research, report only approved observations with a time window and coverage statement. For an authorized commercial evaluation, connect an attributed arrival to a real account or paid activation through the product’s existing first-party system. Keep the original acquisition source; internal research-to-pricing links should not overwrite it with invented campaign parameters."],["Leave missing links missing","Referrer removal, shared devices, consent choices, blocked analytics, and incomplete logging can leave gaps. Record coverage limits rather than manufacturing a deterministic identity from IP address or timing. In particular, do not join unrelated people across websites simply because those sites share an operator."],["An actionable scorecard","Use separate rows for public scan starts, completed reports, connected-site activations, paid activations, verified automated requests, and attributed human referrals. If one increases while another does not, that is a finding to investigate. It is not a reason to rename the first metric until the story looks more successful."]],"next":"credential-minimized-observation","published":"2026-09-26","updated":"2026-09-26","status":"published","author":"AIWebSignals Research","url":"https://aiwebsignals.com/research/regulated-commerce/ai-referrals-vs-crawlers","example":{"type":"synthetic","heading":"Requests are not customers","text":"An invented report contains twelve crawler requests and three sessions with an AI-service referrer. Neither record links those sessions to the requests. Report each measure separately; do not turn the difference into nine abandoned customers. These numbers illustrate incompatible denominators, not industry performance.","exercise":"Separate request counts, attributed human arrivals, completed scans, connected-site activations and paid outcomes. Mark unavailable stages unknown rather than estimating a conversion funnel from unrelated totals."},"methodology":"https://aiwebsignals.com/research/regulated-commerce/methodology"}