Technical guide · original methodology
Robots.txt permission is not evidence of an AI visit
A practical evidence model for declared crawler policy, actual responses, verified requests, and downstream outcomes.
Four records, four meanings
Consider four statements: a robots rule permits an article, an audit retrieved the article, a verified provider request reached the article, and a person arrived from an AI service. Each statement has a different source and denominator. Reporting them under a single visibility metric makes failures hard to diagnose and successes easy to exaggerate.
Start with the declared rule
RFC 9309 describes cooperative crawler rules and explicitly distinguishes them from access authorization. Record the hostname, path, relevant crawler token, fetched policy version, and interpretation in the working evidence record. That record establishes what the policy declares. It does not tell you whether a provider scheduled a visit, whether a firewall challenged the request, or whether retrieved content was used in an answer.
Then inspect the response
A bounded public test can establish that one observer received one response at a particular time. Preserve the observer and conditions. If the response is a challenge page, classify the result as challenge-limited rather than successful simply because the HTTP status was 200. If the test times out, the result is timeout from that vantage, not a universal outage.
Ask for connected evidence when needed
Actual request history requires first-party or otherwise authorized logs. Identity claims require the method used to validate them. Even verified retrieval does not reveal training use, private prompts, model reasoning, citation, or commercial impact. Those remain different questions with different evidence requirements.
Use the distinction in decisions
A policy-response contradiction warrants an infrastructure investigation. A healthy response with no observed provider traffic suggests a different question about demand, discovery, or observation coverage. A referral without a matching crawler event should remain a referral, not be assigned to a guessed bot. This separation makes the observatory useful even when the answer is unknown.
Sources and scope
IETF: Robots Exclusion Protocol, RFC 9309 · MachineRealms: published interoperability evidence
Provider documents support the referenced technical facts. The interpretation, examples and proposed checks are AIWebSignals methodology, not provider endorsement or measured industry results.
Review the assessment method and its limits · Request a correction privately · Read the JSON version