Publishing
The BBC example: separate AI search policy from training controls
A 93/100 public snapshot with purpose-specific restrictions shows why publishers need more than an allow-all or block-all AI policy.
bbc.com · observed 2026-10-02T18:24:23.388Z
Which use does a publisher actually intend to allow?
What the snapshot recorded
The October 2 bbc.com report returned 93/100 technical readiness and 58% evidence coverage. It recorded OAI-SearchBot and PerplexityBot as blocked in the search-related policy group. GPTBot, ClaudeBot and Google-Extended appeared in the separate training or product-control group. An llms.txt-not-found note was informational.
These are observations of public declarations as parsed by the scanner. They do not establish the BBC’s private strategy, whether every named system attempted a request, whether any request honored a rule or whether a policy affected commercial performance.
Why the purpose distinction matters
A publisher can evaluate search discovery differently from content collection for model development. A user-directed retrieval has another context again. A single provider name does not describe all these uses, and a robots control token is not necessarily an HTTP user-agent identity.
For example, Google documents Google-Extended as a product-control token, not a separate crawler user agent. Its declaration should not be counted as a visit. The same discipline applies to the overall policy table: configuration is an input to a cooperative access policy, while requests and responses are observations of what actually occurred.
Use the operator’s objective as the decision rule
The first question is not how to remove every warning. It is which audience and use the publisher wants to support. Review provider documentation and the relevant hostname/path policy, then compare the intended access with the public declaration. If the desired audience should have access, inspect the actual response and edge behavior separately.
If a restriction is deliberate, preserve that choice rather than recommend reversing it merely to make a report look more positive. If the declaration is ambiguous or inconsistent with the stated goal, describe the exact inconsistency and a verification step. Neither situation justifies inventing a traffic or revenue impact.
Build an evidence loop around the policy
Preserve a timestamped baseline, make a deliberate authorized change only when the business objective calls for it, and record the resulting public policy. Then compare appropriately collected first-party requests over a defined observation window. Keep a configuration change, an access change, a referral and a paid outcome as separate claims.
A public scan can begin that conversation. Connected evidence is the next layer when an operator asks what followed the policy in practice. That is a more useful acquisition and product story than an instruction that every publisher should expose everything to every crawler.
Find the evidence on your own site
Run a free public scan, inspect one specific finding, and identify the next useful check. No private account access is required for the public baseline.
Evidence and primary sources
- BBC: saved purpose-specific policy findings
- Google: common crawlers and Google-Extended
- RFC 9309: Robots Exclusion Protocol
Observations can change. An interstitial or truncation classification can also require collector review. Submit a reproducible correction rather than treating a snapshot as a permanent company-wide verdict.
What should the next study answer?
Share a useful lesson, tell us what remains unclear, or send a reproducible question. Questions are reviewed; they do not automatically become published claims.
Continue with a related question

Together AI: a blocked token does not have to mean a broken site
Together AI returned 100/100 while the scan recorded Google-Extended as blocked. Learn why policy choices and technical readiness should stay separate.
Read the evidence and next step →NASA’s scan and the problem with mandatory-sounding AI checklists
NASA’s public snapshot returned 91/100 with an informational llms.txt note. Emerging discovery proposals should not become universal requirements.
Read the evidence and next step →
We scanned 100 websites. A score was not always an answer.
99 reports, 55 numeric scores, 44 withheld scores: what a varied public-web sample teaches about evidence, access, and useful AI-readiness assessments.
Read the evidence and next step →