Crawler policy
Together AI: a blocked token does not have to mean a broken site
Together AI returned 100/100 while the scan recorded Google-Extended as blocked. Learn why policy choices and technical readiness should stay separate.
together.ai · observed 2026-10-02T18:18:41.990Z

Would you change a deliberate policy just to make a dashboard greener?
The observed combination
Together AI’s October 2 public assessment returned 100/100 technical readiness and 58% evidence coverage. In that same report, the crawler-policy evidence recorded Google-Extended=blocked under an informational training-policy restriction. Seven of the eight monitored policy tokens resolved as allowed and one as blocked.
The score did not need to call that configuration broken. A machine-readable policy can be clear and intentional while restricting a particular use. The public scan establishes what its policy parser observed, not why the organization chose the rule or what requests subsequently reached its servers.
Google-Extended is a control token, not another request identity
Google documents Google-Extended as a robots.txt product token. It does not have a separate HTTP request user-agent string. The token controls specified Gemini training and grounding uses of content that Google crawls; Google says it does not control inclusion in Google Search and is not a Search ranking signal.
This makes a common reporting mistake easy to spot. A robots declaration involving Google-Extended is not a log of a visit by a crawler with that name. The evidence belongs in a policy table. Actual request identity, response status and timing belong in request telemetry. Joining those layers requires care rather than matching names and assuming an event happened.
The right remediation starts with a goal
Before recommending a policy change, ask which use the operator intends to allow. Supporting search discovery, permitting model-development collection and serving an explicit user request are not interchangeable objectives. A report that tells every site to allow every named bot would erase the very distinctions the operator needs.
Inspect the relevant robots group on the exact hostname and path. Confirm the provider’s current documentation. Where access is intended, separately check the returned page, authentication and edge behavior. A cooperative permission rule does not guarantee that the origin delivers useful content, just as a restriction does not prove a commercial loss.
Use a counterexample to improve the conversation
This case is useful because it challenges an oversimplified success criterion: green everywhere. A better question is whether the published rule is understandable, matches the desired use and is supported by an observation. Reproducible policy examples can produce more useful operator discussion than another blanket AI-visibility score.
Run a public baseline for a site you operate and inspect its purpose-specific policy. Keep the result timestamped. If the business question is what followed the rule in practice, use connected request evidence; do not treat the rule itself as proof of traffic, citation or sales.
Find the evidence on your own site
Run a free public scan, inspect one specific finding, and identify the next useful check. No private account access is required for the public baseline.
Evidence and primary sources
- Together AI: saved policy evidence
- Google: common crawlers and Google-Extended
- RFC 9309: Robots Exclusion Protocol
Observations can change. An interstitial or truncation classification can also require collector review. Submit a reproducible correction rather than treating a snapshot as a permanent company-wide verdict.
What should the next study answer?
Share a useful lesson, tell us what remains unclear, or send a reproducible question. Questions are reviewed; they do not automatically become published claims.
Continue with a related question
The BBC example: separate AI search policy from training controls
A 93/100 public snapshot with purpose-specific restrictions shows why publishers need more than an allow-all or block-all AI policy.
Read the evidence and next step →
Mistral scored 100. Why was evidence coverage only 58%?
A concrete explanation of technical readiness versus evidence coverage, using Mistral’s October 2 public assessment and its 20 unknown checks.
Read the evidence and next step →NASA’s scan and the problem with mandatory-sounding AI checklists
NASA’s public snapshot returned 91/100 with an informational llms.txt note. Emerging discovery proposals should not become universal requirements.
Read the evidence and next step →