Original AIWebSignals evidence · robots.txt product-token cohort

Google-Extended Policy Cohort: Explicit Product-Token Controls Across 10 Public Technology Sites

Google-Extended looks like a robots.txt user-agent token, but Google explicitly says it has no separate HTTP request user-agent string. It is a product-control token. We measured how often the same fixed ten-site technology cohort used in our OpenAI and Anthropic studies names Google-Extended explicitly versus leaving its homepage policy to broader wildcard rules.

Aggregate result

  • 10/10 cohort robots.txt files were reproducibly retrieved.
  • 10/10 had a permissive Google-Extended robots policy at / under the observed named or inherited wildcard rules.
  • 2/10 explicitly named the Google-Extended product token.
  • 8/10 left Google-Extended to inherited wildcard/default policy at the homepage path.

The defensible finding is about policy expression. It does not mean ten distinct Google-Extended crawlers could reach these sites, because there is no separate Google-Extended HTTP user agent to observe. Google says crawling is performed with existing Google user-agent strings and the Google-Extended token is used in a control capacity.

Google-Extended is not a separate crawler identity

Google's current crawling documentation distinguishes the robots.txt token from HTTP request identity: Google-Extended has no separate HTTP request user-agent string. The token lets publishers manage whether content Google crawls may be used for training future Gemini models and for specified grounding uses in Gemini Apps and Vertex AI. Google also states that Google-Extended does not affect inclusion in Google Search and is not a Google Search ranking signal.

That means a server log cannot establish a distinct "Google-Extended visit" by matching a Google-Extended HTTP user agent. Operators should treat the robots configuration question and the actual Google request-observation question as separate evidence layers.

Reproducible observations

Observed at . Scope: homepage path /. "Named Google-Extended group" means the exact product token appeared in a matching robots group. "Wildcard Allow" means it inherited an explicit User-agent: * allowance. "Wildcard/default" means no matching root-path disallow was observed, so the homepage remained permitted under ordinary robots matching semantics.

DomainGoogle-Extended robots policy at /
cloudflare.com
robots.txt
Allowed · named Google-Extended group
vercel.com
robots.txt
Allowed · wildcard/default
wordpress.com
robots.txt
Allowed · wildcard/default
github.com
robots.txt
Allowed · wildcard/default
stripe.com
robots.txt
Allowed · wildcard/default
shopify.com
robots.txt
Allowed · wildcard/default
hubspot.com
robots.txt
Allowed · wildcard/default
zapier.com
robots.txt
Allowed · named Google-Extended group
notion.so
robots.txt
Allowed · wildcard Allow
wix.com
robots.txt
Allowed · wildcard/default

Cohort methodology

Technical observability cohort of ten public technology/marketing platforms selected before provider-policy inspection and reused from the preceding OpenAI and Anthropic studies. Inclusion required ordinary public retrieval of /robots.txt; the cohort is descriptive and not intended to estimate prevalence across the web.

The cohort consists of Cloudflare, Vercel, WordPress.com, GitHub, Stripe, Shopify, HubSpot, Zapier, Notion, and Wix. Reusing the same cohort reduces selection drift across the OpenAI, Anthropic, and Google policy studies.

This is a technical observability cohort, not a statistically representative sample of the web. The counts describe only these ten public robots files at the timestamp above.

Evidence boundary

Google-Extended is a robots.txt product token, not a separate HTTP crawler user-agent identity. named_group means the exact Google-Extended product token appeared in a matching robots group. wildcard_allow and wildcard_default are inherited robots policy at the homepage path. This study does not establish a distinct Google-Extended request, crawler visit, Gemini or Vertex AI use, training, grounding, citation, referral traffic, ranking effect, or commercial outcome.

This study did not impersonate a Google crawler, bypass a block, or access an authenticated surface. It observed public robots.txt configuration only. A permissive Google-Extended token does not prove that Google fetched a page, that content was used by Gemini or Vertex AI, or that any training or grounding event occurred.

Primary references

Last reviewed September 8, 2026.

Related AIWebSignals evidence

Related technical evidence

Why Google-Extended is not a crawler · Google-Extended in the AI Bot Index · AI website readiness scanner

Separate control tokens from request identities.

AIWebSignals keeps robots policy, HTTP request identity, public edge observations, and first-party traffic evidence as distinct layers.

Open the scanner