Google-Extended Policy Cohort: Explicit Product-Token Controls Across 10 Public Technology Sites
Google-Extended looks like a robots.txt user-agent token, but Google explicitly says it has no separate HTTP request user-agent string. It is a product-control token. We measured how often the same fixed ten-site technology cohort used in our OpenAI and Anthropic studies names Google-Extended explicitly versus leaving its homepage policy to broader wildcard rules.
Aggregate result
- 10/10 cohort robots.txt files were reproducibly retrieved.
- 10/10 had a permissive Google-Extended robots policy at
/under the observed named or inherited wildcard rules. - 2/10 explicitly named the Google-Extended product token.
- 8/10 left Google-Extended to inherited wildcard/default policy at the homepage path.
The defensible finding is about policy expression. It does not mean ten distinct Google-Extended crawlers could reach these sites, because there is no separate Google-Extended HTTP user agent to observe. Google says crawling is performed with existing Google user-agent strings and the Google-Extended token is used in a control capacity.
Google-Extended is not a separate crawler identity
Google's current crawling documentation distinguishes the robots.txt token from HTTP request identity: Google-Extended has no separate HTTP request user-agent string. The token lets publishers manage whether content Google crawls may be used for training future Gemini models and for specified grounding uses in Gemini Apps and Vertex AI. Google also states that Google-Extended does not affect inclusion in Google Search and is not a Google Search ranking signal.
That means a server log cannot establish a distinct "Google-Extended visit" by matching a Google-Extended HTTP user agent. Operators should treat the robots configuration question and the actual Google request-observation question as separate evidence layers.
Reproducible observations
Observed at . Scope: homepage path /. "Named Google-Extended group" means the exact product token appeared in a matching robots group. "Wildcard Allow" means it inherited an explicit User-agent: * allowance. "Wildcard/default" means no matching root-path disallow was observed, so the homepage remained permitted under ordinary robots matching semantics.
| Domain | Google-Extended robots policy at / |
|---|---|
| cloudflare.com robots.txt | Allowed · named Google-Extended group |
| vercel.com robots.txt | Allowed · wildcard/default |
| wordpress.com robots.txt | Allowed · wildcard/default |
| github.com robots.txt | Allowed · wildcard/default |
| stripe.com robots.txt | Allowed · wildcard/default |
| shopify.com robots.txt | Allowed · wildcard/default |
| hubspot.com robots.txt | Allowed · wildcard/default |
| zapier.com robots.txt | Allowed · named Google-Extended group |
| notion.so robots.txt | Allowed · wildcard Allow |
| wix.com robots.txt | Allowed · wildcard/default |
Cohort methodology
Technical observability cohort of ten public technology/marketing platforms selected before provider-policy inspection and reused from the preceding OpenAI and Anthropic studies. Inclusion required ordinary public retrieval of /robots.txt; the cohort is descriptive and not intended to estimate prevalence across the web.
The cohort consists of Cloudflare, Vercel, WordPress.com, GitHub, Stripe, Shopify, HubSpot, Zapier, Notion, and Wix. Reusing the same cohort reduces selection drift across the OpenAI, Anthropic, and Google policy studies.
This is a technical observability cohort, not a statistically representative sample of the web. The counts describe only these ten public robots files at the timestamp above.
Evidence boundary
Google-Extended is a robots.txt product token, not a separate HTTP crawler user-agent identity. named_group means the exact Google-Extended product token appeared in a matching robots group. wildcard_allow and wildcard_default are inherited robots policy at the homepage path. This study does not establish a distinct Google-Extended request, crawler visit, Gemini or Vertex AI use, training, grounding, citation, referral traffic, ranking effect, or commercial outcome.
This study did not impersonate a Google crawler, bypass a block, or access an authenticated surface. It observed public robots.txt configuration only. A permissive Google-Extended token does not prove that Google fetched a page, that content was used by Gemini or Vertex AI, or that any training or grounding event occurred.
Primary references
- Google Crawling Infrastructure: Google-Extended
- Google: How Google interprets the robots.txt specification
- RFC 9309: Robots Exclusion Protocol
Last reviewed September 8, 2026.
Related AIWebSignals evidence
Related technical evidence
Why Google-Extended is not a crawler · Google-Extended in the AI Bot Index · AI website readiness scanner