Field study
We scanned 100 websites. A score was not always an answer.
99 reports, 55 numeric scores, 44 withheld scores: what a varied public-web sample teaches about evidence, access, and useful AI-readiness assessments.
mistral.ai · observed 2026-10-02T18:18:25.625Z
docs.stripe.com · observed 2026-10-02T18:19:01.034Z

When does an AI-readiness score stop being enough?
The result, with the denominator intact
On October 2, 2026, we submitted a fixed roster of 100 hostnames to the public AIWebSignals scanner. Ninety-nine requests returned usable Evidence Report payloads. Fifty-five received a numeric technical-readiness score; 44 returned a report with that score withheld. One request, for hilton.com, timed out without a usable report. That last outcome describes our observation, not the availability of Hilton to every visitor.
The roster deliberately included software, commerce, beauty, publishing, travel, education and standards sites. It was selected for useful contrasts, not drawn randomly from the web. Two affiliated controls, AIWebSignals and Machine Realms, are disclosed in the methodology. Some hostnames belong to related organizations, such as Stripe and its documentation host. It would be misleading to turn these counts into an estimate of how ready the entire internet is.
A perfect technical score still leaves questions
Mistral returned 100/100 for the bounded public technical model. Its broader report contained evidence for 28 of 48 applicable checks, rounded to 58% coverage. Twenty remained unknown. Those are compatible results: the score evaluates a narrower set of public technical checks, while coverage describes how much of the larger assessment has supporting evidence.
The reverse matters too. A withheld score is not zero. It can mean the response was not safely identified as the origin page, a probe timed out, or the available evidence did not justify an aggregate. A convincing report should let the reader inspect those limits instead of hiding them behind a confident color.
What this changes for an operator
Start with one consequential question. Are public pages discoverable? Does the intended crawler policy match the business goal? Can a machine discover authorization requirements? Is there evidence of actual requests? Those questions require different evidence. A homepage scan can help with the first questions, but private execution and customer outcomes require operator-supplied records or an explicitly authorized test.
Use the report as a decision aid: preserve the exact hostname, observation date, method and missing evidence; identify a specific next check; then verify the result after a change. Do not remove an intentional restriction merely to improve a number. Do not claim recovered traffic or sales without measuring them.
A practical next test
Run the same bounded check on a hostname you operate, read the evidence behind the headline, and choose one reproducible action. A second report can then tell you whether the public signal changed. Connected first-party observation is the next step when your question becomes what actually reached the site, rather than what could be discovered publicly.
This publication is the beginning of a research loop, not its proof of commercial success. Reader questions and reproducible counterexamples can improve the next study. Views, shares and scans remain distinct from a paid activation; none of these 100 observations establishes customer demand or revenue lift.
Find the evidence on your own site
Run a free public scan, inspect one specific finding, and identify the next useful check. No private account access is required for the public baseline.
Evidence and primary sources
Observations can change. An interstitial or truncation classification can also require collector review. Submit a reproducible correction rather than treating a snapshot as a permanent company-wide verdict.
What should the next study answer?
Share a useful lesson, tell us what remains unclear, or send a reproducible question. Questions are reviewed; they do not automatically become published claims.
Continue with a related question

Mistral scored 100. Why was evidence coverage only 58%?
A concrete explanation of technical readiness versus evidence coverage, using Mistral’s October 2 public assessment and its 20 unknown checks.
Read the evidence and next step →
Stripe’s homepage and documentation produced different answers. Here is why that matters.
A 100/100 homepage result and a withheld documentation score show why hostname, response classification and inspection limits belong in every report.
Read the evidence and next step →Supabase scored 100—and the scan still hit an inspection limit
Why a homepage-truncated finding belongs beside a strong score, and how to distinguish bounded technical checks from exhaustive inspection.
Read the evidence and next step →