Field study

We scanned 100 websites. A score was not always an answer.

99 reports, 55 numeric scores, 44 withheld scores: what a varied public-web sample teaches about evidence, access, and useful AI-readiness assessments.

AIWebSignals Research3 minute readPublic-assessment case study
Report availability, technical quality and completeness of evidence are different results. This is a timestamped public assessment, not a claim of customer adoption, traffic loss, or revenue lift.

mistral.ai · observed 2026-10-02T18:18:25.625Z

100/100Bounded technical readiness
58%Evidence coverage, not a quality grade
20Unknown checks of 48 applicable
Inspect the saved mistral.ai report · Source JSON

docs.stripe.com · observed 2026-10-02T18:19:01.034Z

WithheldBounded technical readiness
27%Evidence coverage, not a quality grade
35Unknown checks of 48 applicable
Inspect the saved docs.stripe.com report · Source JSON
Actual AIWebSignals report screenshot for the saved mistral.ai observation
Actual report interface captured October 2. Resized for display; the saved evidence and scores are unchanged. Open the source report above for full context.

When does an AI-readiness score stop being enough?

The result, with the denominator intact

On October 2, 2026, we submitted a fixed roster of 100 hostnames to the public AIWebSignals scanner. Ninety-nine requests returned usable Evidence Report payloads. Fifty-five received a numeric technical-readiness score; 44 returned a report with that score withheld. One request, for hilton.com, timed out without a usable report. That last outcome describes our observation, not the availability of Hilton to every visitor.

The roster deliberately included software, commerce, beauty, publishing, travel, education and standards sites. It was selected for useful contrasts, not drawn randomly from the web. Two affiliated controls, AIWebSignals and Machine Realms, are disclosed in the methodology. Some hostnames belong to related organizations, such as Stripe and its documentation host. It would be misleading to turn these counts into an estimate of how ready the entire internet is.

A perfect technical score still leaves questions

Mistral returned 100/100 for the bounded public technical model. Its broader report contained evidence for 28 of 48 applicable checks, rounded to 58% coverage. Twenty remained unknown. Those are compatible results: the score evaluates a narrower set of public technical checks, while coverage describes how much of the larger assessment has supporting evidence.

The reverse matters too. A withheld score is not zero. It can mean the response was not safely identified as the origin page, a probe timed out, or the available evidence did not justify an aggregate. A convincing report should let the reader inspect those limits instead of hiding them behind a confident color.

What this changes for an operator

Start with one consequential question. Are public pages discoverable? Does the intended crawler policy match the business goal? Can a machine discover authorization requirements? Is there evidence of actual requests? Those questions require different evidence. A homepage scan can help with the first questions, but private execution and customer outcomes require operator-supplied records or an explicitly authorized test.

Use the report as a decision aid: preserve the exact hostname, observation date, method and missing evidence; identify a specific next check; then verify the result after a change. Do not remove an intentional restriction merely to improve a number. Do not claim recovered traffic or sales without measuring them.

A practical next test

Run the same bounded check on a hostname you operate, read the evidence behind the headline, and choose one reproducible action. A second report can then tell you whether the public signal changed. Connected first-party observation is the next step when your question becomes what actually reached the site, rather than what could be discovered publicly.

This publication is the beginning of a research loop, not its proof of commercial success. Reader questions and reproducible counterexamples can improve the next study. Views, shares and scans remain distinct from a paid activation; none of these 100 observations establishes customer demand or revenue lift.

Find the evidence on your own site

Run a free public scan, inspect one specific finding, and identify the next useful check. No private account access is required for the public baseline.

Evidence and primary sources

  1. Cohort method and result counts
  2. Mistral saved evidence
  3. Stripe documentation saved evidence

Observations can change. An interstitial or truncation classification can also require collector review. Submit a reproducible correction rather than treating a snapshot as a permanent company-wide verdict.

What should the next study answer?

Share a useful lesson, tell us what remains unclear, or send a reproducible question. Questions are reviewed; they do not automatically become published claims.

Suggest a question or correction

Continue with a related question