Score interpretation
Mistral scored 100. Why was evidence coverage only 58%?
A concrete explanation of technical readiness versus evidence coverage, using Mistral’s October 2 public assessment and its 20 unknown checks.
mistral.ai · observed 2026-10-02T18:18:25.625Z

What would you need to verify after a 100/100 public scan?
Two numbers answering different questions
The saved Mistral homepage assessment returned a technical-readiness score of 100/100 and high observation confidence. The public report observed the homepage, resolved crawler-policy information and found discovery and metadata signals used by the bounded technical model. It also reported 58% evidence coverage: 28 known checks out of 48 applicable checks.
That is not an inconsistency to smooth away. The technical score asks how the site performed on the public model’s checks. Evidence coverage asks how much of the broader machine-counterparty assessment is supported. A reader needs both before interpreting the headline.
Do not turn an unknown into a failed check
Twenty checks remained unknown in this snapshot. The appropriate response is to say what evidence is missing, not to infer that Mistral lacks the corresponding capabilities. A public homepage probe does not operate an authenticated account, authorize a paid call, inspect internal policy or reconcile a receipt.
For the same reason, it would be wrong to advertise this as an end-to-end agent transaction test. The scan does not demonstrate that an arbitrary agent can select a service, authenticate, accept a quote, complete work and safely recover from a failure. Those are separate tasks with separate authorization and evidence requirements.
Turn the report into an evidence plan
An operator can map the unknown checks to the interfaces they actually provide. For authorization, identify the public description of required credentials and permissions. For commercial actions, expose the unit of work and the applicable pricing rule. For reconciliation, document which operation or receipt identifier a client can retain. Each item needs an observed or declared source appropriate to the claim.
Do not add an interface purely to satisfy a generic checklist. Start with the tasks a real customer or integration should perform. A useful action plan says which task is affected, what the public evidence currently establishes, and how the next observation would change the assessment.
What would make the next report more useful?
The next useful comparison is not simply a second high score. It is a versioned account of which evidence was added, removed or contradicted. Keep the original observation date and compare the same hostname and method. Distinguish an improved declaration from demonstrated execution.
Try that process on your own site. The free scan provides a public baseline. When your question concerns real machine requests rather than public readiness, inspect connected first-party evidence. That progression gives a strong public result a sensible next step without inventing a defect or a sales loss.
Find the evidence on your own site
Run a free public scan, inspect one specific finding, and identify the next useful check. No private account access is required for the public baseline.
Evidence and primary sources
Observations can change. An interstitial or truncation classification can also require collector review. Submit a reproducible correction rather than treating a snapshot as a permanent company-wide verdict.
What should the next study answer?
Share a useful lesson, tell us what remains unclear, or send a reproducible question. Questions are reviewed; they do not automatically become published claims.
Continue with a related question

We scanned 100 websites. A score was not always an answer.
99 reports, 55 numeric scores, 44 withheld scores: what a varied public-web sample teaches about evidence, access, and useful AI-readiness assessments.
Read the evidence and next step →Supabase scored 100—and the scan still hit an inspection limit
Why a homepage-truncated finding belongs beside a strong score, and how to distinguish bounded technical checks from exhaustive inspection.
Read the evidence and next step →
What a strong consumer-storefront scan can—and cannot—tell you
Paula’s Choice returned 100/100 in our public homepage assessment. The useful lesson for ecommerce is separating discovery from a verified shopping journey.
Read the evidence and next step →