Score interpretation

Mistral scored 100. Why was evidence coverage only 58%?

A concrete explanation of technical readiness versus evidence coverage, using Mistral’s October 2 public assessment and its 20 unknown checks.

AIWebSignals Research3 minute readPublic-assessment case study
A score can be excellent without making unobserved authorization and execution facts known. This is a timestamped public assessment, not a claim of customer adoption, traffic loss, or revenue lift.

mistral.ai · observed 2026-10-02T18:18:25.625Z

100/100Bounded technical readiness
58%Evidence coverage, not a quality grade
20Unknown checks of 48 applicable
Inspect the saved mistral.ai report · Source JSON
Actual AIWebSignals report screenshot for the saved mistral.ai observation
Actual report interface captured October 2. Resized for display; the saved evidence and scores are unchanged. Open the source report above for full context.

What would you need to verify after a 100/100 public scan?

Two numbers answering different questions

The saved Mistral homepage assessment returned a technical-readiness score of 100/100 and high observation confidence. The public report observed the homepage, resolved crawler-policy information and found discovery and metadata signals used by the bounded technical model. It also reported 58% evidence coverage: 28 known checks out of 48 applicable checks.

That is not an inconsistency to smooth away. The technical score asks how the site performed on the public model’s checks. Evidence coverage asks how much of the broader machine-counterparty assessment is supported. A reader needs both before interpreting the headline.

Do not turn an unknown into a failed check

Twenty checks remained unknown in this snapshot. The appropriate response is to say what evidence is missing, not to infer that Mistral lacks the corresponding capabilities. A public homepage probe does not operate an authenticated account, authorize a paid call, inspect internal policy or reconcile a receipt.

For the same reason, it would be wrong to advertise this as an end-to-end agent transaction test. The scan does not demonstrate that an arbitrary agent can select a service, authenticate, accept a quote, complete work and safely recover from a failure. Those are separate tasks with separate authorization and evidence requirements.

Turn the report into an evidence plan

An operator can map the unknown checks to the interfaces they actually provide. For authorization, identify the public description of required credentials and permissions. For commercial actions, expose the unit of work and the applicable pricing rule. For reconciliation, document which operation or receipt identifier a client can retain. Each item needs an observed or declared source appropriate to the claim.

Do not add an interface purely to satisfy a generic checklist. Start with the tasks a real customer or integration should perform. A useful action plan says which task is affected, what the public evidence currently establishes, and how the next observation would change the assessment.

What would make the next report more useful?

The next useful comparison is not simply a second high score. It is a versioned account of which evidence was added, removed or contradicted. Keep the original observation date and compare the same hostname and method. Distinguish an improved declaration from demonstrated execution.

Try that process on your own site. The free scan provides a public baseline. When your question concerns real machine requests rather than public readiness, inspect connected first-party evidence. That progression gives a strong public result a sensible next step without inventing a defect or a sales loss.

Find the evidence on your own site

Run a free public scan, inspect one specific finding, and identify the next useful check. No private account access is required for the public baseline.

Evidence and primary sources

  1. Mistral: saved report from October 2
  2. How this cohort was selected

Observations can change. An interstitial or truncation classification can also require collector review. Submit a reproducible correction rather than treating a snapshot as a permanent company-wide verdict.

What should the next study answer?

Share a useful lesson, tell us what remains unclear, or send a reproducible question. Questions are reviewed; they do not automatically become published claims.

Suggest a question or correction

Continue with a related question