Evidence methodology
How MCP tool-selection evidence is measured.
Agent Trust Lab reports measured historical selection reliability from frozen benchmark evidence. The method is designed to make the measured surface inspectable, not to certify an MCP or predict every future model decision.
The measurement sequence
- Define the candidate tools. Tool schemas and the intended selection targets are recorded for the evaluation surface.
- Freeze the benchmark. Cases, expected tools, identities, and relevant measurement settings are fixed before model measurement.
- Measure first-tool selection. The result concerns which tool the model selected first; target business tools are not executed as part of the selection measurement.
- Separate outcomes. Strict accuracy, Wilson confidence intervals where applicable, provider-valid coverage, infrastructure failures, and semantic misses remain distinct.
- Preserve the receipt. The report keeps provenance and immutable measurement details so a reader can inspect what was measured.
S1 and S2 evidence
S1 is selection-tested evidence on a frozen constructed benchmark. S2-backed records use independently sourced external gold cases. Independent batches remain separate; the site does not pool them into a universal headline score.
Observed misses matter
When a model selects the wrong tool, an adjudicated semantic miss remains attached to its benchmark, batch, and case. Provider failures are not silently counted as semantic misses, and misses are not removed to improve a result.
Version-bound source snapshots
A page labeled Version-bound source snapshot describes a source version captured for measurement. It does not claim that the current live endpoint has the same tools or current behavior.
What this does not prove
This evidence does not certify security, business correctness, uptime, regulatory compliance, safety, future model performance, or overall MCP quality. It is selection evidence under stated frozen conditions.