Methodology v1

Naming, metrics & evidence standard.

The standard keeps Agents, MCP Servers and toolkits distinct while making their evidence comparable at the field level.

1. Subject types

AGENTAutonomous or semi-autonomous system that plans and/or acts.
MCP_SERVERCapability or tool service exposed through MCP. It is not automatically an Agent.
TOOLKITDeveloper, benchmark or assurance tooling.
PLATFORMMulti-capability platform that may contain Agents, MCP Servers or toolkits.

2. Canonical public locations

/agent/<id>.html · /mcp/<id>.html · /toolkit/<id>.html · /platform/<id>.html

An MCP endpoint is an interface attribute. It does not determine the subject type.

3. Statistics

PercentageMachine value is ratio 0–1; public display is 0–100% with one decimal by default.
Sample contextNumerator / denominator or sample size must be shown when available.
CountsIntegers.
LatencyMilliseconds.
Fiat costUSD unless another asset or currency is explicitly named.
Missing evidenceNOT_PROVIDED; never converted to zero.
Unknown resultUNKNOWN; never silently converted to FAIL.

4. Evidence grades

E0No evidence.
E1Self-declared evidence.
E2Public artifact with traceable source.
E3Independently reproducible evidence.
E4Verified production outcome or assessor-accepted evidence.

5. Trust Profile required fields

Every profile should expose: subject type, canonical ID, version or observation label, claim status, evidence coverage, evidence freshness, performance/benchmark evidence when applicable, production evidence when applicable, autonomy/governance evidence when applicable, source references and observation date.

Evidence coverage measures how many expected evidence fields are supported. It is not an overall quality score.

6. Claiming and commercial neutrality

A vendor may claim a profile by proving control of a canonical repository or website and may submit additional evidence. Claiming a profile changes identity status only. It does not improve benchmark results, evidence grade, trust status or ranking.

Payment buys verification services, not a better score.

7. Version binding and staleness

Results are bound to the version or observed surface that produced them. If a prompt, model configuration, MCP tool schema, permission surface or other dependency changes, prior assurance may become stale and should be marked for retest, reapproval or review rather than carried forward silently.

8. Evidence intake and review

Public evidence submission creates a review candidate only. It never proves ownership, never auto-publishes a profile change, and never automatically raises an evidence grade or ranking.

Self-published public artifactMay support only the specific public claims it actually documents and is capped at E2 as a source. Product claims in a README are not independent proof of production quality, freshness or outcomes.
Independent reproducible evidenceMay support E3 when a third party can reproduce the relevant result with a documented method.
Verified production evidenceMay support E4 only when the production outcome or assessor acceptance is itself verifiable.

Every approved evidence review records both approved claims and explicit not-supported boundaries. Evidence review is prohibited from changing ownership claim status; identity control is handled only by the separate claim protocol.