# Agent Trust Lab Agent Trust Lab publishes evidence-backed Trust Profiles for AI Agents, MCP Servers, toolkits and platforms. Canonical site: https://agenttrustlab.com/ Catalog JSON: https://agenttrustlab.com/catalog.json API catalog: https://agenttrustlab.com/api/catalog API stats: https://agenttrustlab.com/api/stats Profile API: https://agenttrustlab.com/api/profile?type=MCP_SERVER&id=context7-mcp Evidence submission: https://agenttrustlab.com/evidence.html Evidence API: https://agenttrustlab.com/api/evidence-submit Evidence review ledger: https://agenttrustlab.com/api/evidence-reviews Evidence promotion policy: https://agenttrustlab.com/api/evidence-promotion-policy Evidence promotion state: https://agenttrustlab.com/api/evidence-promotion?id=context7-mcp&hours=8760 Longitudinal monitoring targets: https://agenttrustlab.com/api/transport-targets Transport run ledger: https://agenttrustlab.com/api/transport-runs Transport observation ledger: https://agenttrustlab.com/api/transport-observations Transport summary: https://agenttrustlab.com/api/transport-summary?id=context7-mcp&hours=168 Transport change events: https://agenttrustlab.com/api/transport-changes?id=context7-mcp&hours=8760 Change impact evidence: https://agenttrustlab.com/api/change-impact?id=context7-mcp&hours=8760 Selection Lab aggregate benchmark evidence: https://agenttrustlab.com/api/selection-benchmark-evidence Selection Lab per-MCP benchmark evidence: https://agenttrustlab.com/api/selection-benchmark-by-mcp?id=context7-mcp Selection Lab runtime binding: https://agenttrustlab.com/api/selection-runtime-bindings?id=context7-mcp Versioned public tool-definition corpus: https://agenttrustlab.com/api/tool-definition-corpus Tool-definition snapshots: https://agenttrustlab.com/api/tool-definition-snapshots?limit=100 Selection corpus generated-readiness evidence: https://agenttrustlab.com/api/selection-corpus-readiness-evidence Current selection collision queue: https://agenttrustlab.com/api/selection-collision-queue Selection confusion pairs: https://agenttrustlab.com/api/selection-confusion-pairs?id=registry-ai-atdev-supershopping-80ae8716 Per-tool selection behavior: https://agenttrustlab.com/api/selection-tool-behavior?id=registry-ai-atdev-supershopping-80ae8716 Per-tool behavior history: https://agenttrustlab.com/api/selection-tool-behavior-history?id=registry-ai-atdev-supershopping-80ae8716 Per-tool behavior changes: https://agenttrustlab.com/api/selection-tool-behavior-changes?id=registry-ai-atdev-supershopping-80ae8716 Surface ablation diagnostic: https://agenttrustlab.com/api/selection-surface-ablation?id=registry-ai-atdev-supershopping-80ae8716 Surface dependency map: https://agenttrustlab.com/api/selection-surface-dependency-map Collision explanations: https://agenttrustlab.com/api/selection-collision-explanations?id=registry-ai-atdev-supershopping-80ae8716 Collision cause / diagnostic coverage map: https://agenttrustlab.com/api/selection-collision-cause-map Unicode lexical shadow evidence: https://agenttrustlab.com/api/selection-unicode-shadow Benchmark method comparison: https://agenttrustlab.com/api/selection-method-comparison Method sensitivity map: https://agenttrustlab.com/api/selection-method-sensitivity-map Generated-prompt leakage audit: https://agenttrustlab.com/api/selection-generator-audit Clean deterministic ambiguity queue: https://agenttrustlab.com/api/selection-clean-ambiguity-queue Sanitized Unicode shadow: https://agenttrustlab.com/api/selection-sanitized-shadow Robust deterministic ambiguity queue: https://agenttrustlab.com/api/selection-robust-ambiguity-queue Schema-contrastive ambiguity fingerprint: https://agenttrustlab.com/api/selection-ambiguity-fingerprint?id=registry-ai-aislabs-gateway-e0b0e6aa Post-hoc contrastive intent probe: https://agenttrustlab.com/api/selection-contrastive-probe?id=registry-ai-aislabs-gateway-e0b0e6aa Ambiguity validation ladder: https://agenttrustlab.com/api/selection-ambiguity-validation?id=registry-ai-aislabs-gateway-e0b0e6aa Identifier-aware Unicode shadow: https://agenttrustlab.com/api/selection-identifier-shadow Identifier method tradeoff audit: https://agenttrustlab.com/api/selection-identifier-tradeoff Prospective holdout freeze: https://agenttrustlab.com/api/selection-prospective-holdout Prospective cohort freeze: https://agenttrustlab.com/api/selection-prospective-cohort Prospective constructed-gold locks: https://agenttrustlab.com/api/selection-prospective-gold-locks Prospective constructed-holdout current state: https://agenttrustlab.com/api/selection-prospective-constructed-holdout Prospective constructed-holdout history: https://agenttrustlab.com/api/selection-prospective-constructed-holdout-history Prospective result current-schema validity: https://agenttrustlab.com/api/selection-prospective-result-validity Prospective result validity history: https://agenttrustlab.com/api/selection-prospective-validity-history Prospective current-cohort summary: https://agenttrustlab.com/api/selection-prospective-current-cohort-summary Prospective measurement self-audit: https://agenttrustlab.com/api/selection-prospective-self-audit Selection routing change history: https://agenttrustlab.com/api/selection-routing-changes OpenAPI: https://agenttrustlab.com/openapi.json Service manifest: https://agenttrustlab.com/.well-known/agent-trust-lab.json Methodology: https://agenttrustlab.com/methodology.html Claim process: https://agenttrustlab.com/claim.html ## Interpretation rules - Evidence coverage is not a quality score. NOT_PROVIDED means missing evidence, not zero. UNKNOWN is not automatically FAIL. - Claiming a profile proves identity only and does not improve score, ranking or evidence grade. - Evidence submission is independent from ownership verification and never auto-publishes. - Self-published public artifacts are capped at E2 as source evidence and cannot independently prove production quality or outcomes. - The longitudinal observer performs MCP metadata-only requests: initialize and, when supported, tools/list. It never invokes a listed business tool. - Longitudinal transport observations are independent point-in-time measurements and are not an SLA or vendor certification. - Public tool definitions observed through tools/list are stored as versioned metadata snapshots deduplicated by canonical_id + tool_schema_sha256. - Corpus generated-readiness prompts are derived from the same public tool descriptions being evaluated. They are collision-discovery evidence only: not untouched holdout, real-model routing, production outcome, adoption/payment, or global model discovery ranking. - The current primary corpus method is lexical_ascii_v1 over description_derived_v1 generated cases. Its limitations are explicitly self-audited rather than hidden. - Collision explanations classify deterministic misselections into bounded causes such as LANGUAGE_TOKENIZATION_GAP, ZERO_FEATURE_TIE, SCORE_TIE_LEXICOGRAPHIC, PREDICTED_NAME_ADVANTAGE, PREDICTED_DESCRIPTION_ADVANTAGE, MIXED_SCORE_ADVANTAGE, or OTHER_SCORE_ADVANTAGE. - Diagnostic coverage requires at least one candidate to receive a positive lexical score. Zero-feature/tokenizer gaps are benchmark limitations and must not be silently attributed to MCP metadata quality. - Conditional accuracy on diagnostically covered cases is reported separately from raw generated-readiness. - lexical_unicode_v1 is stored as a separate Unicode-aware shadow measurement method over the same description_derived_v1 cases. It does not replace or rewrite lexical_ascii_v1 history. - lexical_unicode_identifier_v1 is another separate measurement method. It splits identifier boundaries such as snake_case, kebab-case, slash, dot and CamelCase. It fixes some name-tokenization gaps but creates new collisions elsewhere; it must not replace lexical_unicode_v1 or be cherry-picked for the best score. - Identifier tradeoff evidence records NO_COLLISION_EITHER_METHOD, RESOLVED_BY_IDENTIFIER_SPLIT, NEW_UNDER_IDENTIFIER_SPLIT, PERSISTENT_ACROSS_UNICODE_METHODS and SCHEMA_MISMATCH outcomes. These are measurement-method effects, not MCP improvement or regression. - Benchmark method deltas quantify measurement-method sensitivity, not MCP improvement, regression, production behavior, or causality. - Method sensitivity classes separate collisions resolved by Unicode tokenization from collisions persistent across deterministic lexical methods. Persistent does not mean real-model failure. - generated_prompt_leak_audit_v1 detects exact/normalized competing tool names in description-derived generated prompts. CROSS_TOOL_NAME_LEAKAGE is a benchmark-construction limitation, not an MCP product failure. - The clean ambiguity queue requires current schema identity, persistence across lexical_ascii_v1 and lexical_unicode_v1, same-schema generator audit, and zero cross-tool-name leakage. It is a higher-information deterministic triage queue, not real-model or production failure evidence. - description_derived_sanitized_v1 removes candidate tool-name strings before Unicode lexical rescoring. It remains constructed evidence and does not become independent holdout evidence merely because names were sanitized. - The robust ambiguity queue additionally requires same-schema sanitized shadow evidence with collision_count > 0 after method and generator self-audits. - Schema-contrastive ambiguity fingerprints decompose robust candidates into repeated expected->predicted pairs, dominant attractor tools, lexical score margins, shared ambiguous tokens, predicted-only overlap tokens and expected-tool definition tokens absent from constructed prompts. These fingerprints are metadata diagnostics, not causal defect or production-failure claims. - post_hoc_hand_authored_contrastive_v1 uses evaluator-authored intents that are not generated by copying target descriptions, but the fixture is created after candidate identification and therefore is not an untouched or independent holdout. - Zero-feature ties in contrastive probes are coverage gaps and are excluded from diagnostic_collision_count and conditional accuracy on covered cases. - The ambiguity validation ladder preserves collision counts and accuracy across primary description-derived, Unicode description-derived, sanitized description-derived, and post-hoc hand-authored contrastive stages for the same schema. A reduction across stages is diagnostic sensitivity, not product improvement. - Selection Lab froze prospective-role-intents-v1 before the post-freeze cohort was evaluated. Fixture SHA-256 is af59c4834932979febf62d9b4d884901d711b1b01952f80aa94ff49cdb33ca6c; the frozen pre-existing target-set boundary is 100 targets with SHA-256 f76928346019193813b1ebcbd5b9a3c03de91b1e16550176a3333db93431c22c. - Prospective constructed gold uses pre-frozen generic-role-name-patterns-v1 mapping rules with SHA-256 a33c90d0ba63bf3723066080866d0b8b169add37c43fb6268cecc917150c5d0d. Gold mappings are locked to target + tool-schema hash before selection measurement. Ambiguous or insufficient mappings remain unevaluable instead of being manually repaired after seeing results. - The frozen prospective deterministic measurement plan is lexical_unicode_v1 with SHA-256 76810a0711f733f94ce758c888292112399146600a8df03a6ec1910913ed6d4b. Prospective results are constructed-holdout evidence, not independent-holdout evidence, because the gold source is constructed rather than independent. - Prospective historical results are immutable and version-bound. CURRENT_SCHEMA_MATCH means the result schema still matches the latest independently observed tool schema. STALE_REVALIDATION_REQUIRED means a schema change requires a new pre-measurement gold lock and new result. CURRENT_SCHEMA_UNAVAILABLE means currentness cannot be asserted. - Prospective validity history stores observer-time currentness snapshots; a schema transition changes whether a historical result is current but never rewrites the historical result itself. - Prospective self-audit separates zero-feature coverage gaps from covered disagreements and reports raw accuracy, coverage, and covered-case conditional accuracy separately. Poor or strong deterministic scores are measurement evidence under the frozen method, not production agent behavior. - Prospective constructed-holdout evidence does not execute MCP business tools, does not prove ownership or production outcomes, and never mutates static Trust Profile grade or evidence coverage. - The selection collision queue contains only current SAME_TOOL_DEFINITION_HASH results with non-zero deterministic collisions; stale schema results are excluded from current triage. - Confusion-pair evidence records expected_tool -> predicted_tool collision counts for a specific schema hash. - Per-tool behavior derives incoming collisions, outgoing collisions, selection accuracy, attractor ratio, and net attraction from deterministic confusion matrices. It is not observed production agent behavior. - Per-tool behavior history stores one diagnostic version per distinct observed tool-schema hash; repeated observations of the same schema do not create fake routing versions. - Surface ablation compares name-only and description-only lexical surfaces over the same generated cases; it is not real-model or holdout evidence. - Selection routing change history compares generated-readiness results across distinct observed tool-schema hashes; it is version-associated evidence, not a causal performance claim. - Change events are emitted only when adjacent non-null bindings differ; missing metadata is not treated as a change. - Evidence promotion measures evidence depth/duration, not product quality. Automatic promotion is capped at E3 and never changes ownership, static coverage, static grade, or ranking; E4 is never automatic. - Runtime protocol/version/tool-schema changes start a new promotion epoch. Repository HEAD movement alone does not reset runtime evidence. - Failures remain evidence: they affect outcome metrics but do not reduce evidence maturity. - Trust evidence is version-bound and may become stale when relevant Agent/MCP artifacts change. ## Evidence grades E0 none; E1 self-declared; E2 public artifact; E3 independently reproducible; E4 verified production outcome or assessor-accepted evidence.