agiscorecard-web3-evidence
Server Details
Review agent task evidence with sample deduplication, version filters and uncertainty intervals.
- Status
- Healthy
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-11-25
- URL
- Repository
- f-tiger/agi-site
- GitHub Stars
- 0
TDQS
Score is being calculated.
Available Tools
3 toolsfetchRead a Web3 method and sourcesRead-onlyIdempotentInspect
Retrieve a public tool reference by the ID returned from search. Contains methodology, limits, official source links and fictional worked examples; no user records. Cite the returned canonical URL.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| id | Yes | |
| url | Yes | |
| text | Yes | |
| title | Yes | |
| metadata | Yes |
review_agent_evidenceAgent EvidenceARead-onlyIdempotentInspect
Agent Evidence reviews imported task outcomes within one task, version and date window. It deduplicates sample labels and shows failures, exclusions and uncertainty. It helps review a supplier comparison; it does not run evaluations or verify reviewer identities. No public reputation ranking, identity verification, calibrated prediction or Sybil detection. Retrieve evidence with fetch or read its example resource to obtain exact inputs. Parameters are processed remotely without application persistence.
| Name | Required | Description | Default |
|---|---|---|---|
| asOf | Yes | Review date (UTC) (preserve source text; decimal amounts must stay strings) | |
| task | Yes | Task (preserve source text; decimal amounts must stay strings) | |
| records | Yes | Records | |
| version | Yes | Version (preserve source text; decimal amounts must stay strings) | |
| maxAgeDays | Yes | Maximum evidence age (days) |
Output Schema
| Name | Required | Description |
|---|---|---|
| report | Yes | |
| toolId | Yes | |
| version | Yes | |
| citation | Yes | |
| revision | Yes | |
| processing | Yes | |
| limitations | Yes | |
| evidenceStatus | Yes | |
| officialReferences | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered structurally. The description adds meaningful behavioral context beyond that: parameters are processed remotely without application persistence, sample labels are deduplicated, and outputs cover failures/exclusions/uncertainty. It also discloses non-capabilities such as no reputation ranking, identity verification, or Sybil detection. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core purpose and scope are front-loaded in the first sentence, with exclusions and usage pointers following. Nearly every sentence earns its place, though the negative list ('No public reputation ranking, identity verification, calibrated prediction or Sybil detection') partially overlaps with the earlier exclusion about not verifying reviewer identities. Minor redundancy keeps it from a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 5 required parameters, full schema coverage, and an output schema, the description covers the scope, intended review context, processing semantics, persistence behavior, and key limitations. The output schema presumably handles return-value details, so those need not be restated. The only minor gap is the ambiguous 'read its example resource' instruction, which could be more precise.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 100% description coverage with per-parameter descriptions, so the baseline is 3. The description reinforces the task/version/date-window framing that maps to 'task', 'version', 'asOf', and 'maxAgeDays', but it does not add type, format, or constraint details beyond the schema. It earns the baseline but not more, since the schema already carries the parameter semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: it 'reviews imported task outcomes' scoped to 'one task, version and date window,' and enumerates concrete behaviors such as deduplicating sample labels and showing failures, exclusions, and uncertainty. It also differentiates itself from evaluation and identity-verification tools with explicit negative scope. This goes well beyond the title 'Agent Evidence' and gives an agent a clear operational identity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly frames the intended use case ('helps review a supplier comparison') and states what it does not do: run evaluations or verify reviewer identities. It also directs the agent to 'fetch' or 'read its example resource' for exact inputs. It stops short of naming specific sibling alternatives for excluded use cases, so the when-not-to-use guidance is clear but not fully mapped to alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
searchFind a Web3 worksheetARead-onlyIdempotentInspect
Search this server’s public AI/Web3 tools and citation pages. English and Chinese names supported. Empty query lists the available worksheets.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| results | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish read-only, idempotent, non-destructive behavior, and the description adds valuable behavioral context: bilingual (English and Chinese) name support, public scope, and the empty-query listing behavior. These details go beyond what annotations or schema provide without contradicting them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two clean sentences with no filler. The primary action is front-loaded, and the supporting usage details are kept minimal and relevant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter search tool with annotations and an output schema, the description is complete: it defines the search scope, supported query forms, and the empty-query behavior. No essential information for invoking the tool correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description carries the burden of explaining the query parameter. It clarifies that queries are names, that English and Chinese are accepted, and that an empty query has special meaning. It does not fully specify matching behavior (e.g., exact vs partial), but it compensates well for the schema gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Search') and resource ('this server’s public AI/Web3 tools and citation pages'), making the tool’s function immediately clear. It is distinct from all sibling tools, which focus on checking permissions, comparing costs, or reading profiles rather than searching a corpus. The title and description align well.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly indicates when to use the tool: when searching for AI/Web3 tools or citation pages on this server. The note about empty query listing available worksheets provides a concrete usage scenario. It does not explicitly name alternatives or exclusion conditions, so it misses the top score.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
- First observed
fetch - First observed
review_agent_evidence - First observed
search
Related MCP Connectors
Evidence-bound second-opinion audit of an agent conclusion against caller-supplied evidence.
Independent AI-agent reviews: trust checks, evidence scorecards, incident registry, recommendations.
Evidence infrastructure for agents: source-backed company verification and beta import assessment.
Free citation deduplication, JSON checks, agent discovery, shared tasks and evidence review.
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceEnables AI agents to run evidence QC gates mid-workflow—checking claims, sources, prose, and vault hygiene deterministically before research ships.MIT
- AlicenseNot gradedqualityBmaintenanceProvides tools for agents to manage a local review graph, tracking acceptance behaviors, evidence, review passes, and human waivers to decouple review convergence from shipping readiness.3MIT
- AlicenseNot gradedqualityBmaintenanceEnables evidence-first regression testing for AI agents by turning production traces into reviewed, replayable cases that gate releases. It supports reproducible, auditable agent evaluation with controlled tool execution, evidence-based judging, and versioned quality gates.MIT
- FlicenseNot gradedqualityNot gradedmaintenanceA bounded evidence review engine that ingests documents, extracts evidence for a given claim, detects contradictions, and produces auditable evidence packets without hallucinations or open-web research.-
Glama MCP Gateway
Add one secure layer between your agents and this server.