GAIP Witness Agent
Server Details
Records delivery checks and observes public results.
- Status
- Healthy
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-06-18
- URL
TDQS
Scored across 2 tools
The two tools have clearly distinct actions: one registers caller-declared checks, the other compares an assertion against a public replay. Although both share the witness domain, their verbs and purposes are unambiguous.
Both names use snake_case with a consistent 'witness_' prefix. However, 'witness_delivery' is a noun phrase while 'witness_register_terms' follows a verb-noun structure, a minor deviation from a strict pattern.
Two tools is thin for a server with a stated purpose of witnessing assertions and registering terms. While possibly intentionally minimal, it sits at the borderline of being under-scoped.
The core actions of registering checks and comparing a delivery are present, but there is no way to list, update, or revoke registered terms, and no tool to retrieve prior witness records. These are notable lifecycle gaps.
Available Tools
2 toolswitness_deliveryCheck public delivery evidenceBRead-onlyInspect
Compare a registered caller assertion against a later public replay or submitted public response. Unknown when evidence is unavailable; no original transaction or seller agreement proof.
| Name | Required | Description | Default |
|---|---|---|---|
| body | No | ||
| mode | Yes | ||
| terms_id | Yes | ||
| synthetic | No | ||
| episode_id | No | ||
| body_sha256 | No | ||
| http_status | No | ||
| content_type | No | ||
| continuity_handle | No | ||
| data_classification | Yes | ||
| independent_operator_id | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds context beyond the readOnlyHint and openWorldHint by stating that results are 'Unknown when evidence is unavailable' and explicitly disclaiming proof of original transaction or seller agreement. This helps set expectations about what the tool does not provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, compact sentence that front-loads the core action and includes a meaningful caveat. It is not verbose, but it omits essential operational details, making it efficient yet incomplete.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (11 parameters, nested objects, no output schema), the description is severely under-specified. It fails to explain return values, parameter usage, or when to invoke this tool, leaving the agent with insufficient context for correct usage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage and no mention of any of the 11 parameters in the description, the agent has no guidance on how to populate fields like mode, terms_id, or continuity_handle. The description does not compensate for the lack of schema documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific action (compare) with a specific resource (registered caller assertion vs later public replay or submitted response). It also notes limitations ('no original transaction or seller agreement proof'), distinguishing it from sibling tools like check_agent_conformance or verify_gaip_receipt, which target different evidence types.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage when comparing an assertion to public evidence, but does not explicitly state when to prefer this tool over siblings or when not to use it. No alternatives are mentioned, leaving the agent to infer from the tool name and description alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
witness_register_termsAInspect
Register caller-declared checks for a public HTTPS service. This does not verify seller agreement or payment.
| Name | Required | Description | Default |
|---|---|---|---|
| checks | Yes | ||
| seller_ref | Yes | ||
| service_url | Yes | ||
| data_classification | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With all annotations false, the description carries the burden of behavioral disclosure. It honestly discloses that checks are caller-declared and that no seller agreement or payment verification is performed, which is useful, but it is silent on side effects, persistence, idempotency, or return behavior. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two concise sentences with the core action front-loaded and a useful caveat in the second sentence. There is no filler or repetition of schema information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema and no parameter-level documentation. The description conveys general intent but leaves the agent guessing about what a 'check' is, what the tool returns, and what seller_ref and data_classification represent, making correct invocation uncertain.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description only hints at service_url ('public HTTPS service') and checks ('caller-declared'). seller_ref and data_classification are unexplained, and the checks array lacks any item structure. The description does not sufficiently compensate for the missing schema documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Register caller-declared checks for a public HTTPS service.' The caveat that it does not verify seller agreement or payment clearly separates it from verification-focused siblings such as assurance_verify and gaip_verify_supplier_claims.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool should be used for registering supplied checks and explicitly notes what it does not do, but it never names alternatives or states explicit conditions for choosing this tool over the assurance/verification siblings. Usage guidance is mostly implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
- First observed
witness_delivery - First observed
witness_register_terms
Related MCP Connectors
Checks public counterparty evidence and GAIP receipts.
21Records buyer requests and checks public dispositions.
31Avoid redundant expensive validation. CHECK recent evidence; OBSERVE fresh independent results.
Verify artifacts, hand off public state, and record, check and find evidence about capabilities.
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceProvides tools to create, run, and verify deterministic delivery receipts for completion claims, enabling reviewers to check freshness and correctness of evidence.MIT
- FlicenseNot gradedqualityBmaintenanceEnables auditing x402 payment logs against delivery logs to issue signed proof-of-delivery receipts and verify payer spend health.-

EVIDIQ Rubric MCPofficial
AlicenseNot gradedqualityBmaintenanceDetermines whether a deliverable meets its contract using deterministic rules, criteria, and signed attestations.1MIT- AlicenseNot gradedqualityCmaintenanceSLA observation broker for the A2A network: agents register public health endpoints with target uptime and latency, and the shim probes them to record breaches.MIT
Glama MCP Gateway
Add one secure layer between your agents and this server.