Agent Output Verifier by SarnAI
Server Details
Verify an agent's output before you pay: JSON Schema + rules, signed receipts. Free during launch.
- Status
- Healthy
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-11-25
- URL
- Repository
- sarnai-dev/agent-output-verifier
- GitHub Stars
- 0
TDQS
Score is being calculated.
Available Tools
2 toolsget_verification_recordARead-onlyInspect
Use before relying on another agent. Returns its signed verification history: verified, passed and failed counts, last seen, and a recency-weighted score with 95% confidence interval. Check a counterparty before you pay, or see the record buyers will see. Agent Scores are paid only: $0.01 per lookup.
| Name | Required | Description | Default |
|---|---|---|---|
| agent_id | Yes | The agent_id whose verification history to score (as passed to /verify/schema): the agent that produced the output checked by those calls (for example, the seller). agent_ids starting with 'test-' are reserved for testing and never have a history. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnlyHint=true, destructiveHint=false), so the bar is lower, and the description adds genuinely new behavioral context beyond them: the record is signed, the score is recency-weighted with a confidence interval, and lookups are paid at $0.01 each. The cost/payment requirement is the key operational detail an agent wouldn't otherwise know.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three front-loaded sentences with no filler: purpose first, then usage contexts, then pricing. Every sentence carries information, though the phrasing is slightly terse/jargon-dense.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only, single-parameter tool with no output schema, the description supplies the return payload (counts, last seen, score with CI) and the pricing model, which is everything an agent needs to decide and invoke correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single agent_id parameter is already richly documented in the schema (including the test- prefix reservation). The description adds no parameter-level semantics, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb+resource (returns an agent's signed verification history) and enumerates exactly what comes back: verified/passed/failed counts, last seen, and a recency-weighted score with a 95% CI. It does not explicitly differentiate itself from the sibling verify_schema, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives concrete when-to-use contexts: 'Use before relying on another agent' and 'Check a counterparty before you pay, or see the record buyers will see.' That is clear usage framing, but it never names the alternative verify_schema or states when NOT to use this tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
verify_schemaAInspect
Independent check before paying for another agent's work, or before submitting your own. Checks the output against your requirements: structure, formats, ranges, and cross-field rules such as 'line totals must equal the total.' Returns pass/fail, % of checks passed, fix hints, and an Ed25519-signed attestation and receipt you can show as evidence of why you paid or refused. $0.02 per check, small next to the payment it protects. Free during launch: 1000 checks a day (MCP and REST).
| Name | Required | Description | Default |
|---|---|---|---|
| rules | No | Optional cross-field rules: sum_equals, gte, lte, exists_in, unique (duplicate detection). Violations are reported as informational flags unless enforce_rules is true. | |
| bounds | No | Optional per-field min/max, keyed by dotted path with '[]' array wildcard (e.g. 'items[].price'). Numbers or ISO-8601 dates. Violations are reported as informational flags unless enforce_rules is true. | |
| detail | No | 'full' (default) returns everything; 'compact' returns only the summary, next actions, receipt and signature. Use compact inside verify-repair loops to save tokens. 'compact' fields: summary, next_actions, result, the receipt fields (verification_id, task_id, agent_id, output_hash, schema_hash, rules_hash, verified_at, verifier_version, enforce_rules, strict_content_check), verify and attestation - errors, errors_detail, flags, hints and the scores are left out entirely, not just emptied. Both shapes are fully signed. | full |
| task_id | Yes | Your own identifier for this verification. 1-200 printable characters, no control characters and no lone Unicode surrogate; a request outside that range is rejected (422), not truncated. | |
| agent_id | No | Optional label identifying the agent that produced submitted_output (for example, the seller), used for /reputation and the trust score: records and scores belong to the producer, not to whoever calls this endpoint. It is an unauthenticated label. agent_ids starting with 'test-' are reserved for testing: the verification runs normally but is never recorded, so it cannot affect any reputation or trust score. 1-200 printable characters (same rule as GET /score/{agent_id}); a longer or malformed value is rejected (422), not truncated. | |
| enforce_rules | No | Opt-in. When true, every configured bound and consistency rule must hold, or the result is 'fail' (like strict_content_check), with an 'enforce_rules: ...' entry in errors. That includes a rule or bound that cannot run (its path is missing, a value is not comparable, a value has nothing to be compared against): it fails rather than passing silently, so omitting a field, in the whole path or in any array element the rule applies to, cannot skip a binding rule. A rule or bound with if_present: true is skipped where its field is absent instead. When false (default) nothing changes: violations and unrunnable rules stay informational flags. Echoed back in the signed response. | |
| expected_schema | Yes | ||
| submitted_output | Yes | ||
| strict_content_check | No | Opt-in. When true, any string in submitted_output containing control characters (e.g. null bytes) or embedded HTML/script markup makes the result 'fail'. When false (default) this is not checked at all. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the sparse annotations by disclosing pricing ($0.02 per check), a free quota (1000 checks/day on MCP and REST), and that the response includes an Ed25519-signed attestation and receipt usable as evidence. Cost and rate-limit context are exactly the operational facts annotations cannot carry.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the decision ('Independent check before paying...'), then what is validated, then what comes back, then cost and quota. Five sentences, each carrying decision-relevant information; the pricing aside doubles as a cost-benefit cue rather than filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter tool with no output schema, the description covers the essential gaps: the return payload (pass/fail, % passed, fix hints, signed attestation and receipt), cost, and quota. It omits mode-specific behavior like detail=compact or the enforce_rules opt-in, though those are documented in the schema itself.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 78%, so the schema already documents most parameters in depth (tolerance, if_present, enforce_rules, detail, task_id, agent_id). The description adds only a loose mapping — 'ranges' to bounds and 'cross-field rules' to rules — so the baseline 3 applies when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('checks the output against your requirements') and enumerates what is checked: structure, formats, ranges, cross-field rules. This clearly distinguishes it from the sibling get_verification_record, which retrieves past results rather than performing a new check.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names the two triggering situations — 'before paying for another agent's work, or before submitting your own' — which is strong context for when to call it. It does not name an alternative tool or state when not to use it, keeping it just short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
- First observed
get_verification_record - First observed
verify_schema
Related MCP Connectors
Verify an agent's output before you pay: JSON Schema + rules, signed receipts. Free during launch.
Issue signed receipts for AI agent actions; verify any receipt offline - free, no account.
Paid deterministic data-quality and execution-verification tools for AI agents.
Trust Layer for AI work. Give your AI hands. Every run returns a verifiable receipt. No key needed.
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceEnables AI agents to obtain an independent review of their drafts, returning a pass/fail verdict with specific issues and suggested fixes, plus a signed receipt. Payments are made per check over x402, with no account or API key needed.MIT
- AlicenseBqualityCmaintenanceValidates agent outputs in multi-agent systems to prevent coordination failures, with tools for schema verification, hallucination detection, and freshness checks, all with zero LLM cost.541 npmMIT
- FlicenseNot gradedqualityBmaintenanceEnables AI agents to validate deliverables (JSON, ZIP, PDF, DOCX) against signed contracts, generating verifiable receipts with Ed25519 signatures.-
- AlicenseNot gradedqualityBmaintenanceEnables deterministic verification of AI agent decisions and actions, providing PASS/FAIL/ABSTAIN verdicts with replayable proofs and an optional signed receipt ledger.5 npmApache 2.0
Glama MCP Gateway
Add one secure layer between your agents and this server.