Skip to main content
Glama

Agent Output Verifier by SarnAI

Server Details

Verify an agent's output before you pay: JSON Schema + rules, signed receipts. Free during launch.

Ownership verified
Status
Healthy
Last Tested
Transport
Streamable HTTP · MCP 2025-11-25
URL
Repository
sarnai-dev/agent-output-verifier
GitHub Stars
0

TDQS

Score is being calculated.

Available Tools

2 tools
get_verification_recordA
Read-only
Inspect

Use before relying on another agent. Returns its signed verification history: verified, passed and failed counts, last seen, and a recency-weighted score with 95% confidence interval. Check a counterparty before you pay, or see the record buyers will see. Agent Scores are paid only: $0.01 per lookup.

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_idYesThe agent_id whose verification history to score (as passed to /verify/schema): the agent that produced the output checked by those calls (for example, the seller). agent_ids starting with 'test-' are reserved for testing and never have a history.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnlyHint=true, destructiveHint=false), so the bar is lower, and the description adds genuinely new behavioral context beyond them: the record is signed, the score is recency-weighted with a confidence interval, and lookups are paid at $0.01 each. The cost/payment requirement is the key operational detail an agent wouldn't otherwise know.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three front-loaded sentences with no filler: purpose first, then usage contexts, then pricing. Every sentence carries information, though the phrasing is slightly terse/jargon-dense.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only, single-parameter tool with no output schema, the description supplies the return payload (counts, last seen, score with CI) and the pricing model, which is everything an agent needs to decide and invoke correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single agent_id parameter is already richly documented in the schema (including the test- prefix reservation). The description adds no parameter-level semantics, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb+resource (returns an agent's signed verification history) and enumerates exactly what comes back: verified/passed/failed counts, last seen, and a recency-weighted score with a 95% CI. It does not explicitly differentiate itself from the sibling verify_schema, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives concrete when-to-use contexts: 'Use before relying on another agent' and 'Check a counterparty before you pay, or see the record buyers will see.' That is clear usage framing, but it never names the alternative verify_schema or states when NOT to use this tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

verify_schemaAInspect

Independent check before paying for another agent's work, or before submitting your own. Checks the output against your requirements: structure, formats, ranges, and cross-field rules such as 'line totals must equal the total.' Returns pass/fail, % of checks passed, fix hints, and an Ed25519-signed attestation and receipt you can show as evidence of why you paid or refused. $0.02 per check, small next to the payment it protects. Free during launch: 1000 checks a day (MCP and REST).

ParametersJSON Schema
NameRequiredDescriptionDefault
rulesNoOptional cross-field rules: sum_equals, gte, lte, exists_in, unique (duplicate detection). Violations are reported as informational flags unless enforce_rules is true.
boundsNoOptional per-field min/max, keyed by dotted path with '[]' array wildcard (e.g. 'items[].price'). Numbers or ISO-8601 dates. Violations are reported as informational flags unless enforce_rules is true.
detailNo'full' (default) returns everything; 'compact' returns only the summary, next actions, receipt and signature. Use compact inside verify-repair loops to save tokens. 'compact' fields: summary, next_actions, result, the receipt fields (verification_id, task_id, agent_id, output_hash, schema_hash, rules_hash, verified_at, verifier_version, enforce_rules, strict_content_check), verify and attestation - errors, errors_detail, flags, hints and the scores are left out entirely, not just emptied. Both shapes are fully signed.full
task_idYesYour own identifier for this verification. 1-200 printable characters, no control characters and no lone Unicode surrogate; a request outside that range is rejected (422), not truncated.
agent_idNoOptional label identifying the agent that produced submitted_output (for example, the seller), used for /reputation and the trust score: records and scores belong to the producer, not to whoever calls this endpoint. It is an unauthenticated label. agent_ids starting with 'test-' are reserved for testing: the verification runs normally but is never recorded, so it cannot affect any reputation or trust score. 1-200 printable characters (same rule as GET /score/{agent_id}); a longer or malformed value is rejected (422), not truncated.
enforce_rulesNoOpt-in. When true, every configured bound and consistency rule must hold, or the result is 'fail' (like strict_content_check), with an 'enforce_rules: ...' entry in errors. That includes a rule or bound that cannot run (its path is missing, a value is not comparable, a value has nothing to be compared against): it fails rather than passing silently, so omitting a field, in the whole path or in any array element the rule applies to, cannot skip a binding rule. A rule or bound with if_present: true is skipped where its field is absent instead. When false (default) nothing changes: violations and unrunnable rules stay informational flags. Echoed back in the signed response.
expected_schemaYes
submitted_outputYes
strict_content_checkNoOpt-in. When true, any string in submitted_output containing control characters (e.g. null bytes) or embedded HTML/script markup makes the result 'fail'. When false (default) this is not checked at all.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the sparse annotations by disclosing pricing ($0.02 per check), a free quota (1000 checks/day on MCP and REST), and that the response includes an Ed25519-signed attestation and receipt usable as evidence. Cost and rate-limit context are exactly the operational facts annotations cannot carry.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the decision ('Independent check before paying...'), then what is validated, then what comes back, then cost and quota. Five sentences, each carrying decision-relevant information; the pricing aside doubles as a cost-benefit cue rather than filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 9-parameter tool with no output schema, the description covers the essential gaps: the return payload (pass/fail, % passed, fix hints, signed attestation and receipt), cost, and quota. It omits mode-specific behavior like detail=compact or the enforce_rules opt-in, though those are documented in the schema itself.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 78%, so the schema already documents most parameters in depth (tolerance, if_present, enforce_rules, detail, task_id, agent_id). The description adds only a loose mapping — 'ranges' to bounds and 'cross-field rules' to rules — so the baseline 3 applies when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('checks the output against your requirements') and enumerates what is checked: structure, formats, ranges, cross-field rules. This clearly distinguishes it from the sibling get_verification_record, which retrieves past results rather than performing a new check.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly names the two triggering situations — 'before paying for another agent's work, or before submitting your own' — which is strong context for when to call it. It does not name an alternative tool or state when not to use it, keeping it just short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 2 tool updates
    • First observedget_verification_record
    • First observedverify_schema

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables AI agents to obtain an independent review of their drafts, returning a pass/fail verdict with specific issues and suggested fixes, plus a signed receipt. Payments are made per check over x402, with no account or API key needed.
    MIT
  • A
    license
    B
    quality
    C
    maintenance
    Validates agent outputs in multi-agent systems to prevent coordination failures, with tools for schema verification, hallucination detection, and freshness checks, all with zero LLM cost.
    5
    41 npm
    MIT
  • F
    license
    Not graded
    quality
    B
    maintenance
    Enables AI agents to validate deliverables (JSON, ZIP, PDF, DOCX) against signed contracts, generating verifiable receipts with Ed25519 signatures.
    -
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables deterministic verification of AI agent decisions and actions, providing PASS/FAIL/ABSTAIN verdicts with replayable proofs and an optional signed receipt ledger.
    5 npm
    Apache 2.0
Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.