Skip to main content
Glama

Agent Output Verifier

Server Details

Independently verify an agent's output before you pay.

If you are the author of this connector, you can claim ownership by verifying the domain or GitHub account it belongs to. Claimed connector authors can inspect health checks, view analytics, and manage their listing.
Status
Healthy
Last Tested
Transport
Streamable HTTP · MCP 2025-11-25
URL
Repository
mglbagi/agent-output-verifier
GitHub Stars
0

TDQS

A4.4/5.0

Scored across 2 tools

Disambiguation5/5

get_verification_record retrieves an agent's reputation history, while verify_schema performs an independent check of a specific output. The two tools target clearly distinct actions and resources, with no overlap in purpose.

Naming Consistency5/5

Both tool names follow a consistent verb_noun snake_case pattern: get_verification_record and verify_schema. The verbs are appropriate and the naming is predictable throughout.

Tool Count3/5

With only two tools, the set feels thin for a server named Agent Output Verifier. While each tool serves a distinct function, the low count is borderline and limits the server's surface area.

Completeness4/5

The core verification lifecycle is covered: verify_schema checks output and returns an attestation, and get_verification_record provides aggregated reputation data. However, there is no way to retrieve individual past attestations or records, which is a minor gap agents can work around.

Available Tools

2 tools
get_verification_recordA
Read-only
Inspect

Use before relying on another agent. Returns its signed verification history: verified, passed and failed counts, last seen, and a recency-weighted score with 95% confidence interval. Check a counterparty before you pay, or see the record buyers will see. $0.01 per lookup. 3 free calls/day.

ParametersJSON Schema
NameRequiredDescriptionDefault
agent_idYesThe agent_id whose verification history to score (as passed to /verify/schema): the agent that produced the output checked by those calls (for example, the seller). agent_ids starting with 'test-' are reserved for testing and never have a history.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered. The description adds genuinely new operational context the annotations lack: '$0.01 per lookup. 3 free calls/day,' which tells the agent about cost and rate limits before invoking. It does not contradict any annotation, though idempotentHint=false on a read tool is unexplained.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences: purpose and return payload first, then the decision context, then cost. Every sentence earns its place and the most important routing information is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, but the description compensates by enumerating the returned fields (counts, last seen, weighted score, CI). Combined with pricing and the usage trigger, an agent has nearly everything needed to call and interpret the result; only score semantics/thresholds are left implicit.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With a single parameter at 100% schema description coverage, the schema already carries semantics, including the 'test-' prefix reservation and what the agent_id refers to. The description adds nothing about agent_id, so this sits at the baseline for schema-documented params.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb (returns) and resource (signed verification history), then enumerates the exact contents: verified, passed and failed counts, last seen, and a recency-weighted score with 95% CI. This is enough for an agent to distinguish it from the sibling verify_schema, which checks output rather than retrieving a stored trust record.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It states explicit when-to-use conditions: 'Use before relying on another agent' and 'Check a counterparty before you pay,' plus the framing 'see the record buyers will see.' That is concrete decision guidance tied to the agent's workflow rather than vague context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

verify_schemaAInspect

Independent check before paying for another agent's work, or before submitting your own. Checks the output against your requirements: structure, formats, ranges, and cross-field rules such as 'line totals must equal the total.' Returns pass/fail, % of checks passed, fix hints, and an Ed25519-signed attestation and receipt you can show as evidence of why you paid or refused. $0.02 per check, small next to the payment it protects. 3 free calls/day.

ParametersJSON Schema
NameRequiredDescriptionDefault
rulesNoOptional cross-field rules: sum_equals, gte, lte, exists_in, unique (duplicate detection). Violations are reported as informational flags unless enforce_rules is true.
boundsNoOptional per-field min/max, keyed by dotted path with '[]' array wildcard (e.g. 'items[].price'). Numbers or ISO-8601 dates. Violations are reported as informational flags unless enforce_rules is true.
detailNo'full' (default) returns everything; 'compact' returns only the summary, next actions, receipt and signature. Use compact inside verify-repair loops to save tokens. 'compact' fields: summary, next_actions, result, the receipt fields (verification_id, task_id, agent_id, output_hash, schema_hash, rules_hash, verified_at, verifier_version, enforce_rules, strict_content_check), verify and attestation - errors, errors_detail, flags, hints and the scores are left out entirely, not just emptied. Both shapes are fully signed.full
task_idYesYour own identifier for this verification. 1-200 printable characters, no control characters and no lone Unicode surrogate; a request outside that range is rejected (422), not truncated.
agent_idNoOptional label identifying the agent that produced submitted_output (for example, the seller), used for /reputation and the trust score: records and scores belong to the producer, not to whoever calls this endpoint. It is an unauthenticated label. agent_ids starting with 'test-' are reserved for testing: the verification runs normally but is never recorded, so it cannot affect any reputation or trust score. 1-200 printable characters (same rule as GET /score/{agent_id}); a longer or malformed value is rejected (422), not truncated.
enforce_rulesNoOpt-in. When true, every configured bound and consistency rule must hold, or the result is 'fail' (like strict_content_check), with an 'enforce_rules: ...' entry in errors. That includes a rule or bound that cannot run (its path is missing, a value is not comparable, a value has nothing to be compared against): it fails rather than passing silently, so omitting a field, in the whole path or in any array element the rule applies to, cannot skip a binding rule. A rule or bound with if_present: true is skipped where its field is absent instead. When false (default) nothing changes: violations and unrunnable rules stay informational flags. Echoed back in the signed response.
expected_schemaYes
submitted_outputYes
strict_content_checkNoOpt-in. When true, any string in submitted_output containing control characters (e.g. null bytes) or embedded HTML/script markup makes the result 'fail'. When false (default) this is not checked at all.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only cover open-world/idempotent/destructive hints, and the description goes well beyond them: it discloses the return payload (pass/fail, % checks passed, fix hints, Ed25519-signed attestation and receipt), the cost model ($0.02 per check), and a rate limit (3 free calls/day). The non-idempotent, billed nature is consistent with idempotentHint=false.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four sentences, each load-bearing: trigger, what is checked, what comes back, and cost/limit. It is front-loaded with the usage condition and contains no restated schema or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 9-parameter, nested-schema tool with no output schema, the description covers the return shape, evidence artifact, and pricing, which is what an agent needs to decide and call. It leaves the bounds/rules parameter surfaces entirely to the schema, which is documented well enough that this is a minor gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 78% and the nested $defs are extensively documented (tolerance, if_present, rule types, detail, enforce_rules). The description adds only the illustrative cross-field rule example ('line totals must equal the total'), not syntax or semantics beyond the schema, so it sits at the baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: it checks a submitted output against stated requirements (structure, formats, ranges, cross-field rules) and returns a signed verdict. The sibling get_verification_record is a retrieval tool, so the action boundary is clear without it being named.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit trigger conditions framed around payment: use it before paying another agent, or before submitting your own work. It does not name an alternative or when-not-to-use case (e.g. reusing a prior verification via get_verification_record), so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool update
    • Changedverify_schema1 field changed
      • addedInput schema / $defs / ConsistencyRule / properties / tolerance / description
        Added value: +"Absolute tolerance for sum_equals and the near-match check gte/lte use. The default (1e-9) suits exact arithmetic but is usually too tight for money that passed through any floating-point math (e.g. percentage discounts, tax, unit-price * quantity): 0.1 + 0.2 != 0.3 in IEEE-754 floats by about 5.5e-17, and real pricing data regularly carries larger rounding drift than that. For a field in major currency units (dollars, euros), 0.01 (one cent) is a reasonable tolerance; for minor units (cents) already stored as integers, the default is fine. Too loose a tolerance can mask a genuine mismatch, so prefer the smallest value that tolerates your data's own rounding, not a large one 'to be safe.'"
  2. 2 tool updates
    • First observedget_verification_record
    • First observedverify_schema

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables deterministic verification of AI agent decisions and actions, providing PASS/FAIL/ABSTAIN verdicts with replayable proofs and an optional signed receipt ledger.
    33 npm
    Apache 2.0
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables AI agents to obtain an independent review of their drafts, returning a pass/fail verdict with specific issues and suggested fixes, plus a signed receipt. Payments are made per check over x402, with no account or API key needed.
    MIT
  • A
    license
    A
    quality
    C
    maintenance
    Enables AI agents to dispatch human verifiers for physical world tasks like product authentication, property inspection, and document verification, returning timestamped evidence reports.
    3
    47 npm
    MIT
Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.