Skip to main content
Glama

Agent Output Verifier

verify_schema

Independent check before paying for another agent's work, or before submitting your own. Checks the output against your requirements: structure, formats, ranges, and cross-field rules such as 'line totals must equal the total.' Returns pass/fail, % of checks passed, fix hints, and an Ed25519-signed attestation and receipt you can show as evidence of why you paid or refused. $0.02 per check, small next to the payment it protects. 3 free calls/day.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
rulesNoOptional cross-field rules: sum_equals, gte, lte, exists_in, unique (duplicate detection). Violations are reported as informational flags unless enforce_rules is true.
boundsNoOptional per-field min/max, keyed by dotted path with '[]' array wildcard (e.g. 'items[].price'). Numbers or ISO-8601 dates. Violations are reported as informational flags unless enforce_rules is true.
detailNo'full' (default) returns everything; 'compact' returns only the summary, next actions, receipt and signature. Use compact inside verify-repair loops to save tokens. 'compact' fields: summary, next_actions, result, the receipt fields (verification_id, task_id, agent_id, output_hash, schema_hash, rules_hash, verified_at, verifier_version, enforce_rules, strict_content_check), verify and attestation - errors, errors_detail, flags, hints and the scores are left out entirely, not just emptied. Both shapes are fully signed.full
task_idYesYour own identifier for this verification. 1-200 printable characters, no control characters and no lone Unicode surrogate; a request outside that range is rejected (422), not truncated.
agent_idNoOptional label identifying the agent that produced submitted_output (for example, the seller), used for /reputation and the trust score: records and scores belong to the producer, not to whoever calls this endpoint. It is an unauthenticated label. agent_ids starting with 'test-' are reserved for testing: the verification runs normally but is never recorded, so it cannot affect any reputation or trust score. 1-200 printable characters (same rule as GET /score/{agent_id}); a longer or malformed value is rejected (422), not truncated.
enforce_rulesNoOpt-in. When true, every configured bound and consistency rule must hold, or the result is 'fail' (like strict_content_check), with an 'enforce_rules: ...' entry in errors. That includes a rule or bound that cannot run (its path is missing, a value is not comparable, a value has nothing to be compared against): it fails rather than passing silently, so omitting a field, in the whole path or in any array element the rule applies to, cannot skip a binding rule. A rule or bound with if_present: true is skipped where its field is absent instead. When false (default) nothing changes: violations and unrunnable rules stay informational flags. Echoed back in the signed response.
expected_schemaYes
submitted_outputYes
strict_content_checkNoOpt-in. When true, any string in submitted_output containing control characters (e.g. null bytes) or embedded HTML/script markup makes the result 'fail'. When false (default) this is not checked at all.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed1 schema field changed
    • addedInput schema / $defs / ConsistencyRule / properties / tolerance / description
      Added value: +"Absolute tolerance for sum_equals and the near-match check gte/lte use. The default (1e-9) suits exact arithmetic but is usually too tight for money that passed through any floating-point math (e.g. percentage discounts, tax, unit-price * quantity): 0.1 + 0.2 != 0.3 in IEEE-754 floats by about 5.5e-17, and real pricing data regularly carries larger rounding drift than that. For a field in major currency units (dollars, euros), 0.01 (one cent) is a reasonable tolerance; for minor units (cents) already stored as integers, the default is fine. Too loose a tolerance can mask a genuine mismatch, so prefer the smallest value that tolerates your data's own rounding, not a large one 'to be safe.'"
  2. First observed

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only cover open-world/idempotent/destructive hints, and the description goes well beyond them: it discloses the return payload (pass/fail, % checks passed, fix hints, Ed25519-signed attestation and receipt), the cost model ($0.02 per check), and a rate limit (3 free calls/day). The non-idempotent, billed nature is consistent with idempotentHint=false.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four sentences, each load-bearing: trigger, what is checked, what comes back, and cost/limit. It is front-loaded with the usage condition and contains no restated schema or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 9-parameter, nested-schema tool with no output schema, the description covers the return shape, evidence artifact, and pricing, which is what an agent needs to decide and call. It leaves the bounds/rules parameter surfaces entirely to the schema, which is documented well enough that this is a minor gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 78% and the nested $defs are extensively documented (tolerance, if_present, rule types, detail, enforce_rules). The description adds only the illustrative cross-field rule example ('line totals must equal the total'), not syntax or semantics beyond the schema, so it sits at the baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: it checks a submitted output against stated requirements (structure, formats, ranges, cross-field rules) and returns a signed verdict. The sibling get_verification_record is a retrieval tool, so the action boundary is clear without it being named.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit trigger conditions framed around payment: use it before paying another agent, or before submitting your own work. It does not name an alternative or when-not-to-use case (e.g. reusing a prior verification via get_verification_record), so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.