Skip to main content
Glama

run_verifier

Run a verifier on agent output to get a pass/fail judgment with reasoning.

verifier_id resolves to an active verifier: UUID string, accessible user-scope name (owned or one unambiguous grant), or platform alias system:<name> for platform-managed judges (no criterion/config leakage via list/get).

inputs keys must match the version's input_fields exactly (no missing or extra). media_url is required for text_image and image contracts, forbidden for text. Caller-shape errors return a 400 with no row written; judge runtime errors persist a row with status="error" and an error_code of runtime_error, verifier_unavailable, or timeout.

Provenance fields (skill_id, skill_version, skill_ref, run_id) are optional and stamped onto the row. skill_id is access-checked: a skill the caller cannot see surfaces as NotFound.

Returns: {verifier_run_id, verifier_id, version, status, passed, reasoning, duration_ms, created_at} on success, plus error_code and error_message on error. passed is null on errors. On a successful run that you attributed to a skill you own, the result also carries skill_id.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
inputsYesField values for the verifier, keyed by the version's input field names (exact match, no missing or extra keys).
run_idNoOptional caller-supplied run identifier to correlate this run, for provenance.
versionNoPin to a specific version number; omit to use the latest version.
skill_idNoOptional id of the skill this run is associated with, for provenance; a skill you cannot see is rejected.
media_urlNoPublic image URL to judge; required for image and text+image verifiers, rejected for text-only ones.
skill_refNoOptional free-form skill label (slug or name) to stamp on the run, for provenance.
verifier_idYesThe verifier to act on, identified by its UUID or, for a verifier you own, its name.
skill_versionNoOptional skill version number to stamp on the run, for provenance.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed6 schema fields changed
    • addedInput schema / properties / skill_id
      Added value: +{
      +  "anyOf": [
      +    {
      +      "type": "string"
      +    },
      +    {
      +      "type": "null"
      +    }
      +  ],
      +  "default": null,
      +  "description": "Optional id of the skill this run is associated with, for provenance; a skill you cannot see is rejected."
      +}
    • addedInput schema / properties / skill_ref
      Added value: +{
      +  "anyOf": [
      +    {
      +      "type": "string"
      +    },
      +    {
      +      "type": "null"
      +    }
      +  ],
      +  "default": null,
      +  "description": "Optional free-form skill label (slug or name) to stamp on the run, for provenance."
      +}
    • addedInput schema / properties / skill_version
      Added value: +{
      +  "anyOf": [
      +    {
      +      "type": "integer"
      +    },
      +    {
      +      "type": "null"
      +    }
      +  ],
      +  "default": null,
      +  "description": "Optional skill version number to stamp on the run, for provenance."
      +}
    • removedInput schema / properties / workflow_id
      Removed value: -{
      -  "anyOf": [
      -    {
      -      "type": "string"
      -    },
      -    {
      -      "type": "null"
      -    }
      -  ],
      -  "default": null,
      -  "description": "Optional id of the workflow this run is associated with, for provenance; a workflow you cannot see is rejected."
      -}
    • removedInput schema / properties / workflow_ref
      Removed value: -{
      -  "anyOf": [
      -    {
      -      "type": "string"
      -    },
      -    {
      -      "type": "null"
      -    }
      -  ],
      -  "default": null,
      -  "description": "Optional free-form workflow label (slug or name) to stamp on the run, for provenance."
      -}
    • removedInput schema / properties / workflow_version
      Removed value: -{
      -  "anyOf": [
      -    {
      -      "type": "integer"
      -    },
      -    {
      -      "type": "null"
      -    }
      -  ],
      -  "default": null,
      -  "description": "Optional workflow version number to stamp on the run, for provenance."
      -}
  2. First observed

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description goes well beyond the annotations (readOnlyHint=false, destructiveHint=false, etc.) by disclosing side effects: successful runs persist a row, judge runtime errors persist a row with status='error', caller-shape errors return 400 with no row written, and provenance fields are stamped. It also explains access-checking behavior for skill_id. This is rich behavioral context that annotations alone cannot convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but well-organized: it front-loads the core purpose, then covers identifier resolution, input constraints, error semantics, provenance, and return shape. Every sentence earns its place. It loses one point because it is quite long and could be slightly tightened, but the density is justified given the tool's complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (8 params, nested inputs object, output schema, error states), the description is remarkably complete. It covers identifier resolution, input validation, media rules, error behavior, provenance, and return values. The output schema exists, so return values are documented, but the description still summarizes them helpfully. Nothing an agent needs to call this correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents all 8 parameters. The description adds value by explaining the verifier_id resolution rules (UUID, user-scope name, or system: alias), the exact-match requirement for inputs keys, and the media_url contract rules. It doesn't repeat every parameter but adds semantic context that the schema lacks. A 4 is appropriate because the description meaningfully supplements the schema without being redundant.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Run a verifier on agent output to get a pass/fail judgment with reasoning.' This clearly distinguishes the tool from siblings like get_verifier, list_verifiers, deploy_verifier, and revoke_verifier. The purpose is unambiguous and action-oriented.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit guidance on when to use this tool and how to avoid errors: it explains verifier_id resolution, input key matching requirements, media_url rules per contract type, and error behavior. It also implicitly distinguishes from get_verifier (which retrieves config) and deploy_verifier (which manages lifecycle). This is comprehensive usage guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.