Skip to main content
Glama

verify_replay

Read-onlyIdempotent

Check replay commitment integrity and compare recorded execution components. Version 4 includes reference context and the final result/status. A match compares commitments; this tool does not rerun the workflow or reconstruct historical reference populations, and does not prove factual correctness.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
api_keyYesGeodesicAI API key (gai_...)
contract_aYesreplay_contract object from one execution
contract_bYesreplay_contract object to compare against contract_a

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed3 schema fields changed
    • addedInput schema / properties / api_key / description
      Added value: +"GeodesicAI API key (gai_...)"
    • addedInput schema / properties / contract_a / description
      Added value: +"replay_contract object from one execution"
    • addedInput schema / properties / contract_b / description
      Added value: +"replay_contract object to compare against contract_a"
  2. Added
  3. Removed
  4. Added

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is carried by structured fields. The description adds value beyond that with the 'does not rerun the workflow or reconstruct historical reference populations' caveat, which prevents an agent from over-interpreting results as fresh executions or authoritative reconstructions. This is genuine behavioral context, though it omits authentication requirements or output behavior, which are minor given the annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is exactly three sentences, front-loaded with purpose and each sentence earns its place: scope, version/status context, and boundary caveats. No filler, no repetition of schema fields, no redundant elaboration. This is appropriately sized for the tool's complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool compares two nested replay_contract objects and has no output schema, so the description should compensate. It covers purpose and limitations well, but never states what the result looks like — a match score, a list of mismatches, or a boolean — nor what 'commitment integrity' means in terms of which components the comparison inspects. Given the deep nested parameters and the absence of an output schema, there are meaningful gaps an agent would face when interpreting the tool's result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all three parameters (api_key, contract_a, contract_b). The description contributes a slight addition by explaining that a 'match compares commitments,' which hints at the relationship between the two opaque nested contract objects. Yet it does not explain what a 'replay_contract' structurally contains or how the api_key is used, leaving the nested objects' semantics under-specified.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear purpose: 'Check replay commitment integrity and compare recorded execution components.' This is a specific verb plus resource pair that is distinguishable from most siblings. However, it relies on the domain term 'commitment integrity' without defining it, and it does not explicitly name siblings like verify_certificate or compare_semantic_equivalence that it must be separated from. The scope is clear overall, but the relationship to the most similar siblings is left to inference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides useful exclusions: it does not rerun the workflow, reconstruct historical reference populations, or prove factual correctness. This tells an agent when NOT to rely on this tool (e.g., when it needs factual proof or a fresh execution). However, it never names the alternative tools for those cases, and the 'Version 4' note mostly addresses version history rather than routing the agent to the right sibling for the covered gaps.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources