Skip to main content
Glama

save_verification

Record verification verdicts for snapshot entries, stamping approved items as verified and flagging failures with notes until corrected.

Instructions

Record verify_snapshot verdicts using each entry’s kind and reviewToken. Missing tokens request a fresh review; changed or deleted entries return conflicts without being stamped. Entries judged ok are stamped verifiedAt; failures are flagged verificationFailed with your note and surface in get_context, get_snapshot, and mason_check_drift until corrected. Verdict notes are required for failures.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
dirYesAbsolute path to the project root directory
verdictsYesEntry name → verdict, exactly as returned by verify_snapshot

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed2 schema fields changedv0.19.0
    • addedInput schema / properties / verdicts / additionalProperties / properties / kind
      Added value: +{
      +  "description": "Entry kind returned by verify_snapshot; required to record a verdict",
      +  "enum": [
      +    "feature",
      +    "flow"
      +  ],
      +  "type": "string"
      +}
    • addedInput schema / properties / verdicts / additionalProperties / properties / reviewToken
      Added value: +{
      +  "description": "Copy the reviewToken from verify_snapshot after inspecting its evidence; required to record a verdict",
      +  "type": "string"
      +}
  2. Addedv0.17.0

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full burden. It discloses side effects (stamped verifiedAt, flagged verificationFailed), persistence of failures across other tools (get_context, get_snapshot, mason_check_drift), and a hard requirement (note required for failures). It also explains the effect of missing tokens. This is exemplary transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a compact paragraph of four sentences, each adding distinct value: core action, edge cases, success/failure outcomes, and a requirement. It front-loads the primary purpose and avoids redundancy. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a write operation with two parameters and nested verdicts, the description fully covers what the tool does, when to use it, edge-case behavior, downstream visibility, and parameter requirements. Without an output schema, the description sufficiently informs the agent about what happens after invocation. Nothing essential is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents all parameters. The description adds behavioral context: it explains that kind and reviewToken are essential for recording, that missing tokens trigger fresh review, and that notes are required for failures. This goes beyond schema definitions and meaningfully aids correct invocation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a clear verb+resource: 'Record verify_snapshot verdicts' and specifies the method (using each entry's kind and reviewToken). It also differentiates from siblings by naming verify_snapshot as the source and listing downstream tools that surface failures, making its unique role unmistakable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It clearly states the primary use case (recording verdicts from verify_snapshot) and explains edge-case behavior (missing tokens request fresh review, conflicts for changed/deleted entries). It doesn't explicitly name alternative tools like save_decision, but the purpose is unambiguous and the conditions are clear, so only a minor gap exists.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.