Skip to main content
Glama

record_verdict

Records a verdict for a claim-key pair, verifying quotes are verbatim from the source PDF before acceptance.

Instructions

Record the verdict for one (claim, key) pair. Every quote is verified against the PDF first; the verdict is rejected if any quote is not verbatim.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
keyYes
issuesNoTags such as overclaim, wrong_number, wrong_population, wrong_condition, misattribution, needs_page, outdated.
quotesNoVerbatim passages from the source (required for supported/partial/contradicted/secondary).
verdictYessupported: The source states the claim as written.; partial: The source supports part of it, or with caveats (overclaim, different number/population/condition).; contradicted: The source says the opposite, or the numbers/facts conflict.; secondary: The source only repeats a claim from another work; the original should be cited.; not_found: No relevant passage found in the available text. NOT the same as unsupported.; unavailable: The source PDF could not be obtained or has no text layer.
claim_idYes
rationaleYesOne or two sentences explaining the verdict, in the manuscript's language.
confidenceNo

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must carry the behavioral load. It does disclose a real trait – every quote is checked against the PDF and the verdict is rejected if a quote is not verbatim – which is more than the schema states. But it omits what 'rejected' looks like (error vs. silent drop), whether re-recording overwrites a prior verdict, and any permission requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, the core action front-loaded and the enforcement rule immediately after. Nothing is padding and nothing needs to be re-read.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-parameter mutation tool with no annotations, no output schema, and 57% schema coverage, the description is adequate but thin. It never says what is returned on success, whether recording is idempotent, or how claim_id/key relate to the sibling lookup tools, leaving real gaps for an agent wiring this into a workflow.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 57%, with the verdict enum, quotes, and issues well documented in the schema itself. The description's verbatim-quote rule adds meaning to the quotes parameter, but claim_id, key, and confidence remain undescribed in both places, so it only partially compensates.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb+resource+scope: 'Record the verdict for one (claim, key) pair.' That is precise enough to distinguish it from read-oriented siblings like get_claim or list_claims, though it never names a sibling explicitly to sharpen the boundary.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied rather than stated: the quote-verification precondition tells the agent when the call will succeed, but there is no explicit 'use this after X' or comparison to alternatives such as verify_quote or audit_status. An agent can infer the workflow but must supply the context itself.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.