Skip to main content
Glama

Verify claims against evidence

jev_verify

Verify claims against cited evidence and return per-claim verdicts, probability distributions, and confidence, flagging whether each verdict stands automatically or needs human review.

Instructions

Check each claim against provided evidence text with TypeSafe Jev. Returns per claim: verdict (verified | contradicted | unsupported), full probability distribution, confidence, and whether the verdict stands on its own (auto) or needs human review. Pattern: docs.typesafe.ai/cookbooks/citation_check. Pass reports, PR descriptions, or agent briefs as claims and their cited sources, diffs, or documents as evidence.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
claimsYesClaims to verify, e.g. individual factual statements from a report.
evidenceYes
subject_atNoA contradiction stands only when same_subject reaches this probability. Default 0.5.
auto_acceptNoVerdicts at or above this confidence stand automatically; below it they are flagged 'review'. Default 0.8.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed1 schema field changedv0.13.0
    • addedInput schema / properties / subject_at
      Added value: +{
      +  "description": "A contradiction stands only when same_subject reaches this probability. Default 0.5.",
      +  "maximum": 1,
      +  "minimum": 0,
      +  "type": "number"
      +}
  2. Changed1 schema field changedv0.10.1
    • changedInput schema / $schema
      Previous value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
  3. First observedv0.1.0

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and largely succeeds: it discloses the exact verdict vocabulary, that a full probability distribution and confidence are returned, and that verdicts are either auto-accepted or flagged for human review. It omits cost/latency expectations and whether any state is mutated, which keeps it out of 5 territory.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with the core action before the return shape and then the input guidance. The bare pattern URL is slightly cryptic but costs little space.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description correctly supplies the return shape and the auto/review semantics; no annotations exist either, and the description covers the main behavioral facts. What remains thin is parameter-specific guidance for subject_at and the evidence array matching behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 75%, so the schema already documents claims, evidence, subject_at and auto_accept in detail. The description only obliquely adds meaning to auto_accept via the 'auto or needs human review' phrasing and says nothing about subject_at or the single-vs-list evidence forms.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Check each claim against provided evidence text') and enumerates the output fields (verdict, probability distribution, confidence, auto vs review). It does not, however, differentiate itself from the many similarly named siblings (jev_audit, jev_review, jev_gate, jev_screen), leaving the agent to infer the boundary.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete usage context by naming what to pass as claims ('reports, PR descriptions, or agent briefs') and as evidence ('cited sources, diffs, or documents'), plus a canonical pattern reference. It stops short of saying when NOT to use it or which sibling to prefer for adjacent tasks.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.