Skip to main content
Glama

Audit extracted values against their source text

jev_audit

Audit extracted values against their claimed source text to catch hallucinations, off-target or incomplete data, wrong formats, and omissions before trusting them.

Instructions

Audit extracted values against the text they claim to come from, before the values are trusted: one request with a per-value failure-mode battery (hallucinated / off-target / incomplete / wrong format, each framed so true = something is wrong) plus a dedicated omission check for empty values. Any value's P(wrong) at wrong_at escalates the whole audit; max-gated, never averaged. For multimodal intake: run your vision or ASR model first to produce a dense transcript of the image, scan, or recording, screen that transcript with jev_screen, then audit the extracted values against it here — the tool never sees pixels or audio, it audits two text artifacts against each other. Schema validation catches structural errors; it can flag a schema-valid fabrication against the supplied source text — a limited cross-check, not verification of the original.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
sourceYesThe text the values claim to come from: a document, or the dense transcript a vision or ASR model produced. Truncated at 50,000 chars.
recordsYesExtracted values to audit against the source. Up to 32 per call.
wrong_atNoA record's P(wrong) at or above this flags the value and escalates. Default 0.7.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Addedv0.12.0

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so: four named failure modes with the polarity spelled out ('true = something is wrong'), a dedicated omission check for empty values, max-gating that 'never averaged', and escalation via wrong_at. It also discloses hard limits (never sees pixels or audio, truncated at 50,000 chars, audits two text artifacts against each other) that an agent could not infer otherwise.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Dense but front-loaded: the core operation and failure-mode battery come first, with workflow and limitations after. The em-dash clauses are information-dense rather than padded, though the closing 'limited cross-check, not verification of the original' restates an idea already implied earlier.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers purpose, failure modes, gating semantics, and workflow for a tool with no annotations and no output schema. The remaining gap is that it never sketches the shape of the result (per-record flags vs a whole-audit escalation verdict), which an agent would want when no output schema exists.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3; the description goes beyond it by explaining wrong_at's behavioral role ('escalates the whole audit; max-gated, never averaged' with a 0.7 default) rather than just its type. It adds meaning for the gating parameter but says nothing extra about source/records beyond what the schema already documents.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('audit extracted values against the text they claim to come from') and scopes it precisely as a pre-trust check on two text artifacts. An agent can distinguish it from sibling jev_screen/jev_verify without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly names the triggering condition ('before the values are trusted') and lays out the multimodal intake sequence: run vision/ASR, screen with jev_screen, then audit here. It also states the boundary versus schema validation, giving both when-to-use and when-not-to-rely-on-it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.