Skip to main content
Glama

Check citation grounding

check_grounding

Detect ungrounded or forged citations in AI-generated text by classifying each sentence against provided evidence before shipping reports, summaries, or answers.

Instructions

Detects ungrounded or forged citations in AI-generated text. Splits text into sentences and classifies each one against evidence: 'grounded' (cites a marker whose evidence plausibly supports it), 'placeholder' (an honest 'TBD'/unknown gap, no fake citation), 'ungrounded' (a claim with no citation at all), or 'invalid' (cites a marker id that is missing from evidence, or whose evidence doesn't plausibly support the sentence under the default matcher -- i.e. a forged or hallucinated citation; this outranks every other status). This is a mechanical/structural check, not a truth checker: the default support check is naive substring/word-overlap matching, not semantic entailment -- it can pass a coincidental word match and can fail a genuine paraphrase, and it cannot verify that the evidence itself is true. Use this before shipping any AI-written report, summary, or answer that cites sources, to catch a model inventing or misattributing a citation. Treat any 'invalid' sentence as a hard stop; treat 'ungrounded' sentences as claims that should probably cite something but currently don't. For higher-stakes content, use grounding-kit directly with a custom supports() function (embedding-similarity or NLI-based) instead of the default matcher.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
textYesThe AI-generated text to check. Citation markers use the kit's default convention `[[cite:id]]` (e.g. "The bridge opened in 1932 [[cite:source-a]]."). This tool uses the default marker/placeholder patterns; if your generator emits a different citation syntax (e.g. "[1]"), rewrite markers to `[[cite:1]]` before calling, or use grounding-kit directly with a custom markerPattern.
evidenceYesMap of citation marker id -> the evidence text/span it claims to support. This is the closed world: a marker cited in `text` whose id is NOT a key here is flagged invalid, and a marker whose evidence text doesn't plausibly support the sentence is also flagged invalid. Pass {} if there is no evidence at all (every citation will then be invalid, and uncited claims will be 'ungrounded').

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so well: it discloses that this is a mechanical/structural check, that the default matcher is naive substring/word-overlap rather than semantic entailment, that it can both false-positive (coincidental word match) and false-negative (genuine paraphrase), and that it cannot verify the truth of the evidence itself. It also states the priority rule ('invalid outranks every other status').

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the purpose, then classification definitions, matcher caveat, and escalation path in a logical order where each sentence carries distinct information. It is long, but the length is justified by the four-status taxonomy and the non-obvious matcher limitations.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, yet the description fully enumerates the return classifications and their meanings, and covers the matcher's failure modes and parameter interplay. For a 2-parameter tool with a nested evidence map and no annotations, nothing an agent needs to invoke and interpret it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real meaning beyond the schema by framing `evidence` as a closed world and explaining the parameter interaction (a marker cited in `text` missing from `evidence` is invalid; empty evidence makes all citations invalid). It reinforces the `[[cite:id]]` marker convention and notes what to do with alternate syntax.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first sentence states a specific verb and resource ('Detects ungrounded or forged citations in AI-generated text') and the body enumerates the four classification outcomes. It is clearly distinguishable from siblings like corroborate_evidence or check_provenance_claims by naming the exact artifact it inspects (sentence-level citation markers vs. evidence map).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says when to use it ('before shipping any AI-written report, summary, or answer that cites sources') and when to escalate ('for higher-stakes content, use grounding-kit directly with a custom supports() function'). It also names the alternative tool and the condition that selects it, leaving nothing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.