Skip to main content
Glama

Check claims for unbacked certainty language

check_provenance_claims

Scans content for certainty claims like '(verified)' or 'guaranteed' that lack a proper provenance tier or source reference before publishing.

Instructions

Scans reader-facing claims for certainty-implying language ("(verified)", "independently verified", "guaranteed", "fact-checked", "100% accurate", ...) that isn't backed by an appropriate provenance tier, plus tiers that require a sourceRef but don't have one, and claims carrying an unrecognized tier. This is the exact pattern that caught ~150 false '(verified)' labels on a live site after they had already shipped -- run it on any copy, marketing page, or AI-drafted content that makes factual-sounding claims before it ships, not after. The default phrase list is a starting point drawn from that one incident, not a taxonomy -- it will miss phrases it doesn't know about (e.g. 'clinically proven', 'third-party tested'); extend certaintyPhrases for your domain. Negation detection is a fixed character window before a match, not a parser, so it can miss a negation in an earlier clause or over-suppress one further away. An empty result means every claim's certainty language (that this tool's phrase list and negation window caught) is backed by its tier.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
claimsYesThe claims to check.
caseSensitiveNoCase-sensitive phrase matching. Default false.
negationWindowNoCharacters before a phrase match to scan for a negation word ('not', 'without', ...). Default 40.
certaintyPhrasesNoOverride the certainty-phrase vocabulary. Defaults to a generic starter list ('(verified)', 'independently verified', 'proprietary dataset', 'guaranteed', 'fact-checked', '100% accurate', ...).
certaintyRequiresTierNoTiers strong enough to back certainty language. Default ['verified'].
requireSourceRefForTiersNoTiers that must carry a non-empty sourceRef. Default ['verified'].

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description carries the full burden and does so unusually well: it discloses the incident-derived provenance of the default phrase list, the phrase-list coverage gap ('will miss phrases it doesn't know about'), the negated-window limitation ('fixed character window... not a parser'), and the precise meaning of an empty result. These are exactly the caveats an agent needs to avoid over-trusting the output.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core detection behavior, then limitations. Dense and every sentence carries signal, though the incident anecdote is longer than strictly required and could be tightened without losing meaning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations and no output schema, the description does most of the work, covering scope, limits, and empty-result semantics. Its one gap is the shape of the returned offenses (it references 'offense messages' but never describes the result structure), which matters for a tool with no output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3; the description adds value beyond the schema by explaining that certaintyPhrases defaults are a starting point drawn from one incident and should be extended per domain, and by clarifying the tier semantics context for certaintyRequiresTier/requireSourceRefForTiers.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('scans reader-facing claims for certainty-implying language') and enumerates the three detection classes (unbacked certainty phrases, tiers requiring a sourceRef, unrecognized tiers). This is clearly distinct from siblings like check_grounding or check_claims_registry without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives strong workflow context ('run it on any copy, marketing page, or AI-drafted content that makes factual-sounding claims before it ships, not after') and a domain-extension hint for certaintyPhrases. It stops short of naming a sibling alternative or the condition that would route to one, so it falls just below the top bar.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.