Skip to main content
Glama

Validate Claim

validate_claim
Read-onlyIdempotent

"Is it true that…" / "fact check" / "verify the claim that…" / "did X really…" / "was Y actually…" / "confirm or refute" / "true or false" — natural-language claim verification against authoritative sources. Use whenever the agent needs to check whether something a user said is factually correct. Company-financial claims (revenue, net income, cash for public US companies) verify via the structured SEC EDGAR + XBRL fast path with exact percent-delta math; ANY OTHER factual claim (macro statistics, rates, prices, drug data, records) automatically falls through to the grounded pipeline — routed to the right live source, answered with verbatim evidence, then judged. Returns a verdict (confirmed / approximately_correct / refuted / inconclusive / unsupported / could_not_verify), the grounded or structured actual value with pipeworx:// citation, and reasoning. IMPORTANT for callers: could_not_verify means the check did not happen (our LLM or source failed) and carries verification_error{stage,detail} — it is NOT evidence for or against the claim, and must not be shown as one. unsupported means we looked and cover no source for it. Replaces 4–6 sequential calls (NL parsing → entity resolution → data lookup → comparison).

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
claimYesNatural-language factual claim, e.g., "Apple's FY2024 revenue was $400 billion" or "Microsoft made about $100B in profit last year".
tolerance_pctNoMax percent deviation still graded approximately_correct (0.5–50). Overrides the tolerance implied by the claim wording — set 1–2 for hallucination detection where any material error must be refuted. Default: implied by wording, capped at 5.

Schema Changelog

Changes observed during successful MCP inspections. Dates show when Glama detected each change.

  1. Changed1 schema field changed
    • addedInput schema / properties / tolerance_pct
      Added value: +{
      +  "description": "Max percent deviation still graded approximately_correct (0.5–50). Overrides the tolerance implied by the claim wording — set 1–2 for hallucination detection where any material error must be refuted. Default: implied by wording, capped at 5.",
      +  "type": "number"
      +}
  2. First observed

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate a safe read-only, idempotent operation, so the description adds valuable context: it details the SEC EDGAR/XBRL path for financial claims, the grounded pipeline for others, and importantly distinguishes could_not_verify (pipeline failure, not evidence) from unsupported (no source covers it). This goes well beyond annotations and prevents misinterpretation of results.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with trigger phrases and provides a well-organized flow: usage, processing paths, return values, caveats. It is somewhat long but dense with necessary information for a complex tool. Every major sentence contributes value; a slight trim could improve conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description fully explains return semantics: enumerated verdicts, grounded actual value with citation, reasoning, and the critical error-handling distinction between could_not_verify and unsupported. It also explains the tool's efficiency benefit. For a tool with this complexity, the description is remarkably complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already provides 100% coverage for both parameters (claim and tolerance_pct) with clear descriptions. The tool description does not add parameter-specific meaning; it only indirectly references tolerance by mentioning 'exact percent-delta math'. With full schema coverage, baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool performs natural-language claim verification against authoritative sources, with specific verbs like 'validate', 'fact check', and 'verify'. It distinguishes itself from sibling research/query tools by framing the purpose as checking factual correctness and noting it replaces 4–6 sequential calls, which uniquely positions its role.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when to use it: 'Use whenever the agent needs to check whether something a user said is factually correct.' It also clarifies the two processing paths (SEC EDGAR fast path vs. grounded pipeline) and notes unsupported means no source exists. It lacks explicit when-not guidance or named alternatives, but the context is strong.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

TDQS

A3.8/5.0
Disambiguation3/5

Tools are individually well-described, but there is overlapping functionality among ask_pipeworx, ask_pipeworx_grounded, and deep_research, as well as multiple prediction market tools. The detailed descriptions help, but the number of similar tools creates some ambiguity.

Naming Consistency4/5

Tool names follow a consistent snake_case pattern with descriptive prefixes (e.g., ask_pipeworx, entity_profile, polymarket_edges). There are minor deviations like 'deep_research' vs 'bet_research' but overall the naming is predictable and clear.

Tool Count2/5

The server is named 'Texas Open Data' but only 3 of 33 tools (datasets, metadata, query) are directly related to Texas open data. The remaining tools cover a much broader domain (Pipeworx ecosystem), making the tool count inappropriate for the stated purpose.

Completeness2/5

For the stated purpose of Texas Open Data, the tool surface is incomplete—only basic query and metadata capabilities are provided, lacking data management, update, or delete operations. As a general pipeworx server it might be more complete, but the name suggests Texas-specific data.