Skip to main content
Glama

Validate Claim

validate_claim
Read-onlyIdempotent

"Is it true that…" / "fact check" / "verify the claim that…" / "did X really…" / "was Y actually…" / "confirm or refute" / "true or false" — natural-language claim verification against authoritative sources. Use whenever the agent needs to check whether something a user said is factually correct. Company-financial claims (revenue, net income, cash for public US companies) verify via the structured SEC EDGAR + XBRL fast path with exact percent-delta math; ANY OTHER factual claim (macro statistics, rates, prices, drug data, records) automatically falls through to the grounded pipeline — routed to the right live source, answered with verbatim evidence, then judged. Returns a verdict (confirmed / approximately_correct / refuted / inconclusive / unsupported / could_not_verify), the grounded or structured actual value with pipeworx:// citation, and reasoning. IMPORTANT for callers: could_not_verify means the check did not happen (our LLM or source failed) and carries verification_error{stage,detail} — it is NOT evidence for or against the claim, and must not be shown as one. unsupported means we looked and cover no source for it. Replaces 4–6 sequential calls (NL parsing → entity resolution → data lookup → comparison).

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
claimYesNatural-language factual claim, e.g., "Apple's FY2024 revenue was $400 billion" or "Microsoft made about $100B in profit last year".
tolerance_pctNoMax percent deviation still graded approximately_correct (0.5–50). Overrides the tolerance implied by the claim wording — set 1–2 for hallucination detection where any material error must be refuted. Default: implied by wording, capped at 5.

Schema Changelog

Changes observed during successful MCP inspections. Dates show when Glama detected each change.

  1. Changed1 schema field changed
    • addedInput schema / properties / tolerance_pct
      Added value: +{
      +  "description": "Max percent deviation still graded approximately_correct (0.5–50). Overrides the tolerance implied by the claim wording — set 1–2 for hallucination detection where any material error must be refuted. Default: implied by wording, capped at 5.",
      +  "type": "number"
      +}
  2. Added

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations (readOnly, openWorld, idempotent), the description reveals critical runtime behavior: the two-pipeline routing, the verdict enum, and especially the distinction between could_not_verify (check did not happen, with verification_error) and unsupported (no source), which is essential for correct interpretation. No contradictions with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

At ~170 words, the description is longer than strictly necessary, but the example trigger phrases and the detailed caveats about error conditions earn their place. It is well-organized and front-loaded with purpose, though a bit expansive.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has no output schema and only two parameters, so the description carries the full burden of explaining return values. It covers the verdict list, the actual value with citation, reasoning, and explicitly distinguishes could_not_verify from unsupported, which is critical for safe use. This is a complete and trustworthy description.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%: claim and tolerance_pct are both well-described with examples and ranges. The description adds a brief reference to tolerance logic but no substantive new parameter semantics, so it meets the baseline of 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with natural-language trigger phrases and clearly defines the tool as claim verification against authoritative sources. It details two distinct processing paths (SEC EDGAR for company-financial claims, grounded pipeline for other factual claims), which distinguishes it from generic Q&A tools in the sibling list.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states 'Use whenever the agent needs to check whether something a user said is factually correct' and provides guidance on how different claim types are routed. It does not name alternatives or state when not to use it, so it falls short of a 5, but the usage context is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

TDQS

A3.9/5.0
Disambiguation2/5

ask_pipeworx, ask_pipeworx_beta, and ask_pipeworx_grounded form a tight cluster, and the beta variant is explicitly identical to ask_pipeworx today. Several meta-tools like discover_tools, suggest_questions, and deep_research also overlap in discovery-oriented usage, so an agent must read carefully to pick the right one.

Naming Consistency3/5

Names are uniformly lowercase with underscores and mostly descriptive, but the conventions are mixed: verb_noun tools like validate_claim and list_subscriptions sit alongside noun_phrase tools like entity_profile and polymarket_arbitrage, plus bare verbs like remember and subscribe. It is readable but not a single predictable pattern.

Tool Count2/5

36 tools is well above the 25+ threshold, and the surface spans unrelated domains: EPA ECHO data, general Pipeworx research, Polymarket betting, memory, npm scanning, and AI visibility checks. For a server named 'Epa Echo', most tools feel out of scope and the collection seems like several separate servers merged together.

Completeness4/5

The EPA ECHO subset provides a solid facility-search, violations, compliance-history, and enforcement-action lifecycle. The broader Pipeworx surface also covers lookups, grounded answers, deep research, entity profiling, subscriptions, and memory, with only minor workaround-level gaps such as no dedicated ECHO permit/emissions detail tool.