Skip to main content
Glama

Validate Claim

validate_claim
Read-onlyIdempotent

"Is it true that…" / "fact check" / "verify the claim that…" / "did X really…" / "was Y actually…" / "confirm or refute" / "true or false" — natural-language claim verification against authoritative sources. Use whenever the agent needs to check whether something a user said is factually correct. Company-financial claims (revenue, net income, cash for public US companies) verify via the structured SEC EDGAR + XBRL fast path with exact percent-delta math; ANY OTHER factual claim (macro statistics, rates, prices, drug data, records) automatically falls through to the grounded pipeline — routed to the right live source, answered with verbatim evidence, then judged. Returns a verdict (confirmed / approximately_correct / refuted / inconclusive / unsupported / could_not_verify), the grounded or structured actual value with pipeworx:// citation, and reasoning. IMPORTANT for callers: could_not_verify means the check did not happen (our LLM or source failed) and carries verification_error{stage,detail} — it is NOT evidence for or against the claim, and must not be shown as one. unsupported means we looked and cover no source for it. Replaces 4–6 sequential calls (NL parsing → entity resolution → data lookup → comparison).

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
claimYesNatural-language factual claim, e.g., "Apple's FY2024 revenue was $400 billion" or "Microsoft made about $100B in profit last year".
tolerance_pctNoMax percent deviation still graded approximately_correct (0.5–50). Overrides the tolerance implied by the claim wording — set 1–2 for hallucination detection where any material error must be refuted. Default: implied by wording, capped at 5.

Schema Changelog

Changes observed during successful MCP inspections. Dates show when Glama detected each change.

  1. Changed1 schema field changed
    • addedInput schema / properties / tolerance_pct
      Added value: +{
      +  "description": "Max percent deviation still graded approximately_correct (0.5–50). Overrides the tolerance implied by the claim wording — set 1–2 for hallucination detection where any material error must be refuted. Default: implied by wording, capped at 5.",
      +  "type": "number"
      +}
  2. First observed

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare it read-only, open-world, and idempotent. The description adds crucial behavior beyond this: it defines the meaning of each verdict, especially that 'could_not_verify' means the check failed (not evidence against) and 'unsupported' means no source was found. It also discloses tolerance_pct semantics and the error field, giving the agent a clear mental model.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but well-structured: trigger phrases, purpose, routing, return values, and caller warnings are logically ordered. Each section earns its place—the caller warning about 'could_not_verify' is critical. It's slightly verbose but not wasteful given the tool's complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description thoroughly covers return values (verdicts), the verification_error field, and domain-specific routing, fully compensating for the lack of an output schema. It also explains edge cases (whether a check happened) that could otherwise mislead an agent. This is complete for a tool of this complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Both parameters have full descriptions in the schema (100% coverage). The description adds value by explaining tolerance_pct's default behavior (implied by wording, capped at 5), how it overrides that default, and recommending 1-2% for hallucination detection—details not present in the schema. This goes beyond the baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's function: verifying natural-language factual claims against authoritative sources, with specific trigger phrases like 'Is it true that…' and 'fact check.' It also emphasizes its role as a composite tool replacing 4-6 sequential calls, distinguishing it from single-step tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit usage context: 'Use whenever the agent needs to check whether something a user said is factually correct.' It also explains the routing logic for company-financial vs. other claims, giving clear criteria for when it applies. However, it doesn't name sibling alternatives or explicitly state when not to use it, so it's not a perfect 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

TDQS

A3.6/5.0
Disambiguation2/5

There are three ask_pipeworx variants that overlap heavily, plus five polymarket_* tools scanning similar market edges, and meta-tools like discover_tools, suggest_questions, and deep_research that blur together. ai_visibility_check and scan_competitor_ai_presence also overlap. The Cloudflare Radar tools are distinct, but they are a small minority in a sea of overlapping data/prediction-market tools.

Naming Consistency3/5

Most tools use snake_case, but the pattern is inconsistent: some are verb_noun (list_subscriptions, resolve_entity, scan_dependency), some are noun_phrase (bgp_leaks, internet_quality, radar_domain_rank), and some use vendor-prefixed naming inconsistently (ask_pipeworx vs pipeworx_feedback vs pipeworx_trending). It is readable but not predictable.

Tool Count2/5

37 tools is heavy for any single server, and the collection spans unrelated domains: Cloudflare Radar, Pipeworx data lookup, Polymarket betting, memory, subscriptions, and npm scanning. The count feels like a bundled platform rather than a focused tool set, and many tools could be split into separate servers.

Completeness3/5

The Pipeworx data-research side is quite complete (ask, grounded, deep_research, entity_profile, compare_entities, validate_claim, resolve_entity, discover_tools, recent_changes), and the prediction-market side has good coverage (edges, arbitrage, fill risk, cross-venue spread, tracking). However, the Cloudflare Radar portion is thin—only six tools cover a service known for many more traffic/attack/outage metrics—and the overall surface has no cohesive domain to judge completeness against.