Skip to main content
Glama

sanity_check

Verify a claim before you act on it or hand it downstream. Every verdict comes from a real model inference run at request time (HHEM entailment) — never cached or guessed.

Two modes:
  - "grounded" (default): is the claim faithful to the `context` you
    supply? `context` is REQUIRED. This is a scoreable question.
  - "open": is the claim consistent with live web evidence found right
    now (real Tavily search, then the same model)? `context` is ignored.
    This is consistency with current web content, NOT objective truth —
    a claim can be well-supported by the web and still be wrong.

Reading the result (in `data`):
  - `verdict`: "supported" (consistent with the source), "contradicted"
    (clearly inconsistent), "unsupported" (not supported but not clearly
    contradicted — treat as unverified), or "insufficient_evidence"
    (open mode found nothing usable; says nothing about whether the
    claim is true).
  - `confidence_score`: the raw consistency score in [0, 1] — near 1 =
    consistent with the source, near 0 = inconsistent. It is NOT
    confidence-in-the-verdict: a "contradicted" verdict has a score near
    0. It is 0.0 for "insufficient_evidence".
  - `evidence`: the audit trail (context snippet, or web citations).
Suggested policy: act on "supported"; treat "unsupported" and
"insufficient_evidence" as unverified; reject or regenerate on
"contradicted".

Access: the public mcp.wickedapi.com server uses a shared, rate-limited
key, so this returns real verdicts directly -- but the shared budget is
small (about 20 requests/minute and 30 open-mode checks/day across ALL
users; grounded mode is not counted against the daily limit). An
http_status 429 means that shared budget is used up: retry after the
Retry-After seconds, run this server locally with your own
SANITY_API_KEY, or pay per call via x402 directly. A server with no key
configured returns an http_status 402 carrying x402 payment instructions
under `payment_required` instead.

Args:
    claim: The single statement to verify (max 5,000 characters).
    context: Source text the claim must be faithful to — a string or a
        list of strings. Required for mode "grounded".
    mode: "grounded" (check against `context`) or "open" (check against
        live web evidence).

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
modeNogrounded
claimYes
contextNo

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observed

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so thoroughly: it discloses that verdicts come from real-time HHEM inference (never cached), the exact rate limits (20/min, 30 open-mode/day), 429/Retry-After handling, and the 402 x402 payment path. It also clarifies the subtle semantics of confidence_score and the four verdict values.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core purpose and uses clear labeled sections (Two modes, Reading the result, Access, Args) so it is scannable despite its length. It is on the long side, but almost every line carries non-obvious operational value, so waste is minimal.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a multi-mode verification tool, the description covers mode selection, verdict interpretation, evidence output, and failure/rate-limit handling, which is more than sufficient given an output schema already exists. An agent has everything needed to call it correctly and interpret results.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate, and it does: claim is capped at 5,000 characters, context accepts a string or list and is REQUIRED for grounded mode but ignored in open mode, and mode defaults to grounded. This fully documents behavior the schema leaves bare.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (verify) and resource (a claim) with clear scope ('before you act on it or hand it downstream'). The args section distinguishes it from the sibling sanity_check_batch by emphasizing 'The single statement to verify'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly defines the two modes and their selection condition (grounded = faithful to supplied context, context REQUIRED; open = live web evidence, context ignored). Adds a 'suggested policy' telling the agent exactly how to act on each verdict, leaving nothing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.