Skip to main content
Glama

hallucination check

hallucination_check

Entailment scoring: does the provided context support this claim? Research agents use this to self-check outputs. [price: $0.001/call USDC via x402]

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
claimYesClaim to check
contextYesEvidence/context text

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden. It explains the tool performs entailment scoring and includes a clear price/payment detail ($0.001/call USDC via x402). It does not specify the exact output format, but 'scoring' implies a result that evaluates claim support.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single tightly-worded sentence with no redundancy. It front-loads the core behavior, follows with the intended use case, and appends the cost detail compactly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has no output schema, and the description does not state the shape or range of the correspondence score. 'Entailment scoring' gives some hint, but an agent may not know whether the result is a boolean, a probability, or a score on a specific scale.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the parameter descriptions already document `context` and `claim`. The tool description adds little beyond paraphrasing the schema ('provided context' and 'this claim'), so the baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Entailment scoring: does the provided context support this claim?' This identifies a specific action and resource. It doesn't explicitly distinguish from sibling tools like search_verify, but its meaning is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Research agents use this to self-check outputs' gives clear context for when to use the tool. It does not mention exclusions or alternatives, but the intended use case is explicit enough for an agent to decide appropriately.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

TDQS

B3.2/5.0
Disambiguation3/5

Most tools have distinct purposes, but several clusters overlap: domain_facts, page_meta, and scrape all return page title information, and search_verify, hallucination_check, and sweep all target claim validation. The descriptions usually clarify the use case, but the boundaries are not always obvious.

Naming Consistency3/5

All names use lowercase snake_case, so there is a baseline consistency, but the pattern is mixed: bare verbs like scrape, summarize, and sweep sit alongside noun+noun forms like domain_facts and noun+verb forms like entity_find. The names are readable but do not form a predictable verb_noun API convention.

Tool Count3/5

At 26 tools, this is heavy and above the typical well-scoped 3-15 range, though the server is explicitly positioned as a broad shelf of paid utilities. Many tools are small one-purpose endpoints, so the count feels more like a catalog than a focused suite, but it is not an extreme mismatch.

Completeness4/5

The shelf covers the major advertised areas: web page analysis, research verification, text guards and NLP, blockchain reads, and image generation. There are some gaps such as no web search and no transaction sending, but agents can typically work around them or pair this with another server.

Resources