Skip to main content
Glama
fredyee

HallucC MCP Server

verify_agent

Checks agent final responses for hallucinated claims and evaluates execution traces across six dimensions: factuality, sourcing, instruction compliance, tool-claim consistency, task completion, and reflection.

Instructions

Agent 输出与轨迹自检。对最终回答跑逐声明幻觉核验,并对执行轨迹做六维结构化评估(事实性/来源/指令合规/工具声明一致/任务完成/反思)。返回 main.claims、steps[].checks、quota。耗 detect 额度。

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
taskNo原始任务指令(task_completion 裁判用)
speedNo速度模式:fast|standard|deep,默认 standard
stepsNoagent 执行轨迹,逐步六维评估
domainNo领域模式:general|medical|legal|finance|education|government
agent_outputYesagent 最终输出文本(≥1 字,≤50000 字)
tool_schemasNo工具 Schema(与 trajectory 端点对齐)

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It does disclose a meaningful behavioral trait: '耗 detect 额度' (consumes detect quota), effectively a cost/rate-limit warning, and clarifies the evaluation dimensions and return payload. It omits failure modes, error behavior, and how quota is shared across calls, but for a non-mutating analysis tool the disclosed cost and output structure are adequate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three terse sentences in the source language with zero filler. Purpose is front-loaded first, then return values, then cost. Every clause carries information – no repetition of the tool name or schema content. Efficient and well-ordered.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex tool with 6 params and nested objects but no output schema, the description compensates reasonably by naming the return fields (main.claims, steps[].checks, quota) and the quota-cost behavior. It does not fully explain the return structure's semantics or the effect of speed/domain switches, but for an analysis tool the coverage is strong.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema fully documents all six parameters. The description adds only framing: it ties 'task' to task_completion judging and maps the six dimensions to steps[].checks, but does not explain the semantics of speed modes (fast|standard|deep), domain values, or tool_schemas beyond what the schema already provides. Baseline 3 applies since the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource pair ('Agent 输出与轨迹自检' – output and trajectory self-check) and details exactly what it does: per-claim hallucination verification plus a six-dimension structured evaluation (事实性/来源/指令合规/工具声明一致/任务完成/反思). It names the concrete return structures, making the tool's function unambiguous and distinct from simpler verifiers.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The scope is implied through the six evaluation dimensions and the agent_trajectory param, which signals this is for full agent runs rather than single strings. However, it never explicitly contrasts with siblings like verify_text, check_cua_actions, or check_safety, nor states when NOT to use it. The '耗 detect 额度' cost note is useful resource guidance but not a when/when-not rule.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.