HallucC MCP Server
Fact-check claims, evaluate agent trajectories, gate Computer-Use Agent actions by risk level, and scan prompts/outputs for safety threats (plus static agent code audit per README).
verify_text: Claim-by-claim hallucination check; returns red/yellow/green summary, per-claim verdicts (supported/refuted/unverified), confidence, reasons, sources, citations. Costs detect quota.verify_agent: Verifies an agent's final output and evaluates its trajectory across 6 dimensions (factuality, sourcing, instruction compliance, tool-claim consistency, task completion, reflection). Costs detect quota.check_cua_actions: L0–L3 risk classification for CUA actions (free, rule-based); covers sensitive paths, destructive commands, domain grading, payment thresholds, task-scope deviation, credential takeover, and multi-agent delegation-chain taint analysis.check_safety: OWASP LLM Top 10-style safety gateway (40+ rules + optional LLM) for prompt injection, jailbreak, harmful content, PII leakage, and fraud; optional hallucination check. Costs detect quota.check_cua_code_audit(README): Static audit of agent source code for dangerous imports, permission boundaries, injection surfaces, dangerous defaults, and sandbox absence; free.
Makes the HallucC fact-check assistant available in the Coze store, enabling Coze users to perform fact verification and AI safety checks.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@HallucC MCP Serververify this claim: the Great Wall is visible from space"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
HallucC MCP Server
中文文档 | Live Demo | Official MCP Registry
Give your AI a fact-checker and a safety gateway: claim-by-claim hallucination detection, agent trajectory evaluation, L0-L3 risk gating for Computer-Use Agent (CUA) actions, Agent source code static audit, and a 40+ feature prompt-injection / jailbreak guard — one MCP server, five tools, works with Claude Code / Cursor / Claude Desktop.
See it first — no sign-up required
🔍 A real verification report (per-claim verdicts + confidence + sources): https://aihcc.cloud/r/GttouoJfMpkC
🛡️ Try CUA action gating online (rule-based, free, no quota cost): https://aihcc.cloud/cua
📝 Web app (3 free checks per day): https://aihcc.cloud
Related MCP server: Verity MCP
Tools (5)
Tool | Backend endpoint | What it does | Quota |
|
| Claim-by-claim hallucination check: red/yellow/green summary + per-claim status/confidence/reason/sources + citations | ✅ detect |
|
| Text-level verification of an agent's final output + 6-dimension trajectory evaluation (factuality / sourcing / instruction compliance / tool-claim consistency / task completion / reflection) | ✅ detect |
|
| L0-L3 risk classification for Computer-Use Agent (CUA) actions + | ❌ free |
|
| Static audit of Agent source code: dangerous imports / permission boundaries / injection surfaces / dangerous defaults / sandbox absence | ❌ free |
|
| 40+ feature safety gateway: prompt injection / jailbreak / harmful content / PII leakage / fraud | ✅ detect |
Auth model: each client sends
Authorization: Bearer <your HallucC API key>→ the server forwards it as-is to the backend → the backend validates and deducts from the same account quota as the web app. Keys travel only in HTTP headers — never in tool arguments, never in model transcripts.
The tightening capabilities of check_cua_actions
Field | Effect |
| Target URL of a browser action — navigation to an unknown domain escalates to L2; pasting sensitive content to a non-allowlisted domain escalates to L3; blacklisted domains are L3 |
| Declares the intended task, used for deviation detection and audit traceability |
| Allowed apps / domains (domain suffix matching includes subdomains) |
| Forbidden element labels (highest precedence) |
| Multi-agent delegation chain id (uuid hex) — when given, the response gains a |
| The agent owning this action / its delegating parent — pairs cross-agent-boundary propagation chains |
| Delegation depth (root agent = 0); a subagent exfiltrating content matching sensitive patterns gets an L2 |
An action that leaves task_scope is escalated one level (L0→L1, L1→L2, L2 stays L2 with an "outside task scope" reason appended, L3 unchanged). This mechanism only tightens and never loosens; omitting task_scope has no effect. Credential fields and payment/login domains yield a takeover verdict (suspend and wait for the user to type it themselves — the agent never types your password).
Multi-agent setups (Claude Code Task subagents, Codex agents, Kimi, …): a parent agent reading a credential file (L0 on its own) → a subagent posting it out (L2 on its own) is harmless per hop and dangerous as a chain. Attach the delegation-chain fields to each action and the backend appends a chain_trace: when cross-agent propagation exists the chain aggregate level is raised to L3 and risk_diluted is set (the laundering signal). Leave the chain fields empty and behaviour is bit-for-bit unchanged — existing callers need no change.
The response's ruleset_version lets you verify that online detection and the local Gate run the same rule version.
Hosted endpoint (recommended, zero local setup)
The MCP server is hosted at https://aihcc.cloud/mcp (remote, streamable-http). Point your client at it with your API key:
# Claude Code
claude mcp add --transport http hallucc https://aihcc.cloud/mcp \
--header "Authorization: Bearer <your HallucC API key>"Cursor / Claude Desktop work the same way: URL https://aihcc.cloud/mcp, header Authorization: Bearer <key>. Create your API key at https://aihcc.cloud (register → /keys). Free daily quota included.
Or: run via npx (stdio, no clone needed)
{
"mcpServers": {
"hallucc": {
"command": "npx",
"args": ["-y", "hallucc-mcp"],
"env": { "HALLUCC_API_KEY": "<your HallucC API key>" }
}
}
}Same five tools over stdio transport; the key comes from the environment and requests default to the hosted backend https://aihcc.cloud/api.
The base URL must keep the
/apiprefix: nginx only proxies/api/to the backend. A barehttps://aihcc.cloudreaches the frontend instead — the request returns 200 with an HTML body, fails silently (nothing lands in audit or in a fallback file), and the audit replay stays empty.
Run locally (from source)
# 1. Install dependencies
npm install
# 2. (Optional) configure backend address & port; defaults to local :8001
cp .env.example .env
# HALLUCC_BASE_URL=http://127.0.0.1:8001
# PORT=8787
# 3. Make sure the backend is up: curl http://127.0.0.1:8001/health
# 4. Start the MCP server (dev with hot reload)
npm run dev
# or production: npm startThe server listens on http://127.0.0.1:8787/mcp; health check GET /health.
Client setup
Claude Code (CLI)
claude mcp add --transport http hallucc http://127.0.0.1:8787/mcp \
--header "Authorization: Bearer <your HallucC API key>"
claude mcp list
# then in conversation: ask Claude to fact-check a passage or scan a prompt for injectionCursor
Settings → MCP → Add MCP server:
Type:
httpURL:
http://127.0.0.1:8787/mcpHeaders:
{"Authorization": "Bearer <your HallucC API key>"}
Claude Desktop
~/Library/Application Support/Claude/claude_desktop_config.json:
{
"mcpServers": {
"hallucc": {
"type": "http",
"url": "http://127.0.0.1:8787/mcp",
"headers": { "Authorization": "Bearer <your HallucC API key>" }
}
}
}MCP Inspector (debug — tool list & auth work without spending quota)
npx @modelcontextprotocol/inspector
# choose Streamable HTTP, URL http://localhost:8787/mcp
# add Custom Header: Authorization: Bearer <key>Auth & quota
Missing/invalid
Authorizationheader → 401 JSON-RPC error at the MCP layer.Valid key but quota exhausted → backend 429, mapped to "free quota exhausted, resets tomorrow or upgrade".
Every tool response includes a
quotaobject (backendquota_status), same source as the web dashboard, decremented per call.check_cua_actionsandcheck_cua_code_auditare both pure rules — no LLM, no quota cost, safe for high-frequency use.
Project layout
src/
server.ts Express + StreamableHTTP transport (stateless) + auth middleware + tool registration
context.ts config loading + API key extraction
backend.ts BackendClient (key pass-through, unified 401/429/5xx mapping) + toMcpResult
schemas.ts zod input schemas for the 5 tools (aligned with backend Pydantic Fields)
tools/
verifyText.ts → /detect
verifyAgent.ts → /detect-agent
checkCuaActions.ts → /cua/classify
checkCuaCodeAudit.ts → /cua/audit-code
checkSafety.ts → /guard/check (fast → /guard/check-fast)
index.ts registerAllToolsTransport
Streamable HTTP (the remote transport recommended by the current MCP spec, formerly "HTTP+SSE"), stateless mode: each POST spins up a fresh transport + McpServer and closes when done. Natively supported by Claude Code / Cursor / Claude Desktop. To support legacy SSE-only clients, add SSEServerTransport (dual endpoints /sse + /messages).
Listed on
✅ Official MCP Registry:
io.github.fredyee/halluccv0.1.8 (active)✅ ModelScope MCP Square: @hallucC/hallucc
✅ Coze Store: HallucC Fact-Check Assistant
License
Available Tools
4 toolscheck_cua_actionsAInspect
Computer-Use Agent 动作风险分级。提交动作轨迹 → 逐步返回 L0(放行)/L1(放行+记录)/L2(需确认)/L3(阻断) 裁决、matched_rules、outcome,及 L0-L3 汇总计数。纯规则判定,不调 LLM、不耗额度。用于在 agent 执行前/后做安全自检。
| Name | Required | Description | Default |
|---|---|---|---|
| actions | Yes | 待分级的动作序列(1~500 条) | |
| session_id | No | 给了则把分级结果写入审计表,可在 /cua/audit 回放 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and adds meaningful behavioral context: it is rule-based, does not call an LLM, and does not consume quota. It also discloses the return value shape. The optional audit side effect when session_id is supplied is only mentioned in the schema, not the description, but that is a minor gap because the schema documents it.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded: it states the purpose, the output, the key behavioral trait, and the intended usage context in three dense sentences. There is no filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description enumerates the return payload (L0-L3 verdicts, matched_rules, outcome, counts), which compensates for the lack of an output schema. Combined with a fully documented input schema, an agent has enough to call the tool correctly. Slightly more top-level transparency about the session_id audit write would round it out, but that is captured in the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the input schema already documents all parameters and aliases in detail. The description adds little parameter-level meaning beyond 'submit action trajectory', so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: risk-grading CUA action trajectories and returning L0-L3 verdicts, matched_rules, outcome, and counts. It is not a tautology and clearly communicates the tool's function, though it does not explicitly differentiate from sibling tools like check_safety or verify_agent.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives a clear applicability context: '用于在 agent 执行前/后做安全自检' (used for safety self-check before/after agent execution). It also notes that the tool is pure rule-based and consumes no LLM quota, which helps an agent decide when to invoke it. However, it does not mention exclusions or explicitly name alternatives among the sibling tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_safetyAInspect
AI 安全网关检测。对 prompt + output 跑 OWASP LLM Top 10 三层检测:40+ 特征规则引擎(注入/越狱/有害内容/敏感信息泄露/欺诈)+ 可选 LLM 深度分析 + 可选幻觉检测。返回 passed、risk_score、risk_level、threats[](category/severity/match_context/recommendation)、quota。fast=true 走纯规则 <10ms 模式。耗 detect 额度(fast 亦计入)。
| Name | Required | Description | Default |
|---|---|---|---|
| fast | No | true 走纯规则快速模式(<10ms,无 LLM)→ /guard/check-fast;默认 false 走完整 /guard/check | |
| output | Yes | AI 响应(≥1 字,≤50000 字) | |
| prompt | No | 用户原始输入(默认空,≤50000 字) | |
| check_hallucination | No | 是否同时检测幻觉,默认 false |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden and largely delivers: it discloses that the tool runs rule-based plus optional LLM checks, returns specific risk fields, supports a fast <10ms mode, and consumes detect quota even in fast mode. It does not discuss mutability or side effects, but for a detection tool this is adequate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and front-loaded, stating the core purpose first, then detection layers, return schema, fast-mode behavior, and quota impact. Every sentence conveys actionable information without redundancy or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema, the description compensates by naming the main return fields (passed, risk_score, risk_level, threats[], quota) and by explaining fast-mode routing. It does not describe possible errors or quota limits in detail, but for a 4-parameter safety check tool the provided context is largely sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds value beyond the schema by clarifying the fast-mode semantics, the prompt+output relationship, optional hallucination checking, and the returned risk structure. This is more than the schema alone provides, so a 4 is justified.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb and resource: an AI safety gateway that detects OWASP LLM Top 10 risks on prompt and output. It further distinguishes the tool by enumerating rule engine dimensions, optional LLM analysis, and hallucination detection, making it easy to tell apart from siblings like verify_text or verify_agent.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when to use the tool: whenever an AI prompt/output needs safety screening. It also explains the fast-mode vs full-mode choice and quota behavior excluded. However, it does not explicitly mention cases where sibling tools should be used instead, so no exclusions or when-not guidance is provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
verify_agentAInspect
Agent 输出与轨迹自检。对最终回答跑逐声明幻觉核验,并对执行轨迹做六维结构化评估(事实性/来源/指令合规/工具声明一致/任务完成/反思)。返回 main.claims、steps[].checks、quota。耗 detect 额度。
| Name | Required | Description | Default |
|---|---|---|---|
| task | No | 原始任务指令(task_completion 裁判用) | |
| speed | No | 速度模式:fast|standard|deep,默认 standard | |
| steps | No | agent 执行轨迹,逐步六维评估 | |
| domain | No | 领域模式:general|medical|legal|finance|education|government | |
| agent_output | Yes | agent 最终输出文本(≥1 字,≤50000 字) | |
| tool_schemas | No | 工具 Schema(与 trajectory 端点对齐) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It does disclose a meaningful behavioral trait: '耗 detect 额度' (consumes detect quota), effectively a cost/rate-limit warning, and clarifies the evaluation dimensions and return payload. It omits failure modes, error behavior, and how quota is shared across calls, but for a non-mutating analysis tool the disclosed cost and output structure are adequate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three terse sentences in the source language with zero filler. Purpose is front-loaded first, then return values, then cost. Every clause carries information – no repetition of the tool name or schema content. Efficient and well-ordered.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex tool with 6 params and nested objects but no output schema, the description compensates reasonably by naming the return fields (main.claims, steps[].checks, quota) and the quota-cost behavior. It does not fully explain the return structure's semantics or the effect of speed/domain switches, but for an analysis tool the coverage is strong.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema fully documents all six parameters. The description adds only framing: it ties 'task' to task_completion judging and maps the six dimensions to steps[].checks, but does not explain the semantics of speed modes (fast|standard|deep), domain values, or tool_schemas beyond what the schema already provides. Baseline 3 applies since the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource pair ('Agent 输出与轨迹自检' – output and trajectory self-check) and details exactly what it does: per-claim hallucination verification plus a six-dimension structured evaluation (事实性/来源/指令合规/工具声明一致/任务完成/反思). It names the concrete return structures, making the tool's function unambiguous and distinct from simpler verifiers.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The scope is implied through the six evaluation dimensions and the agent_trajectory param, which signals this is for full agent runs rather than single strings. However, it never explicitly contrasts with siblings like verify_text, check_cua_actions, or check_safety, nor states when NOT to use it. The '耗 detect 额度' cost note is useful resource guidance but not a when/when-not rule.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
verify_textBInspect
逐声明幻觉核验。输入文本 → 提取声明 → 逐声明验证(supported/refuted/unverified)+ 来源引用。返回红/黄/绿汇总、每条声明的状态/置信度/理由/来源、citations、quota。耗 detect 额度(与 Web 端同源)。
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | 待核验的文本 | |
| model | No | 被测模型名(选填,分享卡片展示用) | |
| speed | No | 速度模式:fast|standard|deep,默认 standard | |
| domain | No | 领域模式:general|medical|legal|finance|education|government | |
| strict | No | 严格模式(自我一致性),不传走全局默认 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It does disclose that the tool consumes detect quota (a side effect) and describes the return structure (red/yellow/green summary, per-claim details, citations, quota). However, it does not mention potential rate limits, authentication, or any other operational behaviors. The disclosure is partial but not misleading.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, dense paragraph that front-loads the core purpose ('逐声明幻觉核验') and follows with the pipeline and output details. Every clause adds value, and there is no redundant fluff. It is concise yet information-dense, though it could benefit from bullet points for readability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given that there is no output schema, the description appropriately describes the return format (summary, per-claim details, citations, quota). It also mentions the quota cost. However, it does not address error scenarios, retry behavior, or any edge cases like empty text (though schema enforces minLength). For a tool with this complexity and 5 parameters, the description is mostly complete but could mention what happens on failure.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 100% description coverage for all 5 parameters, so the schema itself provides clear meanings for each field. The tool description adds no additional parameter-specific context beyond the overall pipeline. Given the high schema coverage, a baseline of 3 is appropriate; the description does not need to compensate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: verifying statements in input text for hallucinations, with a specific pipeline (extract claims, verify each) and output types. It is specific about the verb and resource. However, it does not explicitly distinguish itself from sibling tools like verify_agent, though the name 'verify_text' suggests a text-focused scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It does not mention when it is appropriate or when to choose verify_agent or other siblings. The only usage-related note is that it consumes detect quota, which is a cost constraint but not a usage guideline.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
v0.1.0- First observed
check_cua_actions - First observed
check_safety - First observed
verify_agent - First observed
verify_text
TDQS
Scored across 4 tools
The tools are mostly distinct, but verify_text, check_safety (with optional hallucination detection), and verify_agent all overlap on claim-level hallucination verification. check_cua_actions is clearly separate, but an agent could struggle to choose between the three text/agent verification tools without reading descriptions closely.
All tool names follow a clear lowercase snake_case verb_noun pattern with either check_ or verify_ prefixes. The naming is predictable, though the repeated verbs slightly undercut distinctiveness because verify_text and verify_agent share the same verb while having some overlapping functionality.
Four tools is a well-scoped surface for a safety/verification server. Each tool represents a distinct capability area: free-text verification, agent action safety, general prompt/output safety, and agent trajectory self-check.
The surface covers the core workflows: verifying claims, checking agent actions, scanning prompts/outputs for safety risks, and self-checking agent trajectories. Minor gaps exist—such as no dedicated quota/status tool or a more granular single-claim verification endpoint—but agents can work around these using the existing tools.
Maintenance
Related MCP Connectors
Security & DLP proxy for MCP: tool-poisoning scans, PII redaction on tool args/results. Beta.
Find, vet, and run MCP tools through a secure audited gateway with prompt-injection risk scoring
Security firewall for AI agents — scans MCP calls for injection, secrets, and risks.
Zero-secret MCP gateway for AI agents: risk-scored, audited calls with human-in-the-loop approval.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceEnables content moderation by checking user input against OpenAI's moderation API via Google ADK, with dual SSE/STDIO transport support and integration with local LLMs.4MIT
- AlicenseAqualityAmaintenanceMCP server exposing Verity's trust checks as tools (verify_fact, detect_injection, moderate_content, redact_pii, guard_action) that any MCP-capable agent can discover and call for fail-closed safety verification before spending or sending.827 PyPIMIT
- AlicenseNot gradedqualityCmaintenanceEnables secure interoperability between LLM agents and MCP tool servers by sanitizing requests and responses, masking sensitive tokens, detecting PII, and performing server reputation scans.MIT
- FlicenseNot gradedqualityCmaintenanceEnables safe use of any MCP server by proxying and live-scanning all tool requests and responses, blocking or redacting poison descriptions, indirect prompt injection, malicious arguments, and unauthorized destinations.-