Skip to main content
Glama
README.md
# mcp-verify

A benchmarked hallucination detector. Give it a SOURCE (the source of truth) and a
DRAFT; it returns the specific claims the source doesn't support, each with
`{claim, reason, source_fact_checked, category, severity}`.

**Benchmarked:** 106 labeled cases across 4 domains + an adversarial red-team
tranche. Over 3 independent runs: precision 97.9% [96.6, 99.1], recall 99.1%,
F1 98.5% (95% CIs) — with every residual failure adjudicated and documented.
See [BENCHMARK.md](BENCHMARK.md).

```python
from mcp_verify import build_default_client, verify

client = build_default_client()  # reads ANTHROPIC_API_KEY
report = verify(client, source="...source of truth...", draft="...text to check...")
for f in report.failures:
    print(f.severity, "|", f.claim, "—", f.reason)

# Reliability mode: run N times, keep only majority-confirmed claims;
# unstable ones land in report.uncertain instead of report.failures.
report = verify(client, source="...", draft="...", consistency=3)
```

## Run as MCP server

One tool, `verify(source, draft, consistency=1)`, over stdio:

```bash
pip install mcp-verify   # or: pip install -e . from a checkout
claude mcp add verify -- mcp-verify
```

Two modes, auto-selected (`MCP_VERIFY_MODE=api|sampling` overrides):

- **sampling** (default, no key): the server asks the HOST's model to do the
  verification via MCP sampling — no ANTHROPIC_API_KEY, no extra cost.
- **api**: if `ANTHROPIC_API_KEY` is set in the server's environment, it calls
  the pinned benchmark model directly — the exact path BENCHMARK.md measures.
  `claude mcp add verify -e ANTHROPIC_API_KEY=sk-... -- mcp-verify`

Both return the same JSON: `{"passed", "failures": [...], "uncertain": [...], "mode"}`.

## The benchmark

The eval suite (106 labeled cases across 4 domains plus a red-team tranche, a
15-type hallucination taxonomy, and a confusion matrix) travels with the
detector. See [BENCHMARK.md](BENCHMARK.md).

```bash
# Offline checks (no API key, no cost):
bash check.sh                      # lint + full test suite
python -m eval.analysis --sample   # failure analysis on a sample report

# Live eval (needs ANTHROPIC_API_KEY):
python -m eval.run                 # full suite
python -m eval.run --smoke         # 5 representative cases
python -m eval.run --judge         # LLM-judge matcher (meaning, not tokens)
python -m eval.ablation            # model/thinking/token-budget ablation
```

TDQS

A4.4/5.0

Scored across 1 tool

Disambiguation5/5

With only one tool, there is no risk of ambiguity or confusion between tools. The 'verify' tool has a clear, singular purpose.

Naming Consistency5/5

A single tool cannot have naming inconsistencies. The name 'verify' is a clear verb that directly describes its action.

Tool Count3/5

The single tool is at the low end of the typical range. While the server is highly focused, a single tool feels thin for a standalone server, as users might expect additional related functionality.

Completeness3/5

The server provides a single verification operation, which is complete for that narrow task, but lacks any supporting tools (e.g., source management, draft editing) that would make it a comprehensive verification service.

Maintenance

ActivityStale
ResponsivenessNo issues