groundcheck-mcp
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@groundcheck-mcpverify this quote against https://example.com/paper: "grounding verification with no LLM""
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Groundcheck, verification with no model in the loop
An MCP connector that checks whether a claim's grounding is real: the quote
is actually on the page, the arXiv id resolves, the code prints what it's said
to, the number is right, with no language model anywhere in the verification
path. Works in any MCP host: Claude, Gemini, or another. Every verdict it
returns is recomputed independently by the programs in verify/, and CI fails
the build if any of them disagrees.
Abstract
Assistant "fact-checking" almost always means asking a second model whether the first was right. That relocates the error rather than removing it, because the checker hallucinates too. This is an MCP connector that verifies grounding against reality instead: five tools that fetch a page, resolve an identifier, execute a snippet, grep a codebase or evaluate an expression, and return one of three verdicts with the concrete evidence attached.
The scope is deliberately narrow and stated as such. Groundcheck confirms that the
evidence a claim rests on is real and says what it is quoted to say. It does not
judge whether a claim is semantically true, "this quote is on the cited page" is
checkable, "the page's argument is correct" is not, and it returnsunverifiable
rather than guessing.
No language model is involved in any verdict.
Contributions. (i) Verification grounded in sources rather than in a second
model's opinion. (ii) A three-verdict contract with an explicitunverifiable, so
refusal is a first-class outcome. (iii) Auditable results, every verdict carries
the quote, stdout, matching line or computed value it was based on.
Related MCP server: math-logic-mcp
1. Why this, and why it's hard
Every "fact-check" built into an assistant today ultimately asks a second model whether the first one was right. That doesn't verify anything, it relocates the error, because the checker hallucinates too. The genuinely hard, under-attempted thing is verification grounded in reality rather than in another model's opinion. That's all this does, and it does only that.
Scope, stated honestly, because over-claiming would defeat the point.
Groundcheck confirms that the evidence a claim rests on is real and says what
it's quoted to say. It does not judge whether a claim is semantically true
"this quote is on the cited page" is checkable; "the page's argument is correct"
is not, and no amount of pretending makes it so. Every tool returns one of three
verdicts, and it saysunverifiable rather than guess:
verdict | meaning |
| the grounding was confirmed against a real source |
| the source exists and contradicts the claim (wrong number, missing quote, dead id, failing code) |
| no source, or it needs judgement this tool refuses to fake |
flowchart LR
A["Assistant makes<br/>a claim"] --> M{"Groundcheck<br/>MCP server"}
M --> Q["check_quote<br/><i>fetch the page</i>"]
M --> C["check_citation<br/><i>query arXiv / Crossref</i>"]
M --> X["check_code<br/><i>run in a subprocess</i>"]
M --> R["check_repo<br/><i>grep the files</i>"]
M --> T["check_math<br/><i>evaluate an AST</i>"]
Q & C --> NET[("the live web")]
X & R --> LOCAL[("local filesystem<br/>and interpreter")]
T --> PURE[("arithmetic")]
NET & LOCAL & PURE --> V{"verdict"}
V --> OK["checked"]
V --> NO["refuted"]
V --> UNK["unverifiable"]
classDef src fill:#4d4d4d,stroke:#2b2b2b,color:#fff
classDef good fill:#1a9850,stroke:#0f6b33,color:#fff
classDef bad fill:#b2182b,stroke:#7f0f20,color:#fff
classDef warn fill:#f4a582,stroke:#c06a4f,color:#000
class NET,LOCAL,PURE src
class OK good
class NO bad
class UNK warnNo box in that diagram is a language model. Every verdict terminates at something that can be looked up, executed or computed.
2. The tools
tool | verifies | how (no LLM) |
| an exact quote is on a page | fetch the page, match the text |
| an arXiv id or DOI resolves | query arXiv / Crossref, return the real title |
| code prints what's claimed | run it in a subprocess, compare stdout |
| a string/regex is in a codebase | grep the files, return real matching lines |
| arithmetic is correct | evaluate an AST (no |
Every result is{status, method, evidence, detail}``evidence is the concrete
thing found (the quote, the stdout, the matching line, the computed value), so a
verdict is auditable, not a black box.
2.1 A live run

Every row above is an actual call to the same function the server exposes,
including the network-dependent arXiv lookups. The refutations are real
refutations rather than illustrations of one, the first row is the3.7 x 1400
error from section 3, reproduced.
2.2 The case corpus
verify/export_cases.py writes these tables, and they are tracked so that CI
can corrupt one and require the harness to notice.
table | cases | checked | refuted | unverifiable |
| 27 | 17 | 3 | 7 |
| 10 | 6 | 2 | 2 |
| 11 | 6 | 4 | 1 |
3. It caught a mistake in its own author's work
check_citation exists because fabricated-but-plausible arXiv ids kept slipping
into research write-ups, an id that looks right and resolves to nothing.
check_math exists because3.7 × 1400 was written as8880 in a hardware deck
(it's 5180).check_repo is the generalisation of a profile-README claim-checker
that verifies every quoted number against its source repo. Each tool is a failure
that actually happened, turned into a check.
There's one honest wrinkle worth reporting: while testing, I assumed arXiv
2606.01992 was fabricated and expectedrefuted, the tool returnedchecked.
The tool was right and I was wrong: it's a real June-2026 paper. The verifier
did its job against my own bad assumption, which is the entire reason to ground
verification in a source rather than a hunch.
4. Use it
pip install -e . # or: pip install -r requirements.txt
python -m pytest tests/ # 19 tests, no network needed (mocked transport)Claude / Claude Code: add to your MCP config:
{
"mcpServers": {
"groundcheck": { "command": "python", "args": ["-m", "src.groundcheck.server"] }
}
}Gemini CLI / any MCP host: same stdio server; point your host's MCP config
atpython -m src.groundcheck.server. MCP is the reason one connector serves
both.
5. Security
check_code executes the code you give it in a subprocess. It uses list-form
subprocess (no shell, so nothing to inject) and kills on timeout, but it is
not sandboxed from the network or filesystem. Only pass code you would run
yourself. The other four tools are read-only (HTTP GET, file read, arithmetic).
6. Limitations
Grounding, not truth. By design, see Scope above.
Quote matching is exact (whitespace-normalised). A paraphrase that means the same thing returns
refuted, because "means the same" needs a judge and a judge is what this tool refuses to be. Match the literal text.JS-rendered pages.
check_quotereads the served HTML; a quote injected by client-side JavaScript won't be found. It fails safe (refuted), never a falsechecked.arXiv/Crossref only for citations. Other registries aren't wired up yet.
7. Licence
MIT, see LICENSE.
This server cannot be deployed
Maintenance
Related MCP Connectors
Experimental MCP server for current empirical verification of explicit public HTTPS endpoint claims.
MCP server for progressive tool usage at any scale (see https://klavis.ai)
MCP server providing access to the Scorecard API to evaluate and optimize LLM systems.
Related MCP Servers
- AlicenseAqualityBmaintenanceMCP server for verifying AI agent claims vs reality — single-transcript inline grounding-check that flags when an agent's response states facts not in the input context, when its code silently swallows exceptions and substitutes mock data, or when its multi-turn transcript contains contradictions or unverified completion claims. Sub-second, local, free, no API calls.41MIT
- AlicenseAqualityDmaintenanceMCP server that gives small LLMs verified symbolic-math & logic tools.61Apache 2.0
- AlicenseNot gradedqualityCmaintenanceAn MCP server that verifies whether a claim is actually supported by the source text at a given citation — independent of what the calling LLM asserts.MIT
- AlicenseNot gradedqualityAmaintenanceProvides an MCP server for storing and querying knowledge as verifiable claims, enforcing evidence-backed assertions with exact quotes and refusing paraphrases or unsupported relations.Apache 2.0