mcp-redline
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@mcp-redlineverify that the contract requires 30 days notice"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
๐ต๐ฑ Wersja polska
mcp-redline
Deterministic Model Context Protocol (MCP) server for document citation and claim verification with verbatim quotes or explicit refusal.
๐ Live Demo: redline.robertgrabowski.com
Runs entirely in the browser with an enforced Content-Security-Policy header: connect-src 'none' (zero outbound network requests โ verifiable in DevTools).
Language models frequently assert false figures, dates, and contractual terms with high confidence. mcp-redline connects the model to a local evidential corpus and enforces that every factual assertion is either confirmed by a verbatim source quote, or explicitly flagged as contradicted or unsupported.
The verification engine uses no LLM, no vector database, and no external network calls: the same corpus and claim deterministically yield the same status, reason codes, and quote. When in doubt, it refuses.
Notice: All entities, individuals, registration numbers (KRS, Company No, VAT), and financial figures in the demo corpus are entirely fictional and designed specifically to benchmark anti-hallucination resilience.
30-Second Quick Start
Requirements: Node.js >= 18, npm >= 9.
git clone https://github.com/robrobgr/mcp-redline.git
cd mcp-redline
npm install
npm run build
npm test # 15 unit tests
npm run eval # 64 benchmark claimsRun Web Workbench Locally (100% Offline)
npm run demo
# Open http://localhost:3333 in any browser (works with Wi-Fi disabled)Related MCP server: science2code-mcp
MCP Tools
Tool | Parameters | Output |
| โ | Corpus files, section counts, and evidential authority tier (1 = contract/ledger, 2 = CRM, 3 = email) |
|
| Sentence- and table-row-level verbatim quotes: |
|
| Exact verbatim section text without paraphrase |
|
|
|
Verification Statuses
Status | Definition | Permitted Agent Action |
| A verbatim source unit matches the claim's concepts and numbers, shares polarity, and is not overridden by a higher tier. | Present as fact, citing the exact quote. |
| An equal- or higher-tier source contradicts the claim (negation, conflicting numeric value, currency, or year). | State that source contradicts the claim, providing counter-quote. |
| The corpus contains insufficient evidence to decide. | Refuse to assert as fact. |
Benchmark & Evaluation
Run the automated evaluation harness over 64 assertions:
npm run evalSplit | Claims (n) | 3-Class Accuracy | False GROUNDED | Facts Confirmed | Non-facts Refused |
| 10 | 90% (9/10) | 0 | 4/4 (100%) | 6/6 (100%) |
| 31 | 94% (29/31) | 0 | 18/18 (100%) | 13/13 (100%) |
| 23 | 87% (20/23) | 0 | 10/10 (100%) | 13/13 (100%) |
TOTAL | 64 | 91% (58/64) | 0 | 32/32 (100%) | 32/32 (100%) |
Safe Asymmetry: All 6 discrepancies out of 64 are conservative refusals (UNSUPPORTED instead of CONTRADICTED, or one CONTRADICTED instead of UNSUPPORTED). False-GROUNDED rate is strictly 0.
Contaminated Holdout Disclosure:
The initial holdout run scored 18/23 (78%) with 1 false GROUNDED (H19). After resolving architectural gaps (party binding and compound clause negation scope), the score reached 20/23 (87%) and 0 false GROUNDED. Because this holdout was authored by the same engineer and informed subsequent code revisions, it is documented as contaminated. Fully independent evaluation requires third-party authored assertions. See prompts_eval/HOLDOUT_LOG.md.
How verify Works
The engine contains no corpus-specific heuristics or hardcoded trap rules. Generality is validated against an independent synthetic corpus (lease agreement) in tests/server.test.ts.
Verbatim Units: Corpus is split into sentence units, key-value rows, and table lines. Every returned quote is an exact substring of the source file.
Weighted Concept Coverage: Rare terms carry higher weight (IDF), supplemented by a bilingual PL/EN domain lexicon and lightweight stemming. Proper nouns must be present in the quoted unit itself.
Number & Currency Matching: Numeric tokens are typed (money, percentage, year, count). Mismatched numbers or currencies produce immediate contradiction.
Negation Scope: Polarity is evaluated within the specific clause governing the matched terms. An affirmative mention elsewhere does not contradict a valid negative fact.
Reported Speech vs Fact: Mere proposals or meeting requests ("COO requested penalty...") do not confirm that an action occurred.
Party Binding: Defined entity tokens (
("Supplier" or "Apex Meridian")) prevent attributing terms of one party to another.Authority Tiers: Tier 1 (executed contracts, invoices, balance sheets) overrides Tier 2 (CRM) and Tier 3 (email correspondence).
Compound Claims: Up to 3 adjacent units of a single record (e.g. invoice number, amount, payment bank) may be combined. Prose sentences are never stitched together.
Limitations
Lexical matching, not deep semantic inference: Paraphrases outside the domain lexicon result in
UNSUPPORTED(a safe failure mode). Adapting to a new domain requires expandingGROUPSinsrc/engine.ts.Single assertion per verify: Sentences combining facts from distinct documents are not combined into a synthetic composite; agents should split them into separate
verifycalls.Table column layout: Numeric values are bound to table rows rather than individual header columns.
Tier extraction: Currently inferred from filename patterns (
Email,CRM, etc.) rather than metadata frontmatter.
Integration Configuration
Claude Desktop (claude_desktop_config.json)
{
"mcpServers": {
"mcp-redline": {
"command": "node",
"args": ["<path-to-repo>/dist/src/index.js"],
"env": {
"MCP_REDLINE_CORPUS_DIR": "<path-to-repo>/corpus",
"MCP_REDLINE_LANG": "en"
}
}
}
}Set MCP_REDLINE_LANG=pl to receive explanations in Polish while preserving identical machine-level statuses and reason codes.
Architecture & Diagrams
Visualizations of the verification mechanism and data boundary:
License
MIT ยท Copyright (c) 2026 Robert Grabowski.
Available Tools
4 toolslist_sourcesA
List all documents in the local corpus with their authority tier (1 = contract/ledger, 2 = CRM, 3 = correspondence)
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It usefully discloses that results carry an authority tier and defines the tier values (1=contract/ledger, 2=CRM, 3=correspondence), which is real domain context, but it says nothing about read-only nature, ordering, result size, or permissions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence, front-loaded with the verb and resource, with the tier legend appended compactly. Nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a parameterless list tool with no output schema, the description covers what is listed and partially what each result contains (authority tier). It is close to complete, though ordering/pagination and the exact shape of a 'document' entry remain unstated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4; there is no parameter syntax the description needs to compensate for. The authority-tier legend adds value about result content rather than input semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List) and resource (all documents in the local corpus) with scope ('all' implies exhaustive, not filtered). It does not explicitly contrast itself with siblings like search, but 'all documents' conveys the enumerating nature of the tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: an agent can infer this is the tool for enumerating everything rather than searching, but the description never states when to prefer it over search, quote, or verify, nor any exclusions or prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
quoteA
Retrieve the exact, verbatim text of a section (page) of a corpus file
| Name | Required | Description | Default |
|---|---|---|---|
| file | Yes | Filename in the corpus (e.g. 01_Apex_VeloNova_MSA_2023.md) | |
| page | Yes | Section number or title |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full disclosure burden. 'Exact, verbatim' usefully signals unmodified text (important for citation), but it says nothing about behavior when a page is missing, invalid, or out of range, nor about auth or output size limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler; the key differentiator ('exact, verbatim') appears early.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter read tool with no output schema, the description adequately characterizes the return value as verbatim section text. It is nearly complete, missing only error/edge-case behavior for invalid sections.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and both parameters are documented there with format hints, so the schema does the heavy lifting. The description adds only the 'section (page)' equivalence, which is marginal beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Retrieve) and a precisely bounded resource (exact, verbatim text of a section/page of a corpus file). The word 'verbatim' meaningfully separates it from sibling 'search', though the description never names the sibling outright.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: an agent can infer this is the tool for exact quotation rather than fuzzy matching, given the emphasis on 'exact, verbatim'. There is no explicit when/when-not statement or reference to 'search', 'list_sources', or 'verify', so the guidance is inferential rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
searchB
Search the corpus; returns verbatim sentence-level quotes with file, section and score
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | Search query |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It usefully discloses the return format (verbatim sentence-level quotes with file, section, and score), which substitutes for the absent output schema, but says nothing about ranking, result limits, pagination, or whether matching is lexical/semantic.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One tight sentence with the verb front-loaded and no filler. It could have absorbed a routing clause at no real length cost, but nothing present is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter search tool with no output schema, the description does supply the essential return shape, so it is minimally complete. It still omits any guidance on result volume, ranking, or how to follow up with quote/verify, leaving the agent to guess at the surrounding workflow.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and there is a single required "query" string, so the schema already documents the only parameter. The description adds no query syntax, operators, or phrasing guidance beyond what the schema provides, matching the baseline for high-coverage schemas.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ("Search the corpus") plus the shape of what comes back. It is clear on its own, but it never names or contrasts with the siblings quote/verify/list_sources, even though the return of "verbatim sentence-level quotes" sits close to what the quote sibling likely does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to use search versus quote, verify, or list_sources. An agent gets no condition, prerequisite, or exclusion to route on โ it must infer that search is the discovery step and quote/verify are follow-ups.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
verifyA
Deterministically verify a claim against the corpus (no LLM). Returns GROUNDED with a verbatim supporting quote, CONTRADICTED with a verbatim counter-quote from an equal-or-higher authority source, or UNSUPPORTED when nothing is decisive. Only GROUNDED may be presented as fact.
| Name | Required | Description | Default |
|---|---|---|---|
| claim | Yes | A single factual statement to verify |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and does so well: it discloses deterministic no-LLM execution, exact tri-state outcomes, quote requirements, authority comparison rules, and a critical presentation constraint. This is unusually complete for a verification tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences with no waste. The verification mechanism, outcome taxonomy, and presentation rule are front-loaded in a logical order.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema or annotations, the description must explain return values and interpretation, and it does so completely: GROUNDED, CONTRADICTED, and UNSUPPORTED are clearly defined. The one-parameter scope is unambiguous given the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single 'claim' parameter is already defined as 'A single factual statement to verify.' The description adds no syntax, format, or constraint beyond what the schema already provides, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb 'verify', resource 'claim against the corpus', and a distinctive mechanism ('no LLM') that separates it from search and quote siblings. It also names the three possible verdicts, making the tool's adjudicative purpose immediately clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool is for checking a single factual claim, but it does not explicitly say when to use it instead of search or quote. The rule 'Only GROUNDED may be presented as fact' is a downstream usage constraint rather than routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
v1.2.0- First observed
list_sources - First observed
quote - First observed
search - First observed
verify
TDQS
Scored across 4 tools
search (fuzzy retrieval returning ranked quotes), quote (exact section retrieval), and verify (claim adjudication) have distinguishable purposes, though search's returned quotes and quote's verbatim text create minor conceptual overlap. Boundaries are clarified by descriptions.
All names are lowercase snake_case and action-oriented, but list_sources is a verb_noun compound while search, quote, and verify are bare verbs, a minor deviation. The convention is still predictable and readable.
Four tools is a tight, well-scoped set for a read-only corpus retrieval and verification server, with each tool earning its place. It sits at the thin end of the ideal 3-15 range but is not under-scoped for the purpose.
The surface covers discovery (list_sources), retrieval (search, quote), and adjudication (verify), which is complete for a read-only corpus server. Minor gaps exist, such as no way to enumerate sections/pages within a document or filter search by authority tier.
Maintenance
Related MCP Connectors
Deterministic prompt-injection detector; signed, offline-verifiable verdicts. Not an LLM.
Fact-checks generated content against your sources of truth showing what to trust, change, & verify.
Verify claims with verdict, confidence & cited sources; batch verify, source checks, daily brief.
Cross-check a factual claim against a verified knowledge graph before you assert it. Never guesses.
Related MCP Servers
- FlicenseNot gradedqualityNot gradedmaintenanceA bounded evidence review engine that ingests documents, extracts evidence for a given claim, detects contradictions, and produces auditable evidence packets without hallucinations or open-web research.-
- AlicenseNot gradedqualityCmaintenanceEnables coding agents to verify verbatim quotes against a local corpus of scientific PDFs, returning exact matches with page locators or typed refusals without using an LLM.AGPL 3.0
- AlicenseAqualityAmaintenanceOpen-source verification for evidence-grounded AI. Runs deterministic grounding checks of AI outputs against source evidence, with no LLM judge.2172 npm2MIT
- AlicenseNot gradedqualityCmaintenanceEnables AI assistants to run local, model-backed decisions on short passages and routing choices, and to check whether exactly supplied text supports a scientific claim, returning the full probability distribution, review-policy status, and provenance details. All evaluation stays on the local machine with no cloud inference, and flagged results are surfaced for human review rather than treated as conclusions.MIT