Skip to main content
Glama

๐Ÿ‡ต๐Ÿ‡ฑ Wersja polska

mcp-redline

Deterministic Model Context Protocol (MCP) server for document citation and claim verification with verbatim quotes or explicit refusal.

๐ŸŒ Live Demo: redline.robertgrabowski.com
Runs entirely in the browser with an enforced Content-Security-Policy header: connect-src 'none' (zero outbound network requests โ€” verifiable in DevTools).

Language models frequently assert false figures, dates, and contractual terms with high confidence. mcp-redline connects the model to a local evidential corpus and enforces that every factual assertion is either confirmed by a verbatim source quote, or explicitly flagged as contradicted or unsupported.

The verification engine uses no LLM, no vector database, and no external network calls: the same corpus and claim deterministically yield the same status, reason codes, and quote. When in doubt, it refuses.

Notice: All entities, individuals, registration numbers (KRS, Company No, VAT), and financial figures in the demo corpus are entirely fictional and designed specifically to benchmark anti-hallucination resilience.


30-Second Quick Start

Requirements: Node.js >= 18, npm >= 9.

git clone https://github.com/robrobgr/mcp-redline.git
cd mcp-redline
npm install
npm run build
npm test        # 15 unit tests
npm run eval    # 64 benchmark claims

Run Web Workbench Locally (100% Offline)

npm run demo
# Open http://localhost:3333 in any browser (works with Wi-Fi disabled)

Related MCP server: science2code-mcp

MCP Tools

Tool

Parameters

Output

list_sources

โ€”

Corpus files, section counts, and evidential authority tier (1 = contract/ledger, 2 = CRM, 3 = email)

search

query

Sentence- and table-row-level verbatim quotes: {file, page, quote, score}

quote

file, page

Exact verbatim section text without paraphrase

verify

claim

status (GROUNDED / CONTRADICTED / UNSUPPORTED), reasonCodes: string[], quote, explanation

Verification Statuses

Status

Definition

Permitted Agent Action

GROUNDED

A verbatim source unit matches the claim's concepts and numbers, shares polarity, and is not overridden by a higher tier.

Present as fact, citing the exact quote.

CONTRADICTED

An equal- or higher-tier source contradicts the claim (negation, conflicting numeric value, currency, or year).

State that source contradicts the claim, providing counter-quote.

UNSUPPORTED

The corpus contains insufficient evidence to decide.

Refuse to assert as fact.


Benchmark & Evaluation

Run the automated evaluation harness over 64 assertions:

npm run eval

Split

Claims (n)

3-Class Accuracy

False GROUNDED

Facts Confirmed

Non-facts Refused

legacy

10

90% (9/10)

0

4/4 (100%)

6/6 (100%)

dev

31

94% (29/31)

0

18/18 (100%)

13/13 (100%)

holdout

23

87% (20/23)

0

10/10 (100%)

13/13 (100%)

TOTAL

64

91% (58/64)

0

32/32 (100%)

32/32 (100%)

NOTE

Safe Asymmetry: All 6 discrepancies out of 64 are conservative refusals (UNSUPPORTED instead of CONTRADICTED, or one CONTRADICTED instead of UNSUPPORTED). False-GROUNDED rate is strictly 0.

WARNING

Contaminated Holdout Disclosure:
The initial holdout run scored 18/23 (78%) with 1 false GROUNDED (H19). After resolving architectural gaps (party binding and compound clause negation scope), the score reached 20/23 (87%) and 0 false GROUNDED. Because this holdout was authored by the same engineer and informed subsequent code revisions, it is documented as contaminated. Fully independent evaluation requires third-party authored assertions. See prompts_eval/HOLDOUT_LOG.md.


How verify Works

The engine contains no corpus-specific heuristics or hardcoded trap rules. Generality is validated against an independent synthetic corpus (lease agreement) in tests/server.test.ts.

  1. Verbatim Units: Corpus is split into sentence units, key-value rows, and table lines. Every returned quote is an exact substring of the source file.

  2. Weighted Concept Coverage: Rare terms carry higher weight (IDF), supplemented by a bilingual PL/EN domain lexicon and lightweight stemming. Proper nouns must be present in the quoted unit itself.

  3. Number & Currency Matching: Numeric tokens are typed (money, percentage, year, count). Mismatched numbers or currencies produce immediate contradiction.

  4. Negation Scope: Polarity is evaluated within the specific clause governing the matched terms. An affirmative mention elsewhere does not contradict a valid negative fact.

  5. Reported Speech vs Fact: Mere proposals or meeting requests ("COO requested penalty...") do not confirm that an action occurred.

  6. Party Binding: Defined entity tokens (("Supplier" or "Apex Meridian")) prevent attributing terms of one party to another.

  7. Authority Tiers: Tier 1 (executed contracts, invoices, balance sheets) overrides Tier 2 (CRM) and Tier 3 (email correspondence).

  8. Compound Claims: Up to 3 adjacent units of a single record (e.g. invoice number, amount, payment bank) may be combined. Prose sentences are never stitched together.


Limitations

  • Lexical matching, not deep semantic inference: Paraphrases outside the domain lexicon result in UNSUPPORTED (a safe failure mode). Adapting to a new domain requires expanding GROUPS in src/engine.ts.

  • Single assertion per verify: Sentences combining facts from distinct documents are not combined into a synthetic composite; agents should split them into separate verify calls.

  • Table column layout: Numeric values are bound to table rows rather than individual header columns.

  • Tier extraction: Currently inferred from filename patterns (Email, CRM, etc.) rather than metadata frontmatter.


Integration Configuration

Claude Desktop (claude_desktop_config.json)

{
  "mcpServers": {
    "mcp-redline": {
      "command": "node",
      "args": ["<path-to-repo>/dist/src/index.js"],
      "env": {
        "MCP_REDLINE_CORPUS_DIR": "<path-to-repo>/corpus",
        "MCP_REDLINE_LANG": "en"
      }
    }
  }
}

Set MCP_REDLINE_LANG=pl to receive explanations in Polish while preserving identical machine-level statuses and reason codes.


Architecture & Diagrams

Visualizations of the verification mechanism and data boundary:


License

MIT ยท Copyright (c) 2026 Robert Grabowski.

Available Tools

4 tools
list_sourcesA

List all documents in the local corpus with their authority tier (1 = contract/ledger, 2 = CRM, 3 = correspondence)

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It usefully discloses that results carry an authority tier and defines the tier values (1=contract/ledger, 2=CRM, 3=correspondence), which is real domain context, but it says nothing about read-only nature, ordering, result size, or permissions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence, front-loaded with the verb and resource, with the tier legend appended compactly. Nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a parameterless list tool with no output schema, the description covers what is listed and partially what each result contains (authority tier). It is close to complete, though ordering/pagination and the exact shape of a 'document' entry remain unstated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline is 4; there is no parameter syntax the description needs to compensate for. The authority-tier legend adds value about result content rather than input semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (List) and resource (all documents in the local corpus) with scope ('all' implies exhaustive, not filtered). It does not explicitly contrast itself with siblings like search, but 'all documents' conveys the enumerating nature of the tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: an agent can infer this is the tool for enumerating everything rather than searching, but the description never states when to prefer it over search, quote, or verify, nor any exclusions or prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

quoteA

Retrieve the exact, verbatim text of a section (page) of a corpus file

ParametersJSON Schema
NameRequiredDescriptionDefault
fileYesFilename in the corpus (e.g. 01_Apex_VeloNova_MSA_2023.md)
pageYesSection number or title

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full disclosure burden. 'Exact, verbatim' usefully signals unmodified text (important for citation), but it says nothing about behavior when a page is missing, invalid, or out of range, nor about auth or output size limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; the key differentiator ('exact, verbatim') appears early.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter read tool with no output schema, the description adequately characterizes the return value as verbatim section text. It is nearly complete, missing only error/edge-case behavior for invalid sections.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and both parameters are documented there with format hints, so the schema does the heavy lifting. The description adds only the 'section (page)' equivalence, which is marginal beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Retrieve) and a precisely bounded resource (exact, verbatim text of a section/page of a corpus file). The word 'verbatim' meaningfully separates it from sibling 'search', though the description never names the sibling outright.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: an agent can infer this is the tool for exact quotation rather than fuzzy matching, given the emphasis on 'exact, verbatim'. There is no explicit when/when-not statement or reference to 'search', 'list_sources', or 'verify', so the guidance is inferential rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

verifyA

Deterministically verify a claim against the corpus (no LLM). Returns GROUNDED with a verbatim supporting quote, CONTRADICTED with a verbatim counter-quote from an equal-or-higher authority source, or UNSUPPORTED when nothing is decisive. Only GROUNDED may be presented as fact.

ParametersJSON Schema
NameRequiredDescriptionDefault
claimYesA single factual statement to verify

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and does so well: it discloses deterministic no-LLM execution, exact tri-state outcomes, quote requirements, authority comparison rules, and a critical presentation constraint. This is unusually complete for a verification tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences with no waste. The verification mechanism, outcome taxonomy, and presentation rule are front-loaded in a logical order.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema or annotations, the description must explain return values and interpretation, and it does so completely: GROUNDED, CONTRADICTED, and UNSUPPORTED are clearly defined. The one-parameter scope is unambiguous given the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single 'claim' parameter is already defined as 'A single factual statement to verify.' The description adds no syntax, format, or constraint beyond what the schema already provides, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb 'verify', resource 'claim against the corpus', and a distinctive mechanism ('no LLM') that separates it from search and quote siblings. It also names the three possible verdicts, making the tool's adjudicative purpose immediately clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the tool is for checking a single factual claim, but it does not explicitly say when to use it instead of search or quote. The rule 'Only GROUNDED may be presented as fact' is a downstream usage constraint rather than routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 4 tool updatesv1.2.0
    • First observedlist_sources
    • First observedquote
    • First observedsearch
    • First observedverify

TDQS

A3.6/5.0

Scored across 4 tools

Disambiguation4/5

search (fuzzy retrieval returning ranked quotes), quote (exact section retrieval), and verify (claim adjudication) have distinguishable purposes, though search's returned quotes and quote's verbatim text create minor conceptual overlap. Boundaries are clarified by descriptions.

Naming Consistency4/5

All names are lowercase snake_case and action-oriented, but list_sources is a verb_noun compound while search, quote, and verify are bare verbs, a minor deviation. The convention is still predictable and readable.

Tool Count4/5

Four tools is a tight, well-scoped set for a read-only corpus retrieval and verification server, with each tool earning its place. It sits at the thin end of the ideal 3-15 range but is not under-scoped for the purpose.

Completeness4/5

The surface covers discovery (list_sources), retrieval (search, quote), and adjudication (verify), which is complete for a read-only corpus server. Minor gaps exist, such as no way to enumerate sections/pages within a document or filter search by authority tier.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • F
    license
    Not graded
    quality
    Not graded
    maintenance
    A bounded evidence review engine that ingests documents, extracts evidence for a given claim, detects contradictions, and produces auditable evidence packets without hallucinations or open-web research.
    -
  • A
    license
    A
    quality
    A
    maintenance
    Open-source verification for evidence-grounded AI. Runs deterministic grounding checks of AI outputs against source evidence, with no LLM judge.
    2
    1
    72 npm
    2
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    Enables AI assistants to run local, model-backed decisions on short passages and routing choices, and to check whether exactly supplied text supports a scientific claim, returning the full probability distribution, review-policy status, and provenance details. All evaluation stays on the local machine with no cloud inference, and flagged results are surfaced for human review rather than treated as conclusions.
    MIT