Skip to main content
Glama
Tausifonly001

FactAnchor-MCP

🔗 FactAnchor-MCP

Reduce AI Hallucinations by up to ~80% (informal estimate) using Local Context Anchoring.

License: MIT Python Zero Cost MCP

FactAnchor-MCP is a zero-cost, fully local Model Context Protocol server that grounds your AI assistant (Claude Desktop, Cursor, VS Code, Claude Code) in real, fetched web text — and forces it to answer only from that text.

  • 💸 ₹0 Hosting Cost — runs entirely on your machine. No cloud, no paid API keys.

  • 🛡️ Strict Guardrails — the LLM must cite sources or say "I cannot find a verified source for this information."

  • 🔎 Free Web Fetching — uses DuckDuckGo's free search + page scraping (no Serper/Google keys).

  • 🧠 Smart Chunking — BM25 semantic relevance scoring keeps only the most useful paragraphs.

  • 💾 Persistent Cache — SQLite disk cache (~/.factanchor/cache.db) survives server restarts.

  • Zero-config setuppip install -e . + connect your MCP client. Browser auto-installs on first start.

If FactAnchor-MCP helps you ship more reliable, hallucination-free AI, please consider starring the repository. It takes one click and helps more developers discover a truly zero-cost way to ground their agents. Thank you! 🙏


🚀 1-Minute Quick Start

git clone https://github.com/Tausifonly001/FactAnchor-MCP.git
cd FactAnchor-MCP
pip install -e .          # installs the `factanchor-mcp` command

Then add this to your claude_desktop_config.json (Settings → Developer → Edit Config):

{
  "mcpServers": {
    "FactAnchor-MCP": {
      "command": "factanchor-mcp"
    }
  }
}

Option B — Simple (use the script path directly)

git clone https://github.com/Tausifonly001/FactAnchor-MCP.git
cd FactAnchor-MCP
pip install -r requirements.txt

Add the absolute path to server.py:

{
  "mcpServers": {
    "FactAnchor-MCP": {
      "command": "python",
      "args": ["/absolute/path/to/FactAnchor-MCP/server.py"]
    }
  }
}
  • Windows: "C:\\Users\\you\\FactAnchor-MCP\\server.py"

  • macOS / Linux: "/Users/you/FactAnchor-MCP/server.py" or "/home/you/FactAnchor-MCP/server.py"

3. Restart your client

You'll now see the fetch_verified_context tool available. Ask a factual question and watch the assistant ground its answer in live, cited sources.

Zero-config: on first launch, FactAnchor silently installs the Playwright Chromium browser in the background. No crawl4ai-setup or manual browser commands required.

💡 Want uv instead of pip? uv pip install -e . works identically, and the factanchor-mcp command lands on your PATH.


Related MCP server: Grounded Code MCP

🧩 Supported Clients (drop-in configs)

FactAnchor-MCP is a standard MCP server, so it works with any MCP-compatible client. Below are ready-to-paste configs. Every client uses the same two shapes:

  • Option A (recommended): "command": "factanchor-mcp" — needs pip install -e . (so the command is on your PATH).

  • Option B (path-based): "command": "python" + "args": ["/abs/path/server.py"] — use this if the factanchor-mcp command isn't found.

File: claude_desktop_config.json (Settings → Developer → Edit Config)

{
  "mcpServers": {
    "FactAnchor-MCP": { "command": "factanchor-mcp" }
  }
}

Restart Claude Desktop. Tool appears in the tools list.

File: .opencode.jsonc (project root)

{
  "mcpServers": {
    "FactAnchor-MCP": { "command": "factanchor-mcp" }
  }
}

Verify with /mcpfetch_verified_context should be listed.

These are Claude-Code-style clients. Use a project .mcp.json:

{
  "mcpServers": {
    "FactAnchor-MCP": { "command": "factanchor-mcp" }
  }
}

Or add it from the CLI (runs the same server):

claude mcp add factanchor -- factanchor-mcp
# Kimi/Qwen/Cline equivalents use the same `mcp add` subcommand

File: ~/.cursor/mcp.json (global) or .cursor/mcp.json (project)

{
  "mcpServers": {
    "FactAnchor-MCP": { "command": "factanchor-mcp" }
  }
}

Enable it in Settings → MCP and restart Cursor.

File: .vscode/mcp.json (note: VS Code uses a "servers" key)

{
  "servers": {
    "FactAnchor-MCP": {
      "type": "stdio",
      "command": "factanchor-mcp"
    }
  }
}

Open the Command Palette → MCP: List Servers to confirm it's connected.

File: .gemini/settings.json

{
  "mcpServers": {
    "FactAnchor-MCP": { "command": "factanchor-mcp" }
  }
}

Or: gemini mcp add factanchor -- factanchor-mcp

If the factanchor-mcp command isn't on your PATH, use the absolute path to server.py on every client above:

{
  "mcpServers": {
    "FactAnchor-MCP": {
      "command": "python",
      "args": ["/absolute/path/to/FactAnchor-MCP/server.py"]
    }
  }
}

Per-OS path examples:

  • Windows: "C:\\Users\\you\\FactAnchor-MCP\\server.py"

  • macOS: "/Users/you/FactAnchor-MCP/server.py"

  • Linux: "/home/you/FactAnchor-MCP/server.py"


🛠️ How It Works

[Claude Desktop / Cursor / VS Code]
       │ (Asks a factual query)
       ▼
[FactAnchor MCP Server]
       │
       ├──► [Free DuckDuckGo Search] (URL Discovery)
       │
       ├──► [Persistent Cache: ~/.factanchor/cache.db] (Fast Repeat Queries)
       │
       ├──► [Crawl4AI in Isolated Subprocess] (Page Scraping)
       │
       ├──► [BM25 Semantic Chunking] (Smart Paragraph Selection)
       │
       └──► [Strict Guardrail Injection] (Verified Context Block)
              │
              ▼
[Assistant answers ONLY from verified context → ~0% Hallucination]
  1. Free Searchfetch_verified_context(query) runs a free DuckDuckGo search to discover URLs. No API keys required.

  2. Persistent Caching — results are cached in ~/.factanchor/cache.db (SQLite) for 24 hours, so repeat queries return instantly even after server restarts.

  3. Page ScrapingCrawl4AI runs in an isolated subprocess (crawl_worker.py) to scrape discovered URLs into clean, LLM-optimized Markdown (navbars, ads, and footers auto-stripped).

  4. Semantic Chunking — BM25 relevance scoring extracts only the paragraphs most relevant to your query, maximizing information density within the context window.

  5. Guardrail Injection — the fetched text is wrapped in a strict directive (see guardrail.py):

    • Answer only from <verified_context>.

    • If unanswerable, reply exactly: "I cannot find a verified source for this information."

    • Cite every claim in brackets like [Source: ...].

    • Never fall back to pre-trained knowledge.

  6. Local-Only — the server uses the stdio transport, so all processing stays on your machine.


📦 Project Structure

File

Purpose

server.py

The MCP server + fetch_verified_context tool (FastMCP).

crawl_worker.py

Headless-scrape worker (Crawl4AI) run in an isolated subprocess for robust MCP stdio.

search_backends.py

Free DuckDuckGo search (no API keys required).

disk_cache.py

Persistent SQLite cache (~/.factanchor/cache.db) for repeat queries.

semantic_chunker.py

BM25 relevance scoring to extract the most useful paragraphs.

guardrail.py

The strict fact-anchoring prompt template.

text_cleaner.py

Markdown cleaning + truncation for Crawl4AI output.

pyproject.toml

Packaging + factanchor-mcp console command.

requirements.txt

Dependencies (mcp, ddgs/duckduckgo_search, crawl4ai).

claude_desktop_config.example.json

Copy-paste config snippet.


🧰 Requirements

  • Python 3.10+

  • Internet access (for the free search/scrape)

🔧 Tool Reference

fetch_verified_context(query: str, max_results: int = 3) -> str

Param

Default

Notes

query

The factual topic or question to ground.

max_results

3

Sources to pull (clamped 1–5).


🐛 Troubleshooting

  • command not found: factanchor-mcp → you used Option A but didn't pip install -e ., or your venv isn't on PATH. Use Option B (script path) instead.

  • Pages return only short snippets (first run) → Chromium is still installing in the background. Wait ~1–2 minutes and retry; subsequent runs are instant.

  • "Browser executable doesn't exist" on Linux → install OS deps once: sudo playwright install-deps chromium (or sudo apt install libnss3 libatk-bridge2.0-0 libdrm2 libxkbcommon0 libgbm1 libasound2).

  • Rate-limit system note from the tool → DuckDuckGo is throttling free search. Wait a few minutes and retry. The server never crashes; it returns a clean system note for the LLM.

  • Tool not appearing in client → restart the client fully after editing the config, and check its MCP/Developer panel for errors.


📈 Virality Strategy

  • Before vs After video (X/Twitter & LinkedIn): show the assistant hallucinating a fake npm feature, then enable FactAnchor-MCP and watch it correctly say "I cannot find a verified source for this information." Tag @AnthropicAI with #MCP and #AI.

  • Open-source launch: submit to the official MCP servers list and awesome-mcp collections.


📊 Evaluation

The "~80%" figure is an informal estimate. A small, hand-runnable eval set lives in eval/sample_queries.json — see eval/README.md for how to reproduce it. Contributions of more queries (or a CI assertion) are very welcome.

🤝 Contributing

See CONTRIBUTING.md. Keep it zero-cost and local-first.

📋 Success Metrics (v1.0)

⚠️ Honesty note: The "~80% reduction" is an informal estimate from manual testing against a small set of factual queries (see eval/sample_queries.json) — it is not a benchmarked or statistically validated result. The guardrail is a prompt directive, not a hard infrastructure constraint, so an LLM can occasionally drift from it in long conversations. FactAnchor reduces hallucination but does not eliminate it; always verify critical claims against the cited sources.

  • 🧪 Up to ~80% fewer made-up facts observed in informal test queries.

  • <3 min user setup time (clone → install → config).

  • ₹0.00 server maintenance bill.

📜 License

MIT © FactAnchor-MCP contributors.

Available Tools

1 tool
fetch_verified_contextA

Fetch real, source-of-truth web text for a topic and wrap it in a strict fact-anchoring guardrail so the LLM answers ONLY from it.

Use this tool BEFORE answering any factual question. Then answer the user strictly from the returned block and follow the embedded STRICT RULES (never guess, cite sources in brackets, do not use pre-trained knowledge for missing info).

Args: query: The factual topic or question to ground. max_results: How many web sources to pull (default 3, max 5).

ParametersJSON Schema
NameRequiredDescriptionDefault
queryYes
max_resultsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full burden. It explains the output format (verified_context block with STRICT RULES) and parameter constraints (max_results default/max). However, it does not explicitly state that the tool is read-only or non-mutating, nor mention failure modes or rate limits. This leaves minor ambiguity, but the overall behavior is well described.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with a purpose sentence, usage paragraph, and parameter definitions. It is front-loaded with the key action. While clear, it could be slightly more concise (e.g., combining the usage instructions). Every sentence is valuable, so minor trim could improve.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations or sibling tools, the description adequately covers all aspects: purpose, usage, parameters, and expected output format (<verified_context> block with rules). It lacks explicit mention of edge cases (e.g., no results) but is sufficient for an agent to understand and use the tool effectively.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate. It does so effectively: query is described as 'the factual topic or question to ground,' and max_results as 'how many web sources to pull (default 3, max 5).' This adds contextual meaning beyond the schema's titles and types, making the parameters fully understandable.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Fetch real, source-of-truth web text for a topic and wrap it in a strict fact-anchoring guardrail.' It specifies the action (fetch), resource (web text), and outcome (guardrail for fact-anchoring). With no sibling tools, differentiation is not needed.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly instructs: 'Use this tool BEFORE answering any factual question. Then answer the user strictly from the returned <verified_context> block.' It includes detailed rules (never guess, cite sources, avoid pre-trained knowledge), providing clear guidance on when and how to use the tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool updatev1.0.0
    • First observedfetch_verified_context

TDQS

A4.7/5.0

Scored across 1 tool

Disambiguation5/5

With only one tool, there is no possibility of confusion between tools. The single tool has a clearly distinct purpose.

Naming Consistency5/5

The single tool name 'fetch_verified_context' follows a clear snake_case convention and is descriptive. Consistency is not an issue with only one tool.

Tool Count4/5

One tool is slightly below the typical 3-15 range, but it is appropriate for this narrowly scoped server. The tool covers the entire workflow of fetching and anchoring factual context.

Completeness5/5

The single tool effectively covers the entire domain of fact-anchoring: fetching verified context and enforcing strict answer rules. No obvious gaps exist for its stated purpose.

Maintenance

ActivitySlowing
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    B
    maintenance
    A local MCP server that gives AI coding assistants retrieval access to your personal knowledge base of books, standards, and docs, grounding their answers in sources you trust.
    MIT
  • A
    license
    Not graded
    quality
    A
    maintenance
    Zero-auth multi-source research MCP server that enables web search, reading URLs, PDFs, GitHub repos, and querying Hacker News, Stack Overflow, Semantic Scholar, and YouTube transcripts without API keys.
    10
    Apache 2.0