Skip to main content
Glama
chetan1521

grounded-rag-mcp

by chetan1521

grounded-rag-mcp

An MCP server that gives any LLM host grounded, cited retrieval over your own documents — hybrid retrieval (BM25 + dense), cross-encoder reranking, citations, and a built-in eval harness.

Point it at a folder of documents. Your MCP host (Claude Desktop, an IDE, a custom agent) can then search and answer over them — grounded in the real text, with citations, and an honest "not in the documents" path.

CI


Why

Most RAG-over-MCP examples are toys. This one is built production-flavored:

  • Hybrid retrieval — BM25 (exact terms) + dense (semantics), fused with Reciprocal Rank Fusion.

  • Cross-encoder reranking — precision on the top candidates without blowing latency.

  • Grounding + citations — answers cite their sources; if the answer isn't in the docs, it says so.

  • Built-in eval — measure retrieval quality (recall@k, MRR, hit-rate), not just vibes.

  • Local-first — the default path runs with no external services or API keys.

  • Both transports — stdio and Streamable HTTP.

Status

🚧 Early development. Building in public, phase by phase (see PROJECT_REQUIREMENTS.md).

  • Phase 0 — scaffold, packaging, CI

  • Phase 1 — core retrieval (chunk → embed → BM25 + dense → RRF)

  • Phase 2 — MCP server (stdio) with ingest / search

  • Phase 3 — rerank + grounding + answer (via MCP sampling)

  • Phase 4 — tests, types, docs, resource + prompt

  • Phase 5 — Streamable HTTP transport + evaluate_retrieval

  • Phase 6 — publish to PyPI

Install

pip install grounded-rag-mcp            # lean, local-first default (no torch)
pip install "grounded-rag-mcp[st]"      # + sentence-transformers for semantic embeddings & reranking

Tools

Tool

What it does

ingest_documents

Chunk, embed, and index files or raw text into a named collection

search

Hybrid / dense / bm25 retrieval, optional rerank, per-stage scores

answer

Grounded, cited answer via MCP sampling; refuses when nothing is found

list_collections

List collections and chunk counts

evaluate_retrieval

hit_rate / MRR / recall@k on labeled cases

Also exposes a resource (rag://collections) and a prompt (grounded_answer).

Use it with an MCP host (e.g. Claude Desktop)

Add to your host's MCP config:

{
  "mcpServers": {
    "grounded-rag": {
      "command": "grounded-rag-mcp"
    }
  }
}

Or run it directly:

grounded-rag-mcp            # stdio (default, for local hosts)
grounded-rag-mcp --http     # Streamable HTTP on 127.0.0.1:8000 (remote / multi-client)

Use the retrieval engine as a Python library

from grounded_rag_mcp.collection import Collection
from grounded_rag_mcp.embeddings import HashingEmbedder
from grounded_rag_mcp.ingest import load_texts
from grounded_rag_mcp.config import RetrievalConfig

col = Collection("kb", HashingEmbedder(dim=512))
col.add(load_texts(["The refund policy allows returns within 30 days of purchase."]))

for hit in col.retrieve("refund policy", RetrievalConfig(top_k=1)):
    print(hit.chunk.source, hit.score, hit.stage_scores)

Development

pip install -e ".[dev]"
ruff check . && ruff format --check . && mypy src && pytest -q

See docs/ARCHITECTURE.md, docs/BUILD_STORY.md, and PUBLISHING.md.

License

MIT © Chetan C