grounded-rag-mcp
# grounded-rag-mcp
An **MCP server that gives any LLM host grounded, cited retrieval over your own documents** — hybrid retrieval (BM25 + dense), cross-encoder reranking, citations, and a built-in eval harness.
> Point it at a folder of documents. Your MCP host (Claude Desktop, an IDE, a custom agent) can then `search` and `answer` over them — grounded in the real text, with citations, and an honest "not in the documents" path.
[](https://github.com/chetan1521/grounded-rag-mcp/actions/workflows/ci.yml)
---
## Why
Most RAG-over-MCP examples are toys. This one is built production-flavored:
- **Hybrid retrieval** — BM25 (exact terms) + dense (semantics), fused with Reciprocal Rank Fusion.
- **Cross-encoder reranking** — precision on the top candidates without blowing latency.
- **Grounding + citations** — answers cite their sources; if the answer isn't in the docs, it says so.
- **Built-in eval** — measure retrieval quality (recall@k, MRR, hit-rate), not just vibes.
- **Local-first** — the default path runs with no external services or API keys.
- **Both transports** — stdio and Streamable HTTP.
## Status
🚧 Early development. Building in public, phase by phase (see `PROJECT_REQUIREMENTS.md`).
- [x] Phase 0 — scaffold, packaging, CI
- [x] Phase 1 — core retrieval (chunk → embed → BM25 + dense → RRF)
- [x] Phase 2 — MCP server (stdio) with `ingest` / `search`
- [x] Phase 3 — rerank + grounding + `answer` (via MCP sampling)
- [x] Phase 4 — tests, types, docs, resource + prompt
- [x] Phase 5 — Streamable HTTP transport + `evaluate_retrieval`
- [ ] Phase 6 — publish to PyPI
## Install
```bash
pip install grounded-rag-mcp # lean, local-first default (no torch)
pip install "grounded-rag-mcp[st]" # + sentence-transformers for semantic embeddings & reranking
```
## Tools
| Tool | What it does |
|---|---|
| `ingest_documents` | Chunk, embed, and index files or raw text into a named collection |
| `search` | Hybrid / dense / bm25 retrieval, optional rerank, per-stage scores |
| `answer` | Grounded, cited answer via MCP sampling; refuses when nothing is found |
| `list_collections` | List collections and chunk counts |
| `evaluate_retrieval` | hit_rate / MRR / recall@k on labeled cases |
Also exposes a resource (`rag://collections`) and a prompt (`grounded_answer`).
## Use it with an MCP host (e.g. Claude Desktop)
Add to your host's MCP config:
```json
{
"mcpServers": {
"grounded-rag": {
"command": "grounded-rag-mcp"
}
}
}
```
Or run it directly:
```bash
grounded-rag-mcp # stdio (default, for local hosts)
grounded-rag-mcp --http # Streamable HTTP on 127.0.0.1:8000 (remote / multi-client)
```
## Use the retrieval engine as a Python library
```python
from grounded_rag_mcp.collection import Collection
from grounded_rag_mcp.embeddings import HashingEmbedder
from grounded_rag_mcp.ingest import load_texts
from grounded_rag_mcp.config import RetrievalConfig
col = Collection("kb", HashingEmbedder(dim=512))
col.add(load_texts(["The refund policy allows returns within 30 days of purchase."]))
for hit in col.retrieve("refund policy", RetrievalConfig(top_k=1)):
print(hit.chunk.source, hit.score, hit.stage_scores)
```
## Development
```bash
pip install -e ".[dev]"
ruff check . && ruff format --check . && mypy src && pytest -q
```
See [docs/ARCHITECTURE.md](docs/ARCHITECTURE.md), [docs/BUILD_STORY.md](docs/BUILD_STORY.md),
and [PUBLISHING.md](PUBLISHING.md).
## License
MIT © Chetan C
TDQS
Scored across 5 tools
Each tool targets a distinct operation: ingestion, search, grounded answering, collection listing, and retrieval evaluation. There is no overlap in purpose, and descriptions clearly delineate when to use each.
Tool names follow a clear imperative, snake_case style. Most use verb_noun (ingest_documents, list_collections, evaluate_retrieval), though search and answer are bare verbs rather than verb_noun, creating a minor inconsistency.
Five tools is well-scoped for a grounded RAG server: ingest, search, answer, list collections, and evaluate retrieval. Each tool covers a necessary part of the workflow without redundancy or bloat.
Core RAG workflows are covered end-to-end, including ingestion, retrieval, grounded answering, and quality evaluation. The main gap is lifecycle management: there is no way to delete or update documents or collections once ingested.