Skip to main content
Glama
chetan1521

grounded-rag-mcp

by chetan1521
README.md
# grounded-rag-mcp

An **MCP server that gives any LLM host grounded, cited retrieval over your own documents** — hybrid retrieval (BM25 + dense), cross-encoder reranking, citations, and a built-in eval harness.

> Point it at a folder of documents. Your MCP host (Claude Desktop, an IDE, a custom agent) can then `search` and `answer` over them — grounded in the real text, with citations, and an honest "not in the documents" path.

[![CI](https://github.com/chetan1521/grounded-rag-mcp/actions/workflows/ci.yml/badge.svg)](https://github.com/chetan1521/grounded-rag-mcp/actions/workflows/ci.yml)

---

## Why

Most RAG-over-MCP examples are toys. This one is built production-flavored:

- **Hybrid retrieval** — BM25 (exact terms) + dense (semantics), fused with Reciprocal Rank Fusion.
- **Cross-encoder reranking** — precision on the top candidates without blowing latency.
- **Grounding + citations** — answers cite their sources; if the answer isn't in the docs, it says so.
- **Built-in eval** — measure retrieval quality (recall@k, MRR, hit-rate), not just vibes.
- **Local-first** — the default path runs with no external services or API keys.
- **Both transports** — stdio and Streamable HTTP.

## Status

🚧 Early development. Building in public, phase by phase (see `PROJECT_REQUIREMENTS.md`).

- [x] Phase 0 — scaffold, packaging, CI
- [x] Phase 1 — core retrieval (chunk → embed → BM25 + dense → RRF)
- [x] Phase 2 — MCP server (stdio) with `ingest` / `search`
- [x] Phase 3 — rerank + grounding + `answer` (via MCP sampling)
- [x] Phase 4 — tests, types, docs, resource + prompt
- [x] Phase 5 — Streamable HTTP transport + `evaluate_retrieval`
- [ ] Phase 6 — publish to PyPI

## Install

```bash
pip install grounded-rag-mcp            # lean, local-first default (no torch)
pip install "grounded-rag-mcp[st]"      # + sentence-transformers for semantic embeddings & reranking
```

## Tools

| Tool | What it does |
|---|---|
| `ingest_documents` | Chunk, embed, and index files or raw text into a named collection |
| `search` | Hybrid / dense / bm25 retrieval, optional rerank, per-stage scores |
| `answer` | Grounded, cited answer via MCP sampling; refuses when nothing is found |
| `list_collections` | List collections and chunk counts |
| `evaluate_retrieval` | hit_rate / MRR / recall@k on labeled cases |

Also exposes a resource (`rag://collections`) and a prompt (`grounded_answer`).

## Use it with an MCP host (e.g. Claude Desktop)

Add to your host's MCP config:

```json
{
  "mcpServers": {
    "grounded-rag": {
      "command": "grounded-rag-mcp"
    }
  }
}
```

Or run it directly:

```bash
grounded-rag-mcp            # stdio (default, for local hosts)
grounded-rag-mcp --http     # Streamable HTTP on 127.0.0.1:8000 (remote / multi-client)
```

## Use the retrieval engine as a Python library

```python
from grounded_rag_mcp.collection import Collection
from grounded_rag_mcp.embeddings import HashingEmbedder
from grounded_rag_mcp.ingest import load_texts
from grounded_rag_mcp.config import RetrievalConfig

col = Collection("kb", HashingEmbedder(dim=512))
col.add(load_texts(["The refund policy allows returns within 30 days of purchase."]))

for hit in col.retrieve("refund policy", RetrievalConfig(top_k=1)):
    print(hit.chunk.source, hit.score, hit.stage_scores)
```

## Development

```bash
pip install -e ".[dev]"
ruff check . && ruff format --check . && mypy src && pytest -q
```

See [docs/ARCHITECTURE.md](docs/ARCHITECTURE.md), [docs/BUILD_STORY.md](docs/BUILD_STORY.md),
and [PUBLISHING.md](PUBLISHING.md).

## License

MIT © Chetan C

TDQS

A4.2/5.0

Scored across 5 tools

Disambiguation5/5

Each tool targets a distinct operation: ingestion, search, grounded answering, collection listing, and retrieval evaluation. There is no overlap in purpose, and descriptions clearly delineate when to use each.

Naming Consistency4/5

Tool names follow a clear imperative, snake_case style. Most use verb_noun (ingest_documents, list_collections, evaluate_retrieval), though search and answer are bare verbs rather than verb_noun, creating a minor inconsistency.

Tool Count5/5

Five tools is well-scoped for a grounded RAG server: ingest, search, answer, list collections, and evaluate retrieval. Each tool covers a necessary part of the workflow without redundancy or bloat.

Completeness4/5

Core RAG workflows are covered end-to-end, including ingestion, retrieval, grounded answering, and quality evaluation. The main gap is lifecycle management: there is no way to delete or update documents or collections once ingested.

Maintenance

ActivityMaintained
ResponsivenessNo issues