claude-rag
by sangelastro
README.md
# claude-rag โ MCP RAG Server for Markdown Knowledge Bases
**Author**: [Sergio Angelastro](https://github.com/sangelastro) โ MIT License
MCP server that indexes a folder of `.md` files into a local SQLite vector store and exposes hybrid search (BM25 keyword + semantic embeddings) as Claude Code tools. The core RAG system (chunking, embedding, SQLite vector store, MCP server, hybrid search, retrieval eval) is original work by the author.
> **Hook system** (`hooks/`) inspired by [agd-memory](https://github.com/Pinperepette/agd-memory) by [@Pinperepette](https://github.com/Pinperepette) (MIT) โ see [CREDITS.md](CREDITS.md)
## Architecture

> ๐ **[Interactive diagram โ](docs/diagram.html)** ยท ๐ **[How `kb_search` finds an answer, step by step โ](docs/how_search_works.html)**
Three components working together:
| | What | When |
|---|---|---|
| **โ Indexing** | `.md` files โ Chunker โ Embedder โ SQLite + JSON | on startup / `kb_reindex()` |
| **โก MCP** | Claude calls `kb_search()` โ BM25 + cosine sim fused with RRF โ top-K chunks | explicit hybrid search |
| **โข Hook** | every prompt intercepted โ keyword score on `kb_chunks.json` โ auto-inject | automatic, ~10ms, no model |
**Stack**: Python ยท fastembed / ONNX Runtime (embedding `paraphrase-multilingual-MiniLM-L12-v2` ~120MB + reranker `mmarco-mMiniLMv2-L12` int8 ~120MB, CPU-only) ยท SQLite ยท MCP stdio
## Tools exposed
| Tool | Description |
|---|---|
| `kb_search(query, top_k=5)` | Hybrid search (BM25 + semantic, RRF) โ returns top-K chunks with source file and section. `KB_SEARCH_MODE=semantic` for cosine only |
| `kb_reindex(force=False)` | Re-indexes files modified since last run (mtime-based) |
| `kb_stats()` | Shows indexed files, chunk counts, last update timestamps |
| `kb_savings()` | Shows cumulative token savings: RAG chunks served vs full-file baseline, broken down by source (`mcp` / `hook`) |
## Setup
### 1. Clone
```bash
# Default layout: repo sits inside the KB folder
# KB files (.md) go in the parent directory
git clone https://github.com/sangelastro/claude-rag ~/.claude/my-kb/rag
```
Or clone anywhere and point to your KB folder via env var (see step 3).
### 2. Install dependencies
```bash
cd ~/.claude/my-kb/rag
pip install -r requirements.txt
```
On first run the model (`paraphrase-multilingual-MiniLM-L12-v2`, ~120MB) is downloaded automatically from HuggingFace. This model supports 50+ languages including Italian natively.
### 3. Register in Claude Code
Add to `~/.claude.json` under `mcpServers`:
```json
"my-kb": {
"command": "python",
"args": ["/absolute/path/to/rag/server.py"],
"env": {
"KB_RAG_DIR": "/absolute/path/to/your/kb/folder"
}
}
```
- **`KB_RAG_DIR`** โ folder containing your `.md` files (default: `../` relative to `server.py`)
- **`KB_RAG_DB`** โ SQLite database path (default: `kb.db` next to `server.py`)
- **`CHUNK_MAX_CHARS`** โ max chars per chunk before splitting (default: `400`)
- **`CHUNK_OVERLAP`** โ overlap in chars between consecutive sub-chunks (default: `80`)
If the repo is cloned inside the KB folder (as in the example above), both env vars can be omitted.
### 4. Restart Claude Code
The server starts automatically. On first launch it indexes all `.md` files in `KB_RAG_DIR`.
## Optional: Claude Code Hooks
Two hooks auto-inject KB context without explicit `kb_search` calls.
Approach inspired by [agd-memory](https://github.com/Pinperepette/agd-memory) (MIT).
| Hook | Script | What it does |
|---|---|---|
| `SessionStart` | `hooks/kb_session_start.py` | Injects KB table of contents at session start |
| `UserPromptSubmit` | `hooks/kb_recall.py` | Auto-injects top matching chunks before each prompt (keyword scoring, ~10ms, no model load) |
### Setup hooks
Copy the relevant sections from `hooks/hooks_example.json` into your `~/.claude/settings.json`, replacing the placeholder paths:
```json
{
"hooks": {
"SessionStart": [{
"matcher": "*",
"hooks": [{"type": "command", "command": "python /absolute/path/to/rag/hooks/kb_session_start.py"}]
}],
"UserPromptSubmit": [{
"matcher": "*",
"hooks": [{"type": "command", "command": "python /absolute/path/to/rag/hooks/kb_recall.py", "timeout": 5}]
}]
}
}
```
`kb_chunks.json` (used by `kb_recall.py`) is auto-generated next to `kb.db` on every reindex. No model loading in hooks โ scoring uses token overlap only.
Hook behaviour is tunable via env vars:
| Variable | Default | Description |
|---|---|---|
| `KB_RAG_HOOK_TOP_K` | `3` | Max chunks injected per prompt |
| `KB_RAG_HOOK_MIN_SCORE` | `0.15` | Minimum score to trigger injection |
| `KB_RAG_HOOK_MIN_WORDS` | `4` | Skip prompts shorter than N words |
| `KB_RAG_HOOK_TOKEN_BUDGET` | `6000` | Max chars injected (~4 chars/token) |
## File structure
```
rag/
โโโ server.py # MCP server
โโโ reranker.py # Local cross-encoder reranker + calibrated confidence
โโโ requirements.txt
โโโ .gitignore
โโโ README.md
โโโ hooks/
โ โโโ kb_session_start.py # SessionStart hook
โ โโโ kb_recall.py # UserPromptSubmit hook
โ โโโ hooks_example.json # Hook config template
โโโ eval/
โ โโโ eval_retrieval.py # Retrieval benchmark against a gold set
โ โโโ gold_set.example.json # Gold set format (the real one stays local)
โโโ docs/how_search_works.html # Interactive walkthrough of the search pipeline
โโโ architecture.html # Technical documentation
โโโ kb_rag_slides.html # Architecture slide deck
```
`kb.db`, `kb_chunks.json`, `eval/gold_set.json` and `eval/results.json` are generated locally and excluded from git: they contain KB content.
## How it works
1. **Chunking** โ each `.md` file is split on `##` headers; frontmatter is stripped; sections longer than `CHUNK_MAX_CHARS` (400) are further split into overlapping sub-chunks with `CHUNK_OVERLAP` (80) chars of context continuity
2. **Embedding** โ chunks are encoded with `paraphrase-multilingual-MiniLM-L12-v2` (384 dimensions, 50+ languages)
3. **Storage** โ vectors stored as `float32` BLOBs in SQLite + `kb_chunks.json` for hooks
4. **Search** โ hybrid: BM25 over an in-memory inverted index + cosine similarity in numpy, the two rankings (top 50 each) fused with Reciprocal Rank Fusion; top-K returned. BM25 catches exact identifiers (table names, ports, hostnames) that a 128-token embedding model blurs; embeddings catch paraphrased natural-language questions
5. **Rerank + confidence** โ the top 10 hybrid candidates are rescored by a local multilingual cross-encoder (ONNX, int8, see `reranker.py`): one forward pass per (query, chunk) pair, no text generated. The top score is mapped to a calibrated probability (Platt scaling) that the answer is among the results; below `KB_MIN_CONFIDENCE` the tool says so instead of silently returning noise
6. **Hooks** โ keyword scoring on `kb_chunks.json` (no model), injected before each prompt
7. **Invalidation** โ mtime-based: only modified files are re-indexed on startup
8. **Savings tracking** โ every search records chars served vs full-file baseline in `search_stats` table; `kb_savings()` aggregates the cumulative token reduction without re-reading any file
## Environment variables
| Variable | Default | Description |
|---|---|---|
| `KB_RAG_DIR` | `../` (relative to `server.py`) | Folder with `.md` files to index |
| `KB_RAG_DB` | `./kb.db` (next to `server.py`) | SQLite database path |
| `KB_RAG_NAME` | `kb-rag` | MCP server name |
| `CHUNK_MAX_CHARS` | `400` | Max chars per chunk; longer sections are split into overlapping sub-chunks |
| `CHUNK_OVERLAP` | `80` | Overlap chars between adjacent sub-chunks to preserve context continuity |
| `KB_SEARCH_MODE` | `hybrid` | `hybrid` (BM25 + semantic, RRF) or `semantic` (cosine only, behaviour up to 1.1.0) |
| `KB_RERANK` | `mmarco-mminilm` | Local cross-encoder that reorders the hybrid candidates: `mmarco-mminilm` (AVX-512 int8), `mmarco-mminilm-avx2` (older CPUs), `bge-m3` (slower, not calibrated) or `off`. If it cannot load, search falls back to hybrid |
| `KB_RERANK_CANDIDATES` | `10` | Hybrid candidates passed to the reranker |
| `KB_RERANK_THREADS` | `4` | ONNX Runtime threads for the reranker |
| `KB_RERANK_CACHE` | `~/.cache/claude-rag/models` | Where reranker models are downloaded |
| `KB_MIN_CONFIDENCE` | `0.5` | Below this calibrated confidence, `kb_search` warns that the answer is probably not in the KB |
## Evaluating retrieval
`eval/eval_retrieval.py` measures how often the right section comes back, on a gold set of real queries labelled with their correct `file` + `section`. It compares semantic, hook lexical, BM25, hybrid, the actual `server.rank()` and (optionally) a local cross-encoder reranker. It only reads `kb.db` and never writes to `search_stats`.
```bash
cp eval/gold_set.example.json eval/gold_set.json # then write queries about your KB
py eval/eval_retrieval.py --no-rerank
```
Include unanswerable queries (`"kind": "neg"`): the script also reports whether a score threshold can tell "not in the KB" apart from a real hit. `--calibrate` fits the Platt parameters of each reranker with 5-fold cross-validation; copy `platt_all` into `reranker.py`.
Measured on the author's KB (5,241 chunks, 50 answerable + 7 unanswerable queries). Reranking 10 candidates with `mmarco-mminilm` takes ~0.4 s on CPU (4 threads; `bge-m3` ~6 s for 20). Its calibrated confidence tells "not in the KB" apart with 0.95 accuracy in cross-validation (calibration error 0.07):
| Retriever | Correct section in top 1 | in top 5 | in top 10 |
|---|---|---|---|
| semantic (โค 1.1.0) | 0.58 | 0.74 | 0.82 |
| **hybrid (1.2.0)** | 0.64 | 0.90 | **1.00** |
| **hybrid + rerank `mmarco-mminilm` (1.3.0, default)** | **0.80** | **0.94** | **1.00** |
| hybrid + rerank `bge-m3` (20 candidates) | 0.74 | 0.96 | 1.00 |
## Credits
Hook architecture inspired by [agd-memory](https://github.com/Pinperepette/agd-memory) by Pinperepette (MIT License) โ in particular the UserPromptSubmit recall pattern and guard rail logic.
This server cannot be deployed
Maintenance
ActivityMaintained
ResponsivenessNo issues