Skip to main content
Glama
README.md
# claude-rag โ€” MCP RAG Server for Markdown Knowledge Bases

**Author**: [Sergio Angelastro](https://github.com/sangelastro) โ€” MIT License

MCP server that indexes a folder of `.md` files into a local SQLite vector store and exposes hybrid search (BM25 keyword + semantic embeddings) as Claude Code tools. The core RAG system (chunking, embedding, SQLite vector store, MCP server, hybrid search, retrieval eval) is original work by the author.

> **Hook system** (`hooks/`) inspired by [agd-memory](https://github.com/Pinperepette/agd-memory) by [@Pinperepette](https://github.com/Pinperepette) (MIT) โ€” see [CREDITS.md](CREDITS.md)

## Architecture

![Architecture Diagram](docs/diagram.png)

> ๐Ÿ“Š **[Interactive diagram โ†’](docs/diagram.html)** ยท ๐Ÿ”Ž **[How `kb_search` finds an answer, step by step โ†’](docs/how_search_works.html)**

Three components working together:

| | What | When |
|---|---|---|
| **โ‘  Indexing** | `.md` files โ†’ Chunker โ†’ Embedder โ†’ SQLite + JSON | on startup / `kb_reindex()` |
| **โ‘ก MCP** | Claude calls `kb_search()` โ†’ BM25 + cosine sim fused with RRF โ†’ top-K chunks | explicit hybrid search |
| **โ‘ข Hook** | every prompt intercepted โ†’ keyword score on `kb_chunks.json` โ†’ auto-inject | automatic, ~10ms, no model |

**Stack**: Python ยท fastembed / ONNX Runtime (embedding `paraphrase-multilingual-MiniLM-L12-v2` ~120MB + reranker `mmarco-mMiniLMv2-L12` int8 ~120MB, CPU-only) ยท SQLite ยท MCP stdio

## Tools exposed

| Tool | Description |
|---|---|
| `kb_search(query, top_k=5)` | Hybrid search (BM25 + semantic, RRF) โ€” returns top-K chunks with source file and section. `KB_SEARCH_MODE=semantic` for cosine only |
| `kb_reindex(force=False)` | Re-indexes files modified since last run (mtime-based) |
| `kb_stats()` | Shows indexed files, chunk counts, last update timestamps |
| `kb_savings()` | Shows cumulative token savings: RAG chunks served vs full-file baseline, broken down by source (`mcp` / `hook`) |

## Setup

### 1. Clone

```bash
# Default layout: repo sits inside the KB folder
# KB files (.md) go in the parent directory
git clone https://github.com/sangelastro/claude-rag ~/.claude/my-kb/rag
```

Or clone anywhere and point to your KB folder via env var (see step 3).

### 2. Install dependencies

```bash
cd ~/.claude/my-kb/rag
pip install -r requirements.txt
```

On first run the model (`paraphrase-multilingual-MiniLM-L12-v2`, ~120MB) is downloaded automatically from HuggingFace. This model supports 50+ languages including Italian natively.

### 3. Register in Claude Code

Add to `~/.claude.json` under `mcpServers`:

```json
"my-kb": {
  "command": "python",
  "args": ["/absolute/path/to/rag/server.py"],
  "env": {
    "KB_RAG_DIR": "/absolute/path/to/your/kb/folder"
  }
}
```

- **`KB_RAG_DIR`** โ€” folder containing your `.md` files (default: `../` relative to `server.py`)
- **`KB_RAG_DB`** โ€” SQLite database path (default: `kb.db` next to `server.py`)
- **`CHUNK_MAX_CHARS`** โ€” max chars per chunk before splitting (default: `400`)
- **`CHUNK_OVERLAP`** โ€” overlap in chars between consecutive sub-chunks (default: `80`)

If the repo is cloned inside the KB folder (as in the example above), both env vars can be omitted.

### 4. Restart Claude Code

The server starts automatically. On first launch it indexes all `.md` files in `KB_RAG_DIR`.

## Optional: Claude Code Hooks

Two hooks auto-inject KB context without explicit `kb_search` calls.
Approach inspired by [agd-memory](https://github.com/Pinperepette/agd-memory) (MIT).

| Hook | Script | What it does |
|---|---|---|
| `SessionStart` | `hooks/kb_session_start.py` | Injects KB table of contents at session start |
| `UserPromptSubmit` | `hooks/kb_recall.py` | Auto-injects top matching chunks before each prompt (keyword scoring, ~10ms, no model load) |

### Setup hooks

Copy the relevant sections from `hooks/hooks_example.json` into your `~/.claude/settings.json`, replacing the placeholder paths:

```json
{
  "hooks": {
    "SessionStart": [{
      "matcher": "*",
      "hooks": [{"type": "command", "command": "python /absolute/path/to/rag/hooks/kb_session_start.py"}]
    }],
    "UserPromptSubmit": [{
      "matcher": "*",
      "hooks": [{"type": "command", "command": "python /absolute/path/to/rag/hooks/kb_recall.py", "timeout": 5}]
    }]
  }
}
```

`kb_chunks.json` (used by `kb_recall.py`) is auto-generated next to `kb.db` on every reindex. No model loading in hooks โ€” scoring uses token overlap only.

Hook behaviour is tunable via env vars:

| Variable | Default | Description |
|---|---|---|
| `KB_RAG_HOOK_TOP_K` | `3` | Max chunks injected per prompt |
| `KB_RAG_HOOK_MIN_SCORE` | `0.15` | Minimum score to trigger injection |
| `KB_RAG_HOOK_MIN_WORDS` | `4` | Skip prompts shorter than N words |
| `KB_RAG_HOOK_TOKEN_BUDGET` | `6000` | Max chars injected (~4 chars/token) |

## File structure

```
rag/
โ”œโ”€โ”€ server.py           # MCP server
โ”œโ”€โ”€ reranker.py         # Local cross-encoder reranker + calibrated confidence
โ”œโ”€โ”€ requirements.txt
โ”œโ”€โ”€ .gitignore
โ”œโ”€โ”€ README.md
โ”œโ”€โ”€ hooks/
โ”‚   โ”œโ”€โ”€ kb_session_start.py   # SessionStart hook
โ”‚   โ”œโ”€โ”€ kb_recall.py          # UserPromptSubmit hook
โ”‚   โ””โ”€โ”€ hooks_example.json    # Hook config template
โ”œโ”€โ”€ eval/
โ”‚   โ”œโ”€โ”€ eval_retrieval.py       # Retrieval benchmark against a gold set
โ”‚   โ””โ”€โ”€ gold_set.example.json   # Gold set format (the real one stays local)
โ”œโ”€โ”€ docs/how_search_works.html  # Interactive walkthrough of the search pipeline
โ”œโ”€โ”€ architecture.html   # Technical documentation
โ””โ”€โ”€ kb_rag_slides.html  # Architecture slide deck
```

`kb.db`, `kb_chunks.json`, `eval/gold_set.json` and `eval/results.json` are generated locally and excluded from git: they contain KB content.

## How it works

1. **Chunking** โ€” each `.md` file is split on `##` headers; frontmatter is stripped; sections longer than `CHUNK_MAX_CHARS` (400) are further split into overlapping sub-chunks with `CHUNK_OVERLAP` (80) chars of context continuity
2. **Embedding** โ€” chunks are encoded with `paraphrase-multilingual-MiniLM-L12-v2` (384 dimensions, 50+ languages)
3. **Storage** โ€” vectors stored as `float32` BLOBs in SQLite + `kb_chunks.json` for hooks
4. **Search** โ€” hybrid: BM25 over an in-memory inverted index + cosine similarity in numpy, the two rankings (top 50 each) fused with Reciprocal Rank Fusion; top-K returned. BM25 catches exact identifiers (table names, ports, hostnames) that a 128-token embedding model blurs; embeddings catch paraphrased natural-language questions
5. **Rerank + confidence** โ€” the top 10 hybrid candidates are rescored by a local multilingual cross-encoder (ONNX, int8, see `reranker.py`): one forward pass per (query, chunk) pair, no text generated. The top score is mapped to a calibrated probability (Platt scaling) that the answer is among the results; below `KB_MIN_CONFIDENCE` the tool says so instead of silently returning noise
6. **Hooks** โ€” keyword scoring on `kb_chunks.json` (no model), injected before each prompt
7. **Invalidation** โ€” mtime-based: only modified files are re-indexed on startup
8. **Savings tracking** โ€” every search records chars served vs full-file baseline in `search_stats` table; `kb_savings()` aggregates the cumulative token reduction without re-reading any file

## Environment variables

| Variable | Default | Description |
|---|---|---|
| `KB_RAG_DIR` | `../` (relative to `server.py`) | Folder with `.md` files to index |
| `KB_RAG_DB` | `./kb.db` (next to `server.py`) | SQLite database path |
| `KB_RAG_NAME` | `kb-rag` | MCP server name |
| `CHUNK_MAX_CHARS` | `400` | Max chars per chunk; longer sections are split into overlapping sub-chunks |
| `CHUNK_OVERLAP` | `80` | Overlap chars between adjacent sub-chunks to preserve context continuity |
| `KB_SEARCH_MODE` | `hybrid` | `hybrid` (BM25 + semantic, RRF) or `semantic` (cosine only, behaviour up to 1.1.0) |
| `KB_RERANK` | `mmarco-mminilm` | Local cross-encoder that reorders the hybrid candidates: `mmarco-mminilm` (AVX-512 int8), `mmarco-mminilm-avx2` (older CPUs), `bge-m3` (slower, not calibrated) or `off`. If it cannot load, search falls back to hybrid |
| `KB_RERANK_CANDIDATES` | `10` | Hybrid candidates passed to the reranker |
| `KB_RERANK_THREADS` | `4` | ONNX Runtime threads for the reranker |
| `KB_RERANK_CACHE` | `~/.cache/claude-rag/models` | Where reranker models are downloaded |
| `KB_MIN_CONFIDENCE` | `0.5` | Below this calibrated confidence, `kb_search` warns that the answer is probably not in the KB |

## Evaluating retrieval

`eval/eval_retrieval.py` measures how often the right section comes back, on a gold set of real queries labelled with their correct `file` + `section`. It compares semantic, hook lexical, BM25, hybrid, the actual `server.rank()` and (optionally) a local cross-encoder reranker. It only reads `kb.db` and never writes to `search_stats`.

```bash
cp eval/gold_set.example.json eval/gold_set.json   # then write queries about your KB
py eval/eval_retrieval.py --no-rerank
```

Include unanswerable queries (`"kind": "neg"`): the script also reports whether a score threshold can tell "not in the KB" apart from a real hit. `--calibrate` fits the Platt parameters of each reranker with 5-fold cross-validation; copy `platt_all` into `reranker.py`.

Measured on the author's KB (5,241 chunks, 50 answerable + 7 unanswerable queries). Reranking 10 candidates with `mmarco-mminilm` takes ~0.4 s on CPU (4 threads; `bge-m3` ~6 s for 20). Its calibrated confidence tells "not in the KB" apart with 0.95 accuracy in cross-validation (calibration error 0.07):

| Retriever | Correct section in top 1 | in top 5 | in top 10 |
|---|---|---|---|
| semantic (โ‰ค 1.1.0) | 0.58 | 0.74 | 0.82 |
| **hybrid (1.2.0)** | 0.64 | 0.90 | **1.00** |
| **hybrid + rerank `mmarco-mminilm` (1.3.0, default)** | **0.80** | **0.94** | **1.00** |
| hybrid + rerank `bge-m3` (20 candidates) | 0.74 | 0.96 | 1.00 |

## Credits

Hook architecture inspired by [agd-memory](https://github.com/Pinperepette/agd-memory) by Pinperepette (MIT License) โ€” in particular the UserPromptSubmit recall pattern and guard rail logic.