claude-rag
Indexes a folder of local Markdown (.md) files into a SQLite vector store, splitting them on ## headers with configurable chunk size/overlap, and exposes hybrid BM25 + semantic search with reranking so sections can be retrieved by exact identifiers or paraphrased questions.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@claude-ragfind notes about deploying on AWS"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
claude-rag โ MCP RAG Server for Markdown Knowledge Bases
Author: Sergio Angelastro โ MIT License
MCP server that indexes a folder of .md files into a local SQLite vector store and exposes hybrid search (BM25 keyword + semantic embeddings) as Claude Code tools. The core RAG system (chunking, embedding, SQLite vector store, MCP server, hybrid search, retrieval eval) is original work by the author.
Hook system (
hooks/) inspired by agd-memory by @Pinperepette (MIT) โ see CREDITS.md
Architecture

๐ Interactive diagram โ ยท ๐ How
kb_searchfinds an answer, step by step โ
Three components working together:
What | When | |
โ Indexing |
| on startup / |
โก MCP | Claude calls | explicit hybrid search |
โข Hook | every prompt intercepted โ keyword score on | automatic, ~10ms, no model |
Stack: Python ยท fastembed / ONNX Runtime (embedding paraphrase-multilingual-MiniLM-L12-v2 ~120MB + reranker mmarco-mMiniLMv2-L12 int8 ~120MB, CPU-only) ยท SQLite ยท MCP stdio
Related MCP server: flightlog
Tools exposed
Tool | Description |
| Hybrid search (BM25 + semantic, RRF) โ returns top-K chunks with source file and section. |
| Re-indexes files modified since last run (mtime-based) |
| Shows indexed files, chunk counts, last update timestamps |
| Shows cumulative token savings: RAG chunks served vs full-file baseline, broken down by source ( |
Setup
1. Clone
# Default layout: repo sits inside the KB folder
# KB files (.md) go in the parent directory
git clone https://github.com/sangelastro/claude-rag ~/.claude/my-kb/ragOr clone anywhere and point to your KB folder via env var (see step 3).
2. Install dependencies
cd ~/.claude/my-kb/rag
pip install -r requirements.txtOn first run the model (paraphrase-multilingual-MiniLM-L12-v2, ~120MB) is downloaded automatically from HuggingFace. This model supports 50+ languages including Italian natively.
3. Register in Claude Code
Add to ~/.claude.json under mcpServers:
"my-kb": {
"command": "python",
"args": ["/absolute/path/to/rag/server.py"],
"env": {
"KB_RAG_DIR": "/absolute/path/to/your/kb/folder"
}
}KB_RAG_DIRโ folder containing your.mdfiles (default:../relative toserver.py)KB_RAG_DBโ SQLite database path (default:kb.dbnext toserver.py)CHUNK_MAX_CHARSโ max chars per chunk before splitting (default:400)CHUNK_OVERLAPโ overlap in chars between consecutive sub-chunks (default:80)
If the repo is cloned inside the KB folder (as in the example above), both env vars can be omitted.
4. Restart Claude Code
The server starts automatically. On first launch it indexes all .md files in KB_RAG_DIR.
Optional: Claude Code Hooks
Two hooks auto-inject KB context without explicit kb_search calls.
Approach inspired by agd-memory (MIT).
Hook | Script | What it does |
|
| Injects KB table of contents at session start |
|
| Auto-injects top matching chunks before each prompt (keyword scoring, ~10ms, no model load) |
Setup hooks
Copy the relevant sections from hooks/hooks_example.json into your ~/.claude/settings.json, replacing the placeholder paths:
{
"hooks": {
"SessionStart": [{
"matcher": "*",
"hooks": [{"type": "command", "command": "python /absolute/path/to/rag/hooks/kb_session_start.py"}]
}],
"UserPromptSubmit": [{
"matcher": "*",
"hooks": [{"type": "command", "command": "python /absolute/path/to/rag/hooks/kb_recall.py", "timeout": 5}]
}]
}
}kb_chunks.json (used by kb_recall.py) is auto-generated next to kb.db on every reindex. No model loading in hooks โ scoring uses token overlap only.
Hook behaviour is tunable via env vars:
Variable | Default | Description |
|
| Max chunks injected per prompt |
|
| Minimum score to trigger injection |
|
| Skip prompts shorter than N words |
|
| Max chars injected (~4 chars/token) |
File structure
rag/
โโโ server.py # MCP server
โโโ reranker.py # Local cross-encoder reranker + calibrated confidence
โโโ requirements.txt
โโโ .gitignore
โโโ README.md
โโโ hooks/
โ โโโ kb_session_start.py # SessionStart hook
โ โโโ kb_recall.py # UserPromptSubmit hook
โ โโโ hooks_example.json # Hook config template
โโโ eval/
โ โโโ eval_retrieval.py # Retrieval benchmark against a gold set
โ โโโ gold_set.example.json # Gold set format (the real one stays local)
โโโ docs/how_search_works.html # Interactive walkthrough of the search pipeline
โโโ architecture.html # Technical documentation
โโโ kb_rag_slides.html # Architecture slide deckkb.db, kb_chunks.json, eval/gold_set.json and eval/results.json are generated locally and excluded from git: they contain KB content.
How it works
Chunking โ each
.mdfile is split on##headers; frontmatter is stripped; sections longer thanCHUNK_MAX_CHARS(400) are further split into overlapping sub-chunks withCHUNK_OVERLAP(80) chars of context continuityEmbedding โ chunks are encoded with
paraphrase-multilingual-MiniLM-L12-v2(384 dimensions, 50+ languages)Storage โ vectors stored as
float32BLOBs in SQLite +kb_chunks.jsonfor hooksSearch โ hybrid: BM25 over an in-memory inverted index + cosine similarity in numpy, the two rankings (top 50 each) fused with Reciprocal Rank Fusion; top-K returned. BM25 catches exact identifiers (table names, ports, hostnames) that a 128-token embedding model blurs; embeddings catch paraphrased natural-language questions
Rerank + confidence โ the top 10 hybrid candidates are rescored by a local multilingual cross-encoder (ONNX, int8, see
reranker.py): one forward pass per (query, chunk) pair, no text generated. The top score is mapped to a calibrated probability (Platt scaling) that the answer is among the results; belowKB_MIN_CONFIDENCEthe tool says so instead of silently returning noiseHooks โ keyword scoring on
kb_chunks.json(no model), injected before each promptInvalidation โ mtime-based: only modified files are re-indexed on startup
Savings tracking โ every search records chars served vs full-file baseline in
search_statstable;kb_savings()aggregates the cumulative token reduction without re-reading any file
Environment variables
Variable | Default | Description |
|
| Folder with |
|
| SQLite database path |
|
| MCP server name |
|
| Max chars per chunk; longer sections are split into overlapping sub-chunks |
|
| Overlap chars between adjacent sub-chunks to preserve context continuity |
|
|
|
|
| Local cross-encoder that reorders the hybrid candidates: |
|
| Hybrid candidates passed to the reranker |
|
| ONNX Runtime threads for the reranker |
|
| Where reranker models are downloaded |
|
| Below this calibrated confidence, |
Evaluating retrieval
eval/eval_retrieval.py measures how often the right section comes back, on a gold set of real queries labelled with their correct file + section. It compares semantic, hook lexical, BM25, hybrid, the actual server.rank() and (optionally) a local cross-encoder reranker. It only reads kb.db and never writes to search_stats.
cp eval/gold_set.example.json eval/gold_set.json # then write queries about your KB
py eval/eval_retrieval.py --no-rerankInclude unanswerable queries ("kind": "neg"): the script also reports whether a score threshold can tell "not in the KB" apart from a real hit. --calibrate fits the Platt parameters of each reranker with 5-fold cross-validation; copy platt_all into reranker.py.
Measured on the author's KB (5,241 chunks, 50 answerable + 7 unanswerable queries). Reranking 10 candidates with mmarco-mminilm takes ~0.4 s on CPU (4 threads; bge-m3 ~6 s for 20). Its calibrated confidence tells "not in the KB" apart with 0.95 accuracy in cross-validation (calibration error 0.07):
Retriever | Correct section in top 1 | in top 5 | in top 10 |
semantic (โค 1.1.0) | 0.58 | 0.74 | 0.82 |
hybrid (1.2.0) | 0.64 | 0.90 | 1.00 |
hybrid + rerank | 0.80 | 0.94 | 1.00 |
hybrid + rerank | 0.74 | 0.96 | 1.00 |
Credits
Hook architecture inspired by agd-memory by Pinperepette (MIT License) โ in particular the UserPromptSubmit recall pattern and guard rail logic.
This server cannot be deployed
Maintenance
Related MCP Connectors
Augments MCP Server - A comprehensive framework documentation provider for Claude Code
Serve a folder of Markdown notes as an MCP server: hybrid search, reading, and sourced answers.
An MCP server that gives your AI access to the source code and docs of all public github repos
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceA local MCP server that indexes and searches your Claude Code conversation history with both keyword and semantic search, fully private and running locally.MIT
- AlicenseAqualityCmaintenanceAn MCP server that indexes Claude Code conversation history into SQLite, enabling full-text search across past sessions for context recovery and cross-agent observability.103MIT
- AlicenseAqualityCmaintenanceA local MCP server that indexes and searches your past Claude sessions using SQLite FTS5. No cloud, runs entirely on your machine.3MIT
- AlicenseAqualityDmaintenanceMCP server that indexes Claude.ai chats and local Claude Code sessions, enabling semantic and keyword search across all your conversations with Claude.65MIT