Skip to main content
Glama

claude-rag โ€” MCP RAG Server for Markdown Knowledge Bases

Author: Sergio Angelastro โ€” MIT License

MCP server that indexes a folder of .md files into a local SQLite vector store and exposes hybrid search (BM25 keyword + semantic embeddings) as Claude Code tools. The core RAG system (chunking, embedding, SQLite vector store, MCP server, hybrid search, retrieval eval) is original work by the author.

Hook system (hooks/) inspired by agd-memory by @Pinperepette (MIT) โ€” see CREDITS.md

Architecture

Architecture Diagram

๐Ÿ“Š Interactive diagram โ†’ ยท ๐Ÿ”Ž How kb_search finds an answer, step by step โ†’

Three components working together:

What

When

โ‘  Indexing

.md files โ†’ Chunker โ†’ Embedder โ†’ SQLite + JSON

on startup / kb_reindex()

โ‘ก MCP

Claude calls kb_search() โ†’ BM25 + cosine sim fused with RRF โ†’ top-K chunks

explicit hybrid search

โ‘ข Hook

every prompt intercepted โ†’ keyword score on kb_chunks.json โ†’ auto-inject

automatic, ~10ms, no model

Stack: Python ยท fastembed / ONNX Runtime (embedding paraphrase-multilingual-MiniLM-L12-v2 ~120MB + reranker mmarco-mMiniLMv2-L12 int8 ~120MB, CPU-only) ยท SQLite ยท MCP stdio

Related MCP server: flightlog

Tools exposed

Tool

Description

kb_search(query, top_k=5)

Hybrid search (BM25 + semantic, RRF) โ€” returns top-K chunks with source file and section. KB_SEARCH_MODE=semantic for cosine only

kb_reindex(force=False)

Re-indexes files modified since last run (mtime-based)

kb_stats()

Shows indexed files, chunk counts, last update timestamps

kb_savings()

Shows cumulative token savings: RAG chunks served vs full-file baseline, broken down by source (mcp / hook)

Setup

1. Clone

# Default layout: repo sits inside the KB folder
# KB files (.md) go in the parent directory
git clone https://github.com/sangelastro/claude-rag ~/.claude/my-kb/rag

Or clone anywhere and point to your KB folder via env var (see step 3).

2. Install dependencies

cd ~/.claude/my-kb/rag
pip install -r requirements.txt

On first run the model (paraphrase-multilingual-MiniLM-L12-v2, ~120MB) is downloaded automatically from HuggingFace. This model supports 50+ languages including Italian natively.

3. Register in Claude Code

Add to ~/.claude.json under mcpServers:

"my-kb": {
  "command": "python",
  "args": ["/absolute/path/to/rag/server.py"],
  "env": {
    "KB_RAG_DIR": "/absolute/path/to/your/kb/folder"
  }
}
  • KB_RAG_DIR โ€” folder containing your .md files (default: ../ relative to server.py)

  • KB_RAG_DB โ€” SQLite database path (default: kb.db next to server.py)

  • CHUNK_MAX_CHARS โ€” max chars per chunk before splitting (default: 400)

  • CHUNK_OVERLAP โ€” overlap in chars between consecutive sub-chunks (default: 80)

If the repo is cloned inside the KB folder (as in the example above), both env vars can be omitted.

4. Restart Claude Code

The server starts automatically. On first launch it indexes all .md files in KB_RAG_DIR.

Optional: Claude Code Hooks

Two hooks auto-inject KB context without explicit kb_search calls. Approach inspired by agd-memory (MIT).

Hook

Script

What it does

SessionStart

hooks/kb_session_start.py

Injects KB table of contents at session start

UserPromptSubmit

hooks/kb_recall.py

Auto-injects top matching chunks before each prompt (keyword scoring, ~10ms, no model load)

Setup hooks

Copy the relevant sections from hooks/hooks_example.json into your ~/.claude/settings.json, replacing the placeholder paths:

{
  "hooks": {
    "SessionStart": [{
      "matcher": "*",
      "hooks": [{"type": "command", "command": "python /absolute/path/to/rag/hooks/kb_session_start.py"}]
    }],
    "UserPromptSubmit": [{
      "matcher": "*",
      "hooks": [{"type": "command", "command": "python /absolute/path/to/rag/hooks/kb_recall.py", "timeout": 5}]
    }]
  }
}

kb_chunks.json (used by kb_recall.py) is auto-generated next to kb.db on every reindex. No model loading in hooks โ€” scoring uses token overlap only.

Hook behaviour is tunable via env vars:

Variable

Default

Description

KB_RAG_HOOK_TOP_K

3

Max chunks injected per prompt

KB_RAG_HOOK_MIN_SCORE

0.15

Minimum score to trigger injection

KB_RAG_HOOK_MIN_WORDS

4

Skip prompts shorter than N words

KB_RAG_HOOK_TOKEN_BUDGET

6000

Max chars injected (~4 chars/token)

File structure

rag/
โ”œโ”€โ”€ server.py           # MCP server
โ”œโ”€โ”€ reranker.py         # Local cross-encoder reranker + calibrated confidence
โ”œโ”€โ”€ requirements.txt
โ”œโ”€โ”€ .gitignore
โ”œโ”€โ”€ README.md
โ”œโ”€โ”€ hooks/
โ”‚   โ”œโ”€โ”€ kb_session_start.py   # SessionStart hook
โ”‚   โ”œโ”€โ”€ kb_recall.py          # UserPromptSubmit hook
โ”‚   โ””โ”€โ”€ hooks_example.json    # Hook config template
โ”œโ”€โ”€ eval/
โ”‚   โ”œโ”€โ”€ eval_retrieval.py       # Retrieval benchmark against a gold set
โ”‚   โ””โ”€โ”€ gold_set.example.json   # Gold set format (the real one stays local)
โ”œโ”€โ”€ docs/how_search_works.html  # Interactive walkthrough of the search pipeline
โ”œโ”€โ”€ architecture.html   # Technical documentation
โ””โ”€โ”€ kb_rag_slides.html  # Architecture slide deck

kb.db, kb_chunks.json, eval/gold_set.json and eval/results.json are generated locally and excluded from git: they contain KB content.

How it works

  1. Chunking โ€” each .md file is split on ## headers; frontmatter is stripped; sections longer than CHUNK_MAX_CHARS (400) are further split into overlapping sub-chunks with CHUNK_OVERLAP (80) chars of context continuity

  2. Embedding โ€” chunks are encoded with paraphrase-multilingual-MiniLM-L12-v2 (384 dimensions, 50+ languages)

  3. Storage โ€” vectors stored as float32 BLOBs in SQLite + kb_chunks.json for hooks

  4. Search โ€” hybrid: BM25 over an in-memory inverted index + cosine similarity in numpy, the two rankings (top 50 each) fused with Reciprocal Rank Fusion; top-K returned. BM25 catches exact identifiers (table names, ports, hostnames) that a 128-token embedding model blurs; embeddings catch paraphrased natural-language questions

  5. Rerank + confidence โ€” the top 10 hybrid candidates are rescored by a local multilingual cross-encoder (ONNX, int8, see reranker.py): one forward pass per (query, chunk) pair, no text generated. The top score is mapped to a calibrated probability (Platt scaling) that the answer is among the results; below KB_MIN_CONFIDENCE the tool says so instead of silently returning noise

  6. Hooks โ€” keyword scoring on kb_chunks.json (no model), injected before each prompt

  7. Invalidation โ€” mtime-based: only modified files are re-indexed on startup

  8. Savings tracking โ€” every search records chars served vs full-file baseline in search_stats table; kb_savings() aggregates the cumulative token reduction without re-reading any file

Environment variables

Variable

Default

Description

KB_RAG_DIR

../ (relative to server.py)

Folder with .md files to index

KB_RAG_DB

./kb.db (next to server.py)

SQLite database path

KB_RAG_NAME

kb-rag

MCP server name

CHUNK_MAX_CHARS

400

Max chars per chunk; longer sections are split into overlapping sub-chunks

CHUNK_OVERLAP

80

Overlap chars between adjacent sub-chunks to preserve context continuity

KB_SEARCH_MODE

hybrid

hybrid (BM25 + semantic, RRF) or semantic (cosine only, behaviour up to 1.1.0)

KB_RERANK

mmarco-mminilm

Local cross-encoder that reorders the hybrid candidates: mmarco-mminilm (AVX-512 int8), mmarco-mminilm-avx2 (older CPUs), bge-m3 (slower, not calibrated) or off. If it cannot load, search falls back to hybrid

KB_RERANK_CANDIDATES

10

Hybrid candidates passed to the reranker

KB_RERANK_THREADS

4

ONNX Runtime threads for the reranker

KB_RERANK_CACHE

~/.cache/claude-rag/models

Where reranker models are downloaded

KB_MIN_CONFIDENCE

0.5

Below this calibrated confidence, kb_search warns that the answer is probably not in the KB

Evaluating retrieval

eval/eval_retrieval.py measures how often the right section comes back, on a gold set of real queries labelled with their correct file + section. It compares semantic, hook lexical, BM25, hybrid, the actual server.rank() and (optionally) a local cross-encoder reranker. It only reads kb.db and never writes to search_stats.

cp eval/gold_set.example.json eval/gold_set.json   # then write queries about your KB
py eval/eval_retrieval.py --no-rerank

Include unanswerable queries ("kind": "neg"): the script also reports whether a score threshold can tell "not in the KB" apart from a real hit. --calibrate fits the Platt parameters of each reranker with 5-fold cross-validation; copy platt_all into reranker.py.

Measured on the author's KB (5,241 chunks, 50 answerable + 7 unanswerable queries). Reranking 10 candidates with mmarco-mminilm takes ~0.4 s on CPU (4 threads; bge-m3 ~6 s for 20). Its calibrated confidence tells "not in the KB" apart with 0.95 accuracy in cross-validation (calibration error 0.07):

Retriever

Correct section in top 1

in top 5

in top 10

semantic (โ‰ค 1.1.0)

0.58

0.74

0.82

hybrid (1.2.0)

0.64

0.90

1.00

hybrid + rerank mmarco-mminilm (1.3.0, default)

0.80

0.94

1.00

hybrid + rerank bge-m3 (20 candidates)

0.74

0.96

1.00

Credits

Hook architecture inspired by agd-memory by Pinperepette (MIT License) โ€” in particular the UserPromptSubmit recall pattern and guard rail logic.

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    C
    maintenance
    An MCP server that indexes Claude Code conversation history into SQLite, enabling full-text search across past sessions for context recovery and cross-agent observability.
    10
    3
    MIT
  • A
    license
    A
    quality
    D
    maintenance
    MCP server that indexes Claude.ai chats and local Claude Code sessions, enabling semantic and keyword search across all your conversations with Claude.
    6
    5
    MIT