corpus-search-mcp
README.md
# corpus-search-mcp
Point it at a folder of markdown and get a local, private, hybrid-search MCP
server over your own notes. No cloud API, no per-query cost: chunk your markdown,
embed it locally with [Ollama](https://ollama.com), search it with BM25 +
vector fusion, and wire it into [Claude Code](https://docs.claude.com/en/docs/claude-code)
(or any MCP client) as three tools: `search`, `list_sections`, `get_section`.
The hard parts are built and corpus-agnostic: hybrid retrieval, Reciprocal Rank
Fusion, an embedding-model-down fallback, fuzzy section lookup, incremental
re-indexing, and a runnable quality gate. The only things you tune are a small
`config.py` and, if your corpus is unusual, a chunker.
A tiny example corpus (coffee-brewing notes) is bundled so the eval gate runs
out of the box.
## How it works
```
your markdown -> chunker -> Ollama embeddings + BM25 -> LanceDB -> MCP tools
```
| File | Role |
|---|---|
| `search_core.py` | The engine: query expansion, BM25 + vector search, RRF, formatting, and a pure-stdlib keyword index used as the dependency-free fallback |
| `server.py` | Generic FastMCP server exposing `search` / `list_sections` / `get_section` |
| `indexer.py` | Embeds chunks into LanceDB with a BM25 index; incremental, hash-tracked re-index |
| `config.py` | Everything corpus-specific, driven by environment variables |
| `chunkers/` | How to split your corpus. `h2` (default) and `dispatch` (advanced) built in |
| `eval.py` | Self-contained retrieval quality gate (BM25 keyword mode, no services) |
| `examples/corpus/` | A small bundled markdown corpus so the eval runs meaningfully |
## Install
```bash
python -m venv .venv
. .venv/bin/activate # Windows: .venv\Scripts\activate
pip install -r requirements.txt
```
You also need Ollama running with an embedding model, for indexing and for
vector search at query time:
```bash
ollama pull mxbai-embed-large
```
Search still works without Ollama (BM25 / keyword only); the vector half simply
switches off.
## Configure
Everything is set by environment variables with sane defaults (see `config.py`).
The one you will always set is the corpus folder:
```bash
export CORPUS_DIR=/path/to/your/markdown # default: examples/corpus
export CHUNKER=h2 # or "dispatch"; default h2
export OLLAMA_URL=http://localhost:11434/api/embed
export EMBED_MODEL=mxbai-embed-large
```
Then open `config.py` and edit `SYNONYMS` to your domain vocabulary (optional;
an empty map disables query expansion).
## Index
```bash
python indexer.py # full build
python indexer.py --changed # re-index only files whose content changed
python indexer.py --file note.md # re-index a single file (h2 chunker)
```
The index and a `_index_meta.json` of per-file content hashes are written under
`CORPUS_DIR/.vectordb` by default.
## Serve
```bash
python server.py # runs the MCP server over stdio
```
## Wire into Claude Code
Register the server, pointing it at your corpus:
```bash
claude mcp add corpus-search \
-e CORPUS_DIR=/path/to/your/markdown \
-e CHUNKER=h2 \
-- python /path/to/corpus-search-mcp/server.py
```
Use the venv's Python if you installed into one. After it registers, the
`search`, `list_sections`, and `get_section` tools are available in your session.
## Eval
```bash
python eval.py # exits 0 when Recall@5 / P@1 / MRR meet the floors
python eval.py -v # also print failing cases
```
The gate is self-contained: it builds an in-memory BM25 index from the bundled
corpus and needs no LanceDB, no embedding model, and no network. Rewrite
`TEST_CASES` (and point `CORPUS_DIR` at your corpus) to make it a real regression
gate for your own material.
## Chunkers
How to split your corpus is the one bespoke piece. Two are built in:
- **`h2` (default)** -- split every markdown file at its `##` headers. Best for a
flat folder of topic notes.
- **`dispatch` (advanced)** -- classify each file by the subfolder it lives in
(`notes/`, `articles/`, `references/`) and apply a different strategy per type
(whole-file, H2, H3). The template for a corpus that mixes document kinds.
To add your own, drop a module in `chunkers/` exposing
`chunk_all(corpus_dir) -> list[dict]` (each dict needs `id`, `file`, `heading`,
`heading_path`, `line_start`, `line_end`, `text`, `hash`) and register it in
`chunkers/__init__.py`.
## License
MIT. See [LICENSE](LICENSE).
This server cannot be deployed
Maintenance
ActivityStale
ResponsivenessNo issues