obsidian-rag-mcp
by AdrianBodrug
README.md
# obsidian-rag-mcp
[](https://github.com/AdrianBodrug/obsidian-rag-mcp/actions/workflows/ci.yml)
Local hybrid search over an [Obsidian](https://obsidian.md) vault, exposed to [Claude Code](https://claude.com/claude-code) through the [Model Context Protocol](https://modelcontextprotocol.io).
Ask Claude *"what did we decide about shipping the battery?"* and it searches your notes, reads the relevant ones, and answers with citations like `Projects/Solar Car Logistics.md#Decisions`, without the notes leaving your machine.
- **Hybrid retrieval**: semantic vector search and BM25 keyword search, merged with reciprocal rank fusion.
- **On-device by default**: embeddings run locally (`bge-small-en-v1.5` via transformers.js). OpenAI is optional.
- **Incremental**: only changed notes are re-read, and only edited sections are re-embedded.
- **Obsidian-aware**: heading-aware chunks, YAML frontmatter, tags, aliases, and `[[wiki links]]` that can pull linked notes into results.
- **Measured**: every ranking change was benchmarked on a reproducible sample vault and on a real 171-note vault.
- **Sandboxed**: Claude can only read Markdown inside the vault and never excluded or hidden folders.
## How it works
```mermaid
flowchart LR
subgraph Index["Indexing (sync)"]
A[Vault .md files] --> B{size / mtime<br/>changed?}
B -- no --> S[skip, no read]
B -- yes --> C{content hash<br/>changed?}
C -- no --> S
C -- yes --> D[Parse frontmatter,<br/>tags, links, aliases]
D --> E[Split by heading,<br/>fence-aware, with overlap]
E --> F{chunk hash<br/>seen before?}
F -- yes --> G[reuse stored vector]
F -- no --> H[embed]
G & H --> I[(LanceDB<br/>vectors + BM25 index)]
end
```
```mermaid
flowchart LR
Q[Query] --> V[Vector search<br/>cosine]
Q --> K[BM25 search<br/>title, aliases, heading, text]
V --> R[Reciprocal rank fusion<br/>max 2 chunks per note]
K --> R
R --> L{expandLinks?}
L -- yes --> W[Append best chunk of<br/>wiki-linked notes]
L -- no --> O[Results with citations]
W --> O
```
**Chunking.** Notes are split at headings, and each chunk keeps its heading path (`Projects > Freight > Sea`). Headings inside fenced code blocks are ignored, and a heading with no text of its own is merged into its first subsection. Sections longer than `RAG_CHUNK_SIZE` are cut at paragraph or sentence boundaries, with overlap between neighbouring pieces.
**What gets indexed.** Each chunk's search text is its note title, aliases, heading path and body. The same string is both embedded and keyword-indexed, so a note is found by its title or alias even when the body never repeats it.
**Ranking.** Vector and keyword results are merged with [reciprocal rank fusion](https://plg.uwaterloo.ca/~gvcormac/cormacksigir09-rrf.pdf): each result scores `weight / (60 + rank)` in each list it appears in. Keyword hits get weight 1.15, a value chosen by measurement (see [Design decisions](#design-decisions)). At most two chunks per note are returned, so one long note cannot fill every slot.
**Incremental sync.** A manifest records each file's size, mtime and content hash. An idle sync only calls `stat`. When a note changes, chunks whose search text is unchanged keep their stored vector, so editing one paragraph re-embeds one chunk. The index rebuilds itself when the embedding model, chunk settings, schema or vault path changes.
## Quick start
Requires Node.js 22.9 or newer.
```bash
git clone https://github.com/AdrianBodrug/obsidian-rag-mcp.git
cd obsidian-rag-mcp
npm install
cp .env.example .env # then set OBSIDIAN_VAULT_PATH
npm run rag:sync # first run downloads the ~33 MB embedding model
npm run rag:search -- "how do we ship the battery" --expand-links
```
### Connect Claude Code
Claude Code picks up the bundled [`.mcp.json`](.mcp.json) when started in this directory. To use the vault from any project:
```bash
claude mcp add obsidian-rag -- node --env-file-if-exists=/path/to/obsidian-rag-mcp/.env /path/to/obsidian-rag-mcp/rag/server.js
```
The server speaks plain MCP over stdio, so any MCP client can run it with the same command.
### MCP tools
| Tool | Purpose |
|---|---|
| `search_vault` | Hybrid search. Options: `limit`, `folders`, `tags`, `expandLinks`. |
| `read_note` | Read the full note behind a search result. |
| `sync_vault` | Incremental sync, or `force` to re-embed everything. |
| `vault_status` | Index size, model and last sync time. |
The server's instructions tell Claude to cite sources and to treat note text as data, never as instructions.
## Benchmarks
`npm run bench` indexes a vault into a throwaway database and reports **hit@1** (the right note ranks first), **hit@6**, **recall@6** (share of all relevant notes retrieved) and **MRR** (mean reciprocal rank).
The *baseline* is the original retrieval logic before the fixes and improvements in this repository's history, measured with the same evaluator.
**Sample vault**: [`examples/vault`](examples/vault), 16 notes, 40 queries in [`examples/evals.jsonl`](examples/evals.jsonl). Reproducible with `npm run bench`.
| Embeddings | Version | hit@1 | hit@6 | recall@6 | MRR |
|---|---|---:|---:|---:|---:|
| local (bge-small) | baseline | 0.925 | 1.000 | 1.000 | 0.956 |
| local (bge-small) | **current** | **0.975** | 1.000 | 1.000 | **0.988** |
| local-hash (keyword only) | baseline | 0.775 | 0.950 | 0.938 | 0.850 |
| local-hash (keyword only) | **current** | **0.850** | **0.975** | **0.975** | **0.900** |
**Real vault**: a personal vault of 171 notes (~920 KB of course notes, lecture summaries and projects), 24 paraphrased queries. The notes and queries are private, so only aggregates are shown.
| Embeddings | Version | hit@1 | hit@6 | recall@6 | MRR |
|---|---|---:|---:|---:|---:|
| local (bge-small) | baseline | 0.958 | 1.000 | 0.704 | 0.979 |
| local (bge-small) | current | 0.958 | 1.000 | **0.754** | 0.979 |
| local (bge-small) | current + `expandLinks` | 0.958 | 1.000 | **0.972** | 0.979 |
| local-hash (keyword only) | baseline | 0.583 | 0.917 | 0.586 | 0.722 |
| local-hash (keyword only) | current | **0.667** | **0.958** | **0.678** | **0.757** |
| local-hash (keyword only) | current + `expandLinks` | 0.667 | 1.000 | 0.885 | 0.762 |
What moved the numbers:
- **Indexing titles and aliases** fixed alias queries such as "branching strategy" and "ML course": sample-vault hit@1 went from 0.925 to 0.975 with local embeddings.
- **Capping chunks per note** raised real-vault recall from 0.704 to 0.754. Several sections of one long note had been crowding out other relevant notes.
- **Link expansion** raised real-vault recall to 0.972. It returns up to three extra notes on top of the limit, so this is not a like-for-like comparison. It is off by default.
## Configuration
Set in `.env` or the environment. Invalid values fail at startup rather than being silently replaced.
| Variable | Default | Purpose |
|---|---|---|
| `OBSIDIAN_VAULT_PATH` | required | Vault root |
| `RAG_EMBEDDING_PROVIDER` | `local` | `local`, `openai`, or `local-hash` (offline, keyword hashing) |
| `RAG_EMBEDDING_MODEL` | per provider | `Xenova/bge-small-en-v1.5` / `text-embedding-3-small` |
| `RAG_EMBEDDING_DIMENSIONS` | per provider | 384 / 1536 / 256 |
| `OPENAI_API_KEY` | | Required only for `openai` |
| `RAG_DB_PATH` | `./rag_data` | Derived index; safe to delete |
| `RAG_CHUNK_SIZE` | `3500` | Maximum characters per chunk |
| `RAG_CHUNK_OVERLAP` | `350` | Characters shared with the next chunk (0 allowed) |
| `RAG_EXCLUDE_FOLDERS` | `Templates` | Comma-separated, vault-relative; nested paths allowed |
| `RAG_MAX_CHUNKS_PER_NOTE` | `2` | Cap per note in search results |
| `RAG_MIN_SIMILARITY` | `0` (off) | Drop vector matches below this cosine similarity |
| `RAG_AUTO_SYNC` | `true` | Sync changed notes before each search |
## Design decisions
- **LanceDB** runs embedded, with no server, and supports vector search and a BM25 full-text index on the same table. The index is a derived cache and the Markdown files remain the source of truth.
- **Hybrid instead of vector-only.** Embeddings blur exact identifiers such as `UN 3480` or `Art. 102`, while BM25 misses paraphrases. Rank fusion needs no score calibration between the two.
- **Keyword weight 1.15** was measured, not guessed. On the real vault, 1.0 lowered hit@1 from 0.958 to 0.917 (local) and from 0.667 to 0.583 (hash), and 1.5 gave no further gain.
- **Chunk size stays at 3,500.** A sweep from 800 to 3,500 characters moved real-vault hit@1 by at most one query in 24: smaller chunks helped recall slightly and hurt hit@1 slightly. That is noise, so the default was not changed. Note that bge-small reads only the first 512 tokens (~2,000 characters) of a chunk. The rest of a long chunk is still found by keyword search but not by vector search, which is the main reason to lower the size for prose-heavy vaults.
- **Minimum similarity is off by default.** With bge-small, relevant queries score about 0.67–0.74 and off-topic ones about 0.51, so `0.6` is a reasonable starting point. The eval set has no unanswerable queries to tune it on.
## Security
- `read_note` normalises the path, resolves symlinks and rejects anything outside the vault, in a hidden folder (`.obsidian`) or in an excluded folder. Its errors never reveal absolute paths.
- Folder and tag filters are escaped for SQL string literals and `LIKE` wildcards.
- The server instructions tell Claude to treat retrieved note text as untrusted data. This is the main defence against prompt injection from note content.
- With the default provider, no note content leaves the machine.
## Development
```bash
npm test # 26 tests: parsing, sync, retrieval, sandboxing, MCP end-to-end
npm run lint
npm run bench # sample-vault benchmark, local and local-hash embeddings
npm run bench -- --vault ~/MyVault --evals my-evals.jsonl --expand-links
```
Evaluation cases are JSON lines: `{"query": "...", "expectedSources": ["Folder/Note.md"], "folders": [], "tags": []}`. Files under `evals/private/` are gitignored.
## Roadmap
- Rerank the top candidates with a cross-encoder.
- Watch the vault for changes instead of syncing before each search.
- Add unanswerable queries to the eval set to tune `RAG_MIN_SIMILARITY`.
## License
[MIT](LICENSE)
This server cannot be deployed
Maintenance
ActivityMaintained
ResponsivenessNo issues