Skip to main content
Glama
mpandudc

Obsidian Hybrid RAG MCP Server

by mpandudc

Obsidian Hybrid RAG MCP Server

CI Python MCP Embedding Reranker Vector Store License: MIT

A local Model Context Protocol (MCP) server that gives agents two-stage hybrid search and safe editing tools over an Obsidian markdown vault.

Search combines SQLite FTS5 (BM25) and sqlite-vec dense vectors (BAAI/bge-m3), fuses them with Reciprocal Rank Fusion and reranks with a jinaai/jina-reranker-v2 cross-encoder — all in one Python process and one SQLite file, without a vector database daemon or a RAG framework.


🏛️ Architecture Overview

                          ┌──────────────────────────┐
                          │   Obsidian Vault (.md)   │
                          └─────────────┬────────────┘
                 .vaultignore / size cap │ (fence-aware heading chunker)
                                        ▼
                 ┌──────────────────────────────────────────────┐
                 │       Single Embedded SQLite Database        │
                 │  ┌────────────────────┐ ┌──────────────────┐ │
                 │  │    SQLite FTS5     │ │    sqlite-vec    │ │
                 │  │ (weighted BM25)    │ │ (1024-dim + path │ │
                 │  │                    │ │  metadata)       │ │
                 │  └─────────┬──────────┘ └─────────┬────────┘ │
                 └────────────┼──────────────────────┼──────────┘
                              └──────────┬───────────┘
                                         ▼
                 ┌──────────────────────────────────────────────┐
                 │ Stage 1: Reciprocal Rank Fusion (RRF, k=60)  │
                 │ + folder / tag / status filters              │
                 └───────────────────────┬──────────────────────┘
                                         ▼
                 ┌──────────────────────────────────────────────┐
                 │ Stage 2: Cross-Encoder Reranker (optional    │
                 │ score floor) + max 2 chunks per note         │
                 └───────────────────────┬──────────────────────┘
                                         ▼
                 ┌──────────────────────────────────────────────┐
                 │   FastMCP (stdio / SSE / streamable HTTP)    │
                 │  (Hermes Agent / Claude Desktop / Cursor)    │
                 └──────────────────────────────────────────────┘

Related MCP server: obsidian-rag-mcp

✨ Key Features

  1. In-process, single-file index. FTS5 and sqlite-vec live in one vault-index.db. The index schema is versioned (PRAGMA user_version); an outdated index is rebuilt automatically.

  2. Fence-aware, line-exact chunking. Sections split on real headings only — a # comment inside a fenced code block is code, not a heading. Chunks never exceed VAULT_CHUNK_CHAR_LIMIT and report exact source line ranges. Frontmatter is parsed (tags, status) but not embedded.

  3. Junk-resistant indexing. Notes matched by .vaultignore, larger than VAULT_MAX_FILE_BYTES, or marked index: false are recorded as skipped. Lines longer than VAULT_MAX_LINE_CHARS (pasted JSON / base64 blobs) are dropped from the indexed text.

  4. Memory-safe re-indexing. By default the server re-indexes in a background thread that reuses the already-loaded embedding model, so a write never loads a second bge-m3. Only one indexer runs at a time (<db>.lock); changes made during a pass trigger exactly one more pass instead of being dropped.

  5. Safe editing for agents. Writes are atomic (temp file + rename), confined to the vault, refuse to overwrite unless asked (with a backup in .trash/vault-mcp/), support optimistic concurrency (expected_hash), keep CRLF line endings, and return a [[wikilink]] report.

  6. Vault hygiene tools. vault_lint finds broken links, orphans, missing hub links / frontmatter and off-vocabulary status: values; vault_move renames a note and rewrites every link to it.

  7. Measurable retrieval. vault-eval reports hit@k, recall@k and MRR per mode on a golden query set, plus the reranker score distribution to calibrate a relevance floor.


🚀 Installation & Quickstart

Python 3.10–3.12. uv recommended.

git clone https://github.com/mpandudc/obsidian-hybrid-rag-mcp.git
cd obsidian-hybrid-rag-mcp
uv venv .venv && source .venv/bin/activate

# CPU-only torch first, then the package with the model extras
uv pip install torch --index-url https://download.pytorch.org/whl/cpu
uv pip install -e ".[models]"

The core install (pip install -e .) is enough for keyword search and the editing / lint tools; semantic and hybrid search and indexing need the models extra.

Download the models once (the server runs with HF_HUB_OFFLINE=1):

python -c "from sentence_transformers import SentenceTransformer; SentenceTransformer('BAAI/bge-m3')"
python -c "from fastembed.rerank.cross_encoder import TextCrossEncoder; TextCrossEncoder('jinaai/jina-reranker-v2-base-multilingual')"

Build the index:

vault-indexer --vault-path "/path/to/vault" --rebuild   # full build
vault-indexer --vault-path "/path/to/vault"             # incremental

Environment configuration

Variable

Default

Purpose

OBSIDIAN_VAULT_PATH (or VAULT_PATH)

~/vaults/pandu-second-brain

Vault root

VAULT_INDEX_DB (or INDEX_DB_PATH)

~/.hermes/vault-index.db

SQLite index file

VAULT_INDEX_MODE

inprocess (command if VAULT_INDEXER_COMMAND is set)

inprocess / command / off — how write tools re-index

VAULT_INDEXER_COMMAND (legacy INDEXER_RUNNER)

—

Command for command mode (e.g. a memory-capped wrapper). Required in that mode; the server refuses to start without it

VAULT_MAX_FILE_BYTES

524288

Notes above this size are skipped

VAULT_MAX_LINE_CHARS

10000

Longer lines are dropped from indexed text

VAULT_CHUNK_CHAR_LIMIT

1500

Max characters per chunk

VAULT_EMBED_MAX_SEQ_LENGTH

1024

Token cap per chunk for bge-m3

VAULT_EMBED_BATCH_SIZE

8

Encode batch size

VAULT_MAX_CHUNKS_PER_NOTE

2

Result diversity cap per note

VAULT_RERANK_CHARS

600

Characters of each candidate shown to the reranker

VAULT_RERANK_POOL

8

Top RRF candidates reranked (at least limit)

VAULT_RERANK_THREADS

CPUs in cpuset (max 4)

Reranker ONNX threads

VAULT_MIN_RERANK_SCORE

unset (off)

Drop reranked results below this score — calibrate with vault-eval

VAULT_READ_MAX_CHARS

8000

Default vault_read page size

VAULT_STATUS_VALUES

draft,active,approved,verified,completed,falsified,superseded,archived

Allowed frontmatter status values for vault_lint

MODEL_IDLE_TIMEOUT

300

Seconds before models are unloaded from RAM

FASTEMBED_CACHE_DIR

~/.cache/fastembed

Reranker model cache

MCP_ALLOW_REMOTE

unset

1 allows binding a non-loopback address

.vaultignore

Optional file in the vault root; one glob per line, # for comments:

# a folder
clippings/
# a path pattern
resources/**/Livro_*.md
# a file-name pattern
*.draft.md

Skipped notes stay readable through vault_read; they are only left out of search. vault_status lists them with the reason.


🔌 MCP Client Configuration

Claude Desktop (claude_desktop_config.json)

{
  "mcpServers": {
    "obsidian-vault": {
      "command": "/path/to/obsidian-hybrid-rag-mcp/.venv/bin/vault-mcp",
      "env": {
        "OBSIDIAN_VAULT_PATH": "/path/to/your/obsidian-vault",
        "VAULT_INDEX_DB": "/path/to/vault-index.db"
      }
    }
  }
}

One daemon means one model copy for every agent profile:

# ~/.config/systemd/user/vault-mcp.service
[Unit]
Description=Obsidian Hybrid RAG FastMCP Daemon (SSE)
After=network.target

[Service]
Type=simple
ExecStart=/path/to/.venv/bin/vault-mcp --transport sse --host 127.0.0.1 --port 8765
Environment=OBSIDIAN_VAULT_PATH=/path/to/vault
# In-process re-indexing shares the daemon's model, so cap the daemon itself:
MemoryMax=4G
MemorySwapMax=512M
Restart=always
RestartSec=5

[Install]
WantedBy=default.target
hermes config set mcp_servers.vault.url http://127.0.0.1:8765/sse
hermes config set mcp_servers.vault.transport sse

Security: the write tools have no authentication. The server refuses to bind anything but loopback unless you pass --allow-remote (or MCP_ALLOW_REMOTE=1) — only do that behind an authenticating proxy.

Cron keeps the index fresh for edits made outside the MCP (Obsidian on phone/PC):

*/30 * * * * systemd-run --user --scope -p MemoryMax=3G -p MemorySwapMax=512M /path/to/.venv/bin/vault-indexer

The CLI indexer and the daemon share <db>.lock, so they never index concurrently.


🛠️ MCP Tools

Tool

Purpose

vault_search(query, limit=5, mode="hybrid", folder="", tags="", status="")

Hybrid / keyword / semantic search. folder filters natively in the vector index; tags (all must match) and status (any) filter on frontmatter.

vault_read(rel_path, heading="", start_line=1, max_chars=8000)

Paged read of a note or section. Header shows the line range and a sha256 prefix; a missing heading lists the available headings instead of dumping the note.

vault_list(folder="", limit=200)

Notes and titles under a folder.

vault_recent(limit=20, folder="", days=0)

Most recently modified notes.

vault_backlinks(note, limit=50)

Notes linking to a note, with the linking line.

vault_write(rel_path, content, title="", tags="", overwrite=False, expected_hash="")

Create a note with frontmatter; replacing one needs overwrite=true and backs it up.

vault_append(rel_path, content, heading="", expected_hash="")

Append at the end or under a heading (created if missing).

vault_edit(rel_path, old_text, new_text, replace_all=False, expected_hash="")

Exact-string replace; refuses missing or ambiguous matches.

vault_move(src, dst, update_links=True)

Rename/move a note and rewrite every wikilink to it (aliases, #heading, ![[embeds]] kept; code blocks untouched).

vault_lint(folder="", limit=30)

Broken links, orphans, missing hub link / frontmatter, invalid status, oversized notes, blob lines.

vault_status()

OK / INDEXING / STALE / INCONSISTENT / OUTDATED SCHEMA, counts, skipped notes, index mode, model RAM state.


📏 Evaluating retrieval

Write a golden set (see eval/golden.example.json) and run:

vault-eval --golden eval/golden.json --k 5

It prints hit@k / recall@k / MRR per mode, every miss, and the reranker score distribution of relevant vs irrelevant results. Use the relevant-score p10 to pick VAULT_MIN_RERANK_SCORE, and re-run after changing chunk size, weights or models.

Reference run on the author's vault (283 notes, 30 queries from eval/golden.example.json, k=5, 3 vCPU, no GPU):

mode / rerank budget

hit@5

MRR

avg latency

keyword

0.97

0.77

3 ms

semantic

1.00

0.93

~0.15 s

hybrid, 20 candidates × 1500 chars

1.00

0.97

11.9 s

hybrid, 12 × 800

1.00

0.93

4.8 s

hybrid, 8 × 600 (default)

1.00

0.96

2.6 s

The cross-encoder is almost all of hybrid latency on CPU; use mode="semantic" when speed matters more than the last few points of ranking. Relevant results scored −0.78 … 2.03 (p10 0.15) with the default budget, so a floor of -1.0 keeps every relevant hit.


🔁 Upgrading from 1.x

  • The package moved from src to obsidian_hybrid_rag_mcp: use the vault-mcp / vault-indexer entry points (or python -m obsidian_hybrid_rag_mcp.server).

  • The index schema is now v2. The first indexer run rebuilds it automatically (vault_status says OUTDATED SCHEMA until then).

  • vault_write no longer overwrites silently: pass overwrite=true.

  • Re-indexing defaults to in-process. To keep an external wrapper, set VAULT_INDEX_MODE=command and VAULT_INDEXER_COMMAND (the legacy INDEXER_RUNNER is still read).

  • Binding a non-loopback address now requires --allow-remote.


🧪 Development

uv pip install -e ".[dev]"
ruff check .
pytest

The tests use a deterministic fake embedder, so they need neither torch nor the models.


📄 License

MIT — see LICENSE.

Developed by Muhammad Pandu Dwi Cahyo.

Available Tools

3 tools
get_noteGet NoteA

Retrieve full markdown content of a specific note with line number pagination.

ParametersJSON Schema
NameRequiredDescriptionDefault
rel_pathYesNote path relative to vault root (e.g. 'trading/cuantum-setup.md')
limit_linesNoMaximum lines to return in single request
offset_lineNoLine number to start reading from (1-indexed)

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the burden of disclosing behavior. It clearly signals a read-only operation ('Retrieve') and describes pagination behavior, which helps an agent understand it may receive partial content with limit/offset. It does not discuss error cases or side effects, but for a read tool the core behavioral traits are covered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no filler. It states the action, resource, and key pagination behavior in a compact way that is easy to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a straightforward read tool with a complete input schema and an output schema, the description covers the essential behavior. It could be more complete by explicitly noting that full content may require multiple paginated calls and by referencing sibling tools, but the current definition is sufficient for correct invocation in most cases.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the input schema already explains rel_path, limit_lines, and offset_line. The description's mention of 'line number pagination' adds conceptual context but not new parameter-level details, so the baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Retrieve'), a clear resource ('full markdown content of a specific note'), and a distinctive mechanism ('line number pagination'). This makes it easy to distinguish from sibling tools search_vault and sync_vault, which are clearly different operations.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the tool should be used when you know the exact note path and need its full content, but it does not explicitly say when to prefer it over search_vault or sync_vault. There is no mention of alternatives or exclusions, so guidance is only implicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_vaultSearch VaultA

Search your Obsidian knowledge base using state-of-the-art Hybrid RAG. Combines SQLite FTS5 (BM25 lexical), BAAI/bge-m3 (1024-dim dense semantic vector), Reciprocal Rank Fusion (RRF), and Jina Reranker v2 Cross-Encoder scoring.

ParametersJSON Schema
NameRequiredDescriptionDefault
queryYesNatural language question or search phrase
top_kNoNumber of highest-ranking passages to return (default 5)
heading_filterNoOptional substring filter for document section headers

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the burden for behavioral disclosure. It does reveal the retrieval pipeline (BM25, dense embeddings, RRF, reranking), which is useful behavioral context. However, it does not explicitly state that the tool is read-only, mention any network/local dependencies, or note potential latency from reranking.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with a clear front-loaded purpose. The second sentence lists the technical underpinnings, which is informative but slightly dense with jargon; it still earns its place by explaining what 'Hybrid RAG' means. No redundant filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a search tool with a full input schema and an output schema, the description covers the core behavior and mechanism sufficiently. The main gap is the missing guidance on tool selection relative to siblings, and the lack of explicit safety/read-only context due to absent annotations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema fully documents all three parameters (query, top_k, heading_filter) with descriptions and defaults, so schema coverage is 100%. The description adds no parameter-level detail, but it doesn't need to since the schema already covers it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Search your Obsidian knowledge base.' It clearly identifies the tool as a search operation, distinct from the sibling tools get_note and sync_vault. The additional technical detail (Hybrid RAG) reinforces its specialized role.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit guidance about when to use search_vault versus get_note or sync_vault. The description states what the tool does but provides no alternatives, exclusions, or contextual conditions that would help an agent decide between the siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sync_vaultSync VaultB

Trigger incremental sync on the Obsidian vault. Scans for created, modified, or deleted notes and updates dense vectors & FTS5 index.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. The description says it triggers a sync and updates vectors and indexes, but it doesn't mention whether it's safe to call (read-only vs mutating), what side effects occur (e.g., does it modify the vault or just the index?), or any performance implications. An agent cannot infer the safety profile from this sparse description.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise, with two sentences: the first states the primary action, and the second elaborates on scope and effects. It's front-loaded with the key verb, and there's no wasted words. However, it could have been structured to include a leading phrase for when to use it, but it's still efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given that this tool has no parameters and an output schema exists, the description doesn't need to explain return values. However, the description lacks context on when to use it relative to siblings (e.g., after editing notes) and what the impact is. It's adequate for a simple trigger tool but could be improved by mentioning that it's for keeping search indexes up to date, which is implied but not explicit.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Since there are no parameters, the description needs to explain the scope of the sync, which it does by mentioning it's incremental and scans for created/modified/deleted notes. With zero parameters, the baseline is 4, and the description doesn't need to compensate for parameter documentation; it provides clear semantics for what the tool does without any input.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's verb ('Trigger incremental sync') and the resource ('Obsidian vault'). It also specifies what the sync does: scanning for created, modified, or deleted notes and updating dense vectors & FTS5 index. This is specific and distinguishes it from search_vault and get_note, which are read operations, though it doesn't explicitly name them as alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use this tool: whenever the vault's notes have changed and you need indexes updated. However, it doesn't explicitly state when not to use it (e.g., for full re-index) or compare it to search_vault or get_note as alternatives for different purposes. The context that it's for syncing is clear, but exclusions are absent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 3 tool updatesv1.0.0
    • First observedget_note
    • First observedsearch_vault
    • First observedsync_vault

TDQS

A3.9/5.0

Scored across 3 tools

Disambiguation5/5

Each tool has a clearly distinct purpose: search_vault queries, get_note retrieves a specific note, and sync_vault updates the index. There is no overlap or ambiguity between them.

Naming Consistency5/5

All tool names follow a consistent verb_noun snake_case pattern: search_vault, get_note, sync_vault. The naming is predictable and uniform.

Tool Count5/5

Three tools is well-scoped for a focused Obsidian RAG server: search, retrieve, and sync. Each tool serves a necessary function without redundancy.

Completeness4/5

The core workflow of searching, retrieving, and syncing is covered. A minor gap is the lack of an explicit listing or status tool, but users can still work around this via search.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    Provides semantic search and keyword search over Obsidian notes, along with direct note retrieval, allowing external AI agents to query and access the vault.
    19
    BSD Zero Clause
  • F
    license
    Not graded
    quality
    A
    maintenance
    Enables semantic search and note management for Obsidian vaults via the Model Context Protocol, allowing LLMs to search, read, and index notes, PDFs, and web pages locally.
    -
  • A
    license
    Not graded
    quality
    A
    maintenance
    Provides a hybrid search engine for Obsidian vaults, enabling LLM agents to query notes with BM25 keyword and vector semantic search, metadata filtering, and sibling-document retrieval.
    124 npm
    MIT