Skip to main content
Glama

Vault RAG MCP

Local multilingual retrieval for Markdown knowledge bases, exposed through a read-only Model Context Protocol server.

Vault RAG MCP retrieving a cited policy from an allowlisted Markdown vault

Point it at the smallest folder an assistant should be allowed to read. It chunks Markdown by heading, generates embeddings locally, stores a persistent SQLite index, and gives MCP clients four narrow tools that return evidence with paths and line numbers.

No API key. No write tool. No bulk upload of the corpus.

Privacy boundary: indexing and embeddings stay local. If the MCP host uses a cloud model, the passages it retrieves can still be sent to that model. This project reduces the data exposed per answer; it does not turn a cloud LLM into a local one.

Why this exists

Giving an assistant filesystem access is easy. Giving it the minimum useful access is the engineering problem.

Vault RAG MCP separates retrieval from generation:

ALLOWLISTED MARKDOWN ROOT
          │
          ▼
 heading-aware chunker ──▶ multilingual local embeddings
          │                         │
          └──────────────▶ SQLite vector index
                                      ▲
                                      │ cosine ranking
MCP HOST ──▶ search_vault(query) ─────┘
    │                 │
    │                 └── path + heading + exact line range
    └────▶ read_vault_note(path, lines)

Every operation ──▶ private JSONL audit trail

The server never calls an LLM. Claude, another MCP host, or a local model owns answer generation; this process only retrieves source evidence.

Related MCP server: team-docs-mcp

Verified sample run

The included five-note corpus produces 14 heading-aware chunks with the default model. A real offline stdio smoke run produced:

Check

Result

Protocol discovery

All four tools listed and callable

Incremental refresh with no changes

5 unchanged files, 0 rewritten chunks, 2 ms

English policy query

Projects/Support-Playbook.md, lines 11–13, score 0.8703

English query over a Spanish note

Projects/Asistente-Universidad.md, lines 5–7, score 0.8460

Timings depend on hardware. Similarity scores rank passages; they are not probabilities or guarantees of correctness.

What it exposes

Tool

Purpose

Source mutation

vault_status

Show index readiness and the enforced boundary

None

refresh_vault_index

Incrementally embed changed Markdown notes

None

search_vault

Rank passages semantically, optionally under a path prefix

None

read_vault_note

Read a bounded line range from one validated note

None

Tool descriptions tell the host to treat note content as untrusted data and to cite paths and line numbers.

Quickstart

Requirements: Python 3.11–3.13 and uv. CI tests all three supported Python versions on Linux; the release smoke run was also verified locally on macOS. The project defaults to Python 3.12 in .python-version, and uv can download it automatically.

git clone https://github.com/rickquant/vault-rag-mcp
cd vault-rag-mcp
uv sync

Build the included sample index:

uv run vault-rag-mcp index --vault sample_vault

The first run downloads intfloat/multilingual-e5-small. Source notes are embedded on the same machine.

Search it:

uv run vault-rag-mcp search \
  "When should a delayed package be investigated?" \
  --vault sample_vault \
  --limit 3

Read the exact source lines returned by search:

uv run vault-rag-mcp read \
  Projects/Support-Playbook.md \
  --vault sample_vault \
  --start-line 10 \
  --end-line 14

Connect it to Claude Desktop

Use absolute paths. Claude Desktop may launch stdio servers from an undefined working directory, and GUI applications often do not inherit the same PATH as a terminal.

Open Claude Desktop → Settings → Developer → Edit Config and add:

{
  "mcpServers": {
    "vault-rag": {
      "command": "/ABSOLUTE/PATH/TO/uv",
      "args": [
        "--directory",
        "/ABSOLUTE/PATH/TO/vault-rag-mcp",
        "run",
        "vault-rag-mcp",
        "serve"
      ],
      "env": {
        "VAULT_ROOT": "/ABSOLUTE/PATH/TO/YOUR/MARKDOWN/FOLDER",
        "VAULT_MODEL_LOCAL_FILES_ONLY": "true"
      }
    }
  }
}

Fully quit and reopen Claude Desktop after changing the config.

Example prompt:

Search my vault for the decision about local embeddings. Answer only from the retrieved evidence and cite the note plus line range.

Connect it to Claude Code

claude mcp add vault-rag \
  --scope user \
  --env VAULT_ROOT=/ABSOLUTE/PATH/TO/YOUR/MARKDOWN/FOLDER \
  --env VAULT_MODEL_LOCAL_FILES_ONLY=true \
  -- /ABSOLUTE/PATH/TO/uv \
  --directory /ABSOLUTE/PATH/TO/vault-rag-mcp \
  run vault-rag-mcp serve

claude mcp get vault-rag

Configuration

Environment variable

Default

Meaning

VAULT_ROOT

required

The single allowlisted Markdown root

VAULT_EMBEDDING_MODEL

intfloat/multilingual-e5-small

Local Sentence Transformers model

VAULT_EMBEDDING_DEVICE

cpu

Torch device used for embeddings

VAULT_MODEL_LOCAL_FILES_ONLY

false

Refuse model-network access; enable after the model is cached

VAULT_EXCLUDE_DIRS

.git,.obsidian,.trash,.venv,__pycache__,node_modules

Directory names skipped recursively

VAULT_MAX_FILE_BYTES

1000000

Per-note indexing and read limit

VAULT_MAX_READ_CHARS

20000

Maximum text returned by one read

VAULT_MAX_RESULTS

10

Hard cap on search results

VAULT_INDEX_PATH

OS cache directory

SQLite index location

VAULT_AUDIT_PATH

OS state directory

JSONL audit location

VAULT_AUDIT_QUERY_TEXT

false

Store raw searches instead of hashes

The default index and audit paths are namespaced by a hash of the resolved vault root. The server refuses to place either file inside VAULT_ROOT.

The first model load needs network access. After that succeeds once, set VAULT_MODEL_LOCAL_FILES_ONLY=true in the MCP server environment to make model loading fail closed instead of checking the network.

Security boundary

Enforced in code:

  • One resolved root is the filesystem allowlist.

  • Only .md files are scanned or read.

  • Absolute paths, .., null bytes, non-Markdown files, oversized files, and every symlinked path are rejected.

  • Hidden and configured directories are skipped.

  • There are no create, edit, move, delete, shell, network, or arbitrary-path tools.

  • Search returns at most 10 bounded excerpts; note reads are character-capped.

  • The SQLite index and JSONL audit log are created with owner-only permissions where the operating system supports them.

  • Audit records hash query text by default.

Not guaranteed:

  • Retrieved passages may leave the machine through the MCP host's model provider.

  • This server cannot constrain other tools enabled in the same host.

  • Local processes running as the same operating-system user may read the source folder, index, or audit log.

  • Semantic retrieval can miss relevant text. Source citations make results inspectable; they do not make ranking infallible.

See docs/SECURITY.md for the threat boundaries and design decisions.

How indexing works

  1. Scan Markdown files without following symlinks.

  2. Remove YAML frontmatter from retrieval text.

  3. Split each note by its Markdown heading hierarchy, then into overlapping bounded chunks.

  4. Prefix passages and queries according to the multilingual E5 retrieval model.

  5. Normalize vectors and store them as float32 blobs in SQLite.

  6. On refresh, re-embed only files whose size or nanosecond modification time changed; remove deleted files.

  7. Rank the in-memory matrix with cosine similarity and return inspectable source metadata.

This deliberate brute-force search is appropriate for personal and small-team knowledge bases. A distributed vector database would add machinery without improving this demo's real use case.

Development

Tests inject a deterministic network-free embedder. They exercise incremental indexing, retrieval, deletion, source reads, path attacks, audit behavior, and all four tools through an in-memory MCP client/server transport.

uv sync --group dev
uv run ruff check .
uv run ruff format --check .
uv run pytest

To verify the packaged command through a real stdio subprocess after the model is cached:

uv run python scripts/smoke_stdio.py

Current scope and non-goals live in docs/SCOPE.md.

Limitations

  • Manual or tool-invoked refresh; no filesystem watcher.

  • One local user and one configured root.

  • Markdown only.

  • Stdio transport only in the portfolio release.

  • No answer generator, UI, OAuth, cloud deployment, or write operations.

  • Intended for small knowledge bases, not millions of chunks.

License

MIT

Available Tools

4 tools
read_vault_noteA

Read a bounded line range from an allowlisted Markdown note.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYes
end_lineNo
start_lineNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the behavioral disclosure burden. It does add meaningful behavior: reads only a bounded range and only from allowlisted notes. However, it does not mention error behavior, what happens when ranges are invalid, or whether the operation is strictly read-only beyond the verb 'Read'.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with no filler. The core action and constraint are front-loaded, making it easy for an agent to parse quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is relatively simple and has an output schema, so the description does not need to explain return values. However, with no annotations and no parameter descriptions, it omits enough behavioral and boundary details that an agent may not understand defaults or edge-case behavior when invoking the tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for undocumented parameters. It only says 'bounded line range', which loosely maps to start_line and end_line, but it does not explain line numbering, inclusivity, the meaning of end_line being null, or the expected path format. The parameter names carry most of the meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb, 'Read', with a specific resource, 'a bounded line range from an allowlisted Markdown note'. This clearly identifies the operation and distinguishes it from sibling tools like search_vault, vault_status, and refresh_vault_index, which involve searching, status, and indexing rather than reading note content.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'bounded line range' and 'allowlisted Markdown note' imply the intended use case: retrieving a specific section of an allowed note. However, it does not explicitly state when to prefer this tool over search_vault or refresh_vault_index, nor does it mention any exclusions beyond the allowlist constraint.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

refresh_vault_indexB

Incrementally index changed Markdown notes. This never modifies source notes.

ParametersJSON Schema
NameRequiredDescriptionDefault
forceNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure. It does state a key guarantee ('never modifies source notes'), which is valuable. However, it does not disclose what it actually modifies (the index), potential side effects, concurrency behavior, error conditions, or whether permissions are required. The description provides one behavioral fact but omits many others that an agent would need for safe invocation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two short sentences with no redundant wording. It front-loads the core action and follows with a critical guarantee. It is concise and well-structured, though it could be slightly expanded to cover the parameter without losing brevity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given there is an output schema, return values are not the issue. However, the tool has an unexplained parameter and no usage guidance. For a simple tool, the lack of parameter semantics and usage context makes it incomplete. The agent knows what the tool does but not when or how to control it, leaving important gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has zero description coverage, and the description does not mention the 'force' parameter at all. An agent cannot infer what 'force' does (e.g., whether it forces a full re-index or overrides caching). The name hints at forcing, but without an explanation, the parameter semantics are completely opaque. The description adds no meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action: 'Incrementally index changed Markdown notes.' This specifies the verb (index), the resource (Markdown notes), and the incremental nature. It also adds a distinct guarantee ('never modifies source notes') that separates it from mutation tools. It is easily distinguishable from sibling tools like vault_status, search_vault, and read_vault_note, which serve different purposes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus the alternatives. It does not say 'use this to update the search index' or mention any prerequisites, exclusions, or typical scenarios. The agent must infer that refresh_vault_index is for maintaining the index, but no explicit direction is given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_vaultA

Find semantically relevant passages with source paths and exact line ranges.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
queryYes
min_scoreNo
path_prefixNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral disclosure burden. It gives some useful detail about output ('source paths and exact line ranges') and establishes this is a read/search operation, but it does not mention behavior around missing index, result limits, or whether the search is exact or fuzzy.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence with no filler. It communicates the primary purpose and a key output characteristic efficiently.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is enough for a basic query-only call, but not for correct use of the optional parameters. With no annotations and no schema descriptions, the agent is missing important context about 'min_score' and 'path_prefix', which are not obvious from the one-line description.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description only loosely implies that 'query' is a semantic search term. It gives no meaning for 'limit', 'min_score', or 'path_prefix', leaving the agent to guess their roles entirely from names and defaults.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific action ('Find'), a resource ('vault'), and defines the core capability: semantic relevance with source paths and exact line ranges. This clearly differentiates from siblings like vault_status, refresh_vault_index, and read_vault_note.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The semantic-search phrasing implies when to use it, but there is no explicit guidance on when to choose this over read_vault_note or when an index refresh might be required first. Usage context is present but not spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

vault_statusA

Show index readiness and the server's read-only security boundary.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the behavioral disclosure burden. The verb 'Show' signals a non-mutating status operation, and the described output is clear, but it does not explicitly state that no changes are made or mention authentication or access considerations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no filler. Each part adds distinct information: readiness status and the read-only security boundary.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With zero parameters and an output schema present, the description does not need to explain return values. It covers the two key aspects of the tool's result and is adequate for a simple status check, though it could add a sentence on how this relates to refreshing or searching the vault.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so there is nothing the description needs to explain beyond the schema. Baseline 4 applies because parameter semantics are moot and no confusion is possible.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Show') with a clear resource: index readiness and the server's read-only security boundary. It is not a tautology and distinguishes this status tool from the sibling mutation/lookup tools without needing to open the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use it: when you need index readiness or security-boundary information. However, it does not explicitly compare against siblings like refresh_vault_index or state when not to use this tool, so the guidance is only implicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 4 tool updatesv0.1.0
    • First observedread_vault_note
    • First observedrefresh_vault_index
    • First observedsearch_vault
    • First observedvault_status

TDQS

A3.7/5.0

Scored across 4 tools

Disambiguation5/5

Each tool has a clearly distinct responsibility: status, indexing, search, and reading. There is no overlap or ambiguity between actions an agent could confuse.

Naming Consistency4/5

Most tools follow a verb_noun pattern (refresh_vault_index, search_vault, read_vault_note). vault_status deviates slightly by using a noun phrase instead of a verb prefix, but the naming style is otherwise uniform and readable.

Tool Count5/5

Four tools is a well-scoped size for a focused RAG-over-vault server. Each tool covers a necessary part of the workflow without redundancy or bloat.

Completeness5/5

For a read-only RAG server, the set covers the full lifecycle: checking index health, refreshing the index, searching semantically, and reading source passages. No obvious missing operations are needed to accomplish its stated purpose.

Maintenance

ActivityStale
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers