Skip to main content
Glama
wirux

mcp-markdown-vault

by wirux
README.md
<div align="center">

# πŸ“ Markdown Vault MCP Server

**Headless semantic [MCP](https://modelcontextprotocol.io/) server for Obsidian, Logseq, Dendron, Foam, and any folder of markdown files.**

`npm install` and point it at a folder. Hybrid search, AST editing, zero-config embeddings. No app, no plugins, no API keys.

[![CI / Release](https://github.com/wirux/mcp-markdown-vault/actions/workflows/release.yml/badge.svg)](https://github.com/wirux/mcp-markdown-vault/actions/workflows/release.yml)
[![PR Check](https://github.com/wirux/mcp-markdown-vault/actions/workflows/pr-check.yml/badge.svg)](https://github.com/wirux/mcp-markdown-vault/actions/workflows/pr-check.yml)
[![npm version](https://img.shields.io/npm/v/@wirux/mcp-markdown-vault?color=cb3837&logo=npm)](https://www.npmjs.com/package/@wirux/mcp-markdown-vault)
[![Docker](https://img.shields.io/badge/ghcr.io-mcp--markdown--vault-blue?logo=docker)](https://github.com/wirux/mcp-markdown-vault/pkgs/container/mcp-markdown-vault)
[![License: MIT](https://img.shields.io/badge/license-MIT-green.svg)](LICENSE)
[![TypeScript](https://img.shields.io/badge/TypeScript-5.x-3178c6?logo=typescript&logoColor=white)](https://www.typescriptlang.org/)
[![Node.js](https://img.shields.io/badge/Node.js-%3E%3D22-339933?logo=node.js&logoColor=white)](https://nodejs.org/)
[![Tests](https://img.shields.io/badge/tests-568%20passed-brightgreen?logo=vitest&logoColor=white)](#-testing)
[![mcp-markdown-vault MCP server](https://glama.ai/mcp/servers/wirux/mcp-markdown-vault/badges/score.svg)](https://glama.ai/mcp/servers/wirux/mcp-markdown-vault)

</div>

<div align="center">

![Markdown Vault MCP Server Demo](assets/demo.gif)

</div>

---

## πŸ’‘ Why this server?

> **TL;DR** β€” One `npx` command. No running app. No plugins. No vector DB. Semantic search works out of the box.

| | Differentiator | Details |
|---|---|---|
| 🚫 | **No app or plugins required** | Most Obsidian MCP servers ([mcp-obsidian](https://github.com/MarkusPfundstein/mcp-obsidian), [obsidian-mcp-server](https://github.com/cyanheads/obsidian-mcp-server)) need Obsidian running with the Local REST API plugin. This server reads and writes `.md` files directly β€” point it at a folder and go. |
| 🧠 | **Built-in semantic search, zero setup** | Hybrid search: cosine-similarity vectors + TF-IDF + word proximity. Local embeddings (`@huggingface/transformers`, `all-MiniLM-L6-v2`, 384d) download on first run. No API keys, no external services. Ollama optional for higher quality. |
| πŸ”¬ | **Surgical AST-based editing** | `remark` AST pipeline patches specific headings or block IDs without touching the rest of the file. Freeform line-range & string replace as fallback. Levenshtein fuzzy matching handles LLM typos. |
| πŸ”“ | **Tool-agnostic** | Obsidian vaults, Logseq graphs, Dendron workspaces, Foam, or any plain folder of `.md` files. If it's markdown, it works. |
| πŸ“¦ | **Single package, no infrastructure** | Unlike Python alternatives that need ChromaDB or other vector stores, everything runs in one Node.js process. `npx @wirux/mcp-markdown-vault` and you're running. Docker image available. |

<div align="center">

πŸ’Ž **Obsidian** Β· πŸ““ **Logseq** Β· 🌳 **Dendron** Β· 🫧 **Foam** Β· πŸ“‚ **Any `.md` folder**

</div>

---

## ✨ Features

| | Feature | Description |
|---|---|---|
| πŸ—‚οΈ | **Headless vault ops** | Read, create, update, edit, delete `.md` notes with strict path traversal protection |
| πŸ“‘ | **Read by heading** | Read a single section by heading title β€” returns only content under that heading (up to the next same-level heading), saving context window space |
| πŸ“¦ | **Bulk read** | Read multiple files and/or heading-scoped sections in a single call β€” reduces MCP round-trips with per-item fault tolerance |
| πŸ”¬ | **Surgical editing** | AST-based patching targets specific headings or block IDs β€” never overwrites the whole file |
| πŸ” | **Fragment retrieval** | Heading-aware chunking + TF-IDF + proximity scoring returns only relevant sections |
| πŸ“‚ | **Scoped search** | Optional directory filter for `global_search` and `semantic_search` β€” restrict results to specific folders to reduce noise |
| 🧠 | **Semantic search** | Hybrid vector + lexical search with background auto-indexing |
| ⚑ | **Zero-setup embeddings** | Built-in local embeddings via `@huggingface/transformers` β€” Ollama optional |
| πŸ”„ | **Workflow tracking** | Petri net state machine with contextual LLM hints |
| 🌐 | **Dual transport** | Stdio (single client) or SSE over HTTP (multi-client, Docker-friendly) |
| ✏️ | **Freeform editing** | Line-range replacement and string find/replace as AST fallback |
| 🏷️ | **Frontmatter management** | AST-based read and update of YAML frontmatter β€” safely manage tags, statuses, and metadata without corrupting file structure |
| πŸ‘€ | **Dry-run / diff preview** | Preview any edit operation as a unified diff without saving β€” set `dryRun=true` on any edit action |
| πŸ“ | **Templating / scaffolding** | Create new notes from template files with `{{variable}}` placeholder injection β€” refuses to overwrite existing files |
| πŸ—ΊοΈ | **Self-orienting vault context** | Assisted or manual `meta/overview.md` with host-visible `vault_scope` and live `vault://overview` context for connected agents |
| πŸ“¦ | **Batch edit** | Apply multiple edit operations in a single call β€” sequential execution, stops on first error, supports `dryRun`, max 50 ops |
| πŸ”— | **Backlinks index** | Find all notes linking to a given path β€” supports wikilinks and markdown links with line numbers and context snippets |
| 🎯 | **Typo resilience** | Levenshtein-based fuzzy matching for edit operations |

---

## πŸ› οΈ MCP Tools

| Tool | Actions | Description |
|---|---|---|
| πŸ“ **vault** | `list` `read` `create` `update` `delete` `stat` `create_from_template` | Full CRUD for vault notes + template scaffolding |
| ✏️ **edit** | `append` `prepend` `replace` `delete` `line_replace` `string_replace` `frontmatter_set` + `operations[]` batch mode | AST-based patching + freeform fallback + frontmatter update + batch edit (supports `dryRun` diff preview) |
| πŸ‘οΈ **view** | `search` `global_search` `semantic_search` `outline` `read` `frontmatter_get` `bulk_read` `backlinks` | Fragment retrieval, cross-vault search, hybrid semantic search, read by heading, frontmatter read, bulk read, backlinks |
| πŸ”„ **workflow** | `status` `transition` `history` `reset` | Petri net state machine control |
| βš™οΈ **system** | `status` `reindex` `overview` `overview_status` `prepare_overview` `save_overview` | Server health, indexing info, vault structure overview, assisted overview rebuild |

> All tool responses include contextual hints based on the current workflow state.

---

## πŸ’‘ Operational Guidance

### πŸ› οΈ Safe Editing
- **`dryRun=true`**: Highly recommended before destructive operations like `delete` or `replace` with `replaceMode="section"`.
- **Heading Disambiguation**: If multiple identical headings are found, the server returns `AMBIGUOUS_HEADING_TARGET` with a list of candidates. Use `blockId` to target specific elements if headings are not unique.
- **AST vs Freeform**: Always prefer AST operations (`append`, `prepend`, `replace`, `delete`) as they are structural. Use `string_replace` only as a last resort; it requires exact literal matches including whitespace and newlines.
- **`replaceMode`**: `replace` defaults to `body` (preserves the heading, replaces content). Set `replaceMode: "section"` to replace the heading node and all its child headings.
- **`returnContent`**: Set to `section` or `file` to see the results of your edit immediately in the tool response (max 8KB).

### πŸš€ Performance & Consistency
- **`bulk_read`**: Use this to read 2 or more files/sections concurrently. It is significantly faster than multiple sequential `view.read` calls.
- **Workflow State**: The `workflow` tool manages session-specific state used for contextual hints. It does **not** modify vault data or search indexes.
- **`system.reindex`**: Only use this for recovery or after making out-of-band file changes (e.g., via external scripts). Normal MCP edits automatically update backlinks and queue vector indexing.
- **`view.outline`**: Supports a `directory` parameter to get a flat list of headings across multiple files in a folder.

### πŸ§ͺ Batch Edits
- **Sequential Execution**: Operations in a batch are executed one by one. If one fails, the remaining are skipped.
- **Dry-run Asymmetry**: In `dryRun=false`, each operation sees the file state after previous operations. In `dryRun=true`, the file is never written, so sequential dependent operations (e.g., editing the same line twice) may produce different results than a live run.

---

## πŸš€ Quick Start

### Prerequisites

- [Node.js](https://nodejs.org/) >= 22
- *(Optional)* [Ollama](https://ollama.com/) for higher-quality embeddings

### πŸ“¦ Install from NPM

```bash
npm install -g @wirux/mcp-markdown-vault
```

Then run directly:

```bash
VAULT_PATH=/path/to/your/vault markdown-vault-mcp
```

### πŸ”Œ MCP Client Configuration

Add to your MCP client config (e.g. Claude Desktop, Claude Code):

```json
{
  "mcpServers": {
    "markdown-vault": {
      "command": "npx",
      "args": ["-y", "@wirux/mcp-markdown-vault"],
      "env": {
        "VAULT_PATH": "/path/to/your/vault"
      }
    }
  }
}
```

> `npx -y` auto-installs the package if not already present β€” no global install needed.

> **Try it in the browser:** You can test this server directly at [Glama Inspector](https://glama.ai/mcp/servers/@wirux/mcp-markdown-vault) β€” no local install required.

### 🐳 Docker

Pull the pre-built multi-arch image from GitHub Container Registry:

```bash
docker pull ghcr.io/wirux/mcp-markdown-vault:latest
```

Or use Docker Compose:

```bash
docker compose up
```

Edit `docker-compose.yml` to point at your markdown vault directory. The default compose file uses SSE transport on port 3000.

### πŸ› οΈ Development (from source)

```bash
git clone https://github.com/wirux/mcp-markdown-vault.git
cd mcp-markdown-vault
npm install
npm run build
VAULT_PATH=/path/to/your/vault node dist/index.js
```

---

## 🌐 Transport Modes

| Mode | Use case | How it works |
|---|---|---|
| πŸ“‘ `stdio` *(default)* | Single-client desktop apps (Claude Desktop) | Reads/writes stdin/stdout; 1:1 connection |
| 🌊 `sse` | Multi-client setups (Docker, Claude Code) | HTTP server with SSE streams; one connection per client |

**SSE** starts an HTTP server on `PORT` (default `3000`):

- `GET /sse` β€” establishes an SSE stream (one per client)
- `POST /messages?sessionId=...` β€” receives JSON-RPC messages

```bash
MCP_TRANSPORT_TYPE=sse PORT=3000 VAULT_PATH=/path/to/vault npx @wirux/mcp-markdown-vault
```

Each SSE client gets its own workflow state. Shared resources (vault, vector index, embedder) are reused across all connections.

---

## 🧠 Embedding Providers

The server selects an embedding provider automatically:

| `OLLAMA_URL` set? | Ollama reachable? | Provider used |
|---|---|---|
| ❌ No | β€” | 🏠 Local (`@huggingface/transformers`, `all-MiniLM-L6-v2`, 384d) |
| βœ… Yes | βœ… Yes | πŸ¦™ Ollama (`nomic-embed-text`, 768d) |
| βœ… Yes | ❌ No | 🏠 Local *(fallback with warning)* |

> No configuration needed for local embeddings β€” the model downloads on first use and is cached automatically.

---

## βš™οΈ Configuration

| Variable | Default | Description |
|---|---|---|
| `VAULT_PATH` | `/vault` | Markdown vault directory |
| `VAULT_CONTEXT_MODE` | `assisted` | Vault orientation mode: `assisted` (host LLM/agent calls `prepare_overview` to gather evidence, then generates prose and calls `save_overview`) or `manual` (you author `meta/overview.md` yourself and the server does not overwrite it). `auto` is a deprecated alias for `assisted`. |
| `VAULT_CONTEXT` | *(deprecated)* | Deprecated and ignored. Use `VAULT_CONTEXT_MODE` instead. |
| `MCP_TRANSPORT_TYPE` | `stdio` | `stdio` (single client) or `sse` (multi-client HTTP) |
| `PORT` | `3000` | HTTP port (SSE mode only) |
| `OLLAMA_URL` | *(unset)* | Set to enable Ollama embeddings |
| `OLLAMA_MODEL` | `nomic-embed-text` | Ollama embedding model name |
| `OLLAMA_DIMENSIONS` | `768` | Ollama embedding vector dimensions |
| `VECTOR_STORE_URL` | *(unset)* | Set to use Qdrant (e.g. `http://localhost:6333`). If unset, local persisted flat store is used. |
| `VECTOR_STORE_COLLECTION` | `markdown_vault` | Qdrant collection name when `VECTOR_STORE_URL` is set. |
| `VECTOR_STORE_RESET` | `false` | Set to `true` to auto-delete a mismatched vector index on startup and rebuild from scratch. |
| `MCP_AUTH_TOKEN` | *(unset)* | Bearer token for SSE transport auth. If set, all SSE endpoints require `Authorization: Bearer <token>`. |
| `HOST_BIND_ADDRESS` | `127.0.0.1` | Bind address for the SSE HTTP server. |
| `BODY_LIMIT_BYTES` | `1mb` | Max JSON request body size for SSE `POST /messages`. |

> **Note**: When using the default local vector store, a `.markdown_vault_mcp` directory will be created in your vault. It's recommended to add this directory to your `.gitignore`.

Use `assisted` mode when you want the connected host LLM/agent to generate and refresh vault context from server-provided evidence. Use `manual` mode when you want to write and maintain `meta/overview.md` yourself; in manual mode, the server creates the file if missing but does not overwrite it.

---

## πŸ—οΈ Architecture

Clean Architecture with strict layer separation:

```
src/
β”œβ”€β”€ domain/           πŸ”· Errors, interfaces (ports), value objects
β”œβ”€β”€ use-cases/        πŸ”Ά Business logic (AST, chunking, search, workflow)
β”œβ”€β”€ infrastructure/   🟒 Adapters (file system, Ollama, vector store)
└── presentation/     🟣 MCP tool bindings, transport layer (stdio/SSE)
```

See [CLAUDE.md](CLAUDE.md) for detailed architecture docs and [CHANGELOG.md](CHANGELOG.md) for implementation history.

---

## πŸ—ΊοΈ Self-Orienting Context Layer

Connected agents automatically discover **when** to query this vault and **how** to use its tools β€” no explicit user instructions needed.

> **Quick start:** Run the **`rebuild-overview`** MCP prompt after adding notes to your vault. This generates context that helps agents route queries to the right vault.

### How it works

The server delivers vault context through multiple mechanisms (graceful degradation across clients):

| Mechanism | When | What the agent sees |
|---|---|---|
| `instructions` field | MCP handshake | `vault_scope` + tool summary |
| MCP Resources | On-demand | `vault://overview` (stats + overview + conventions) |
| First-call priming | First tool call per session | `vault_scope` + hint to read `vault://overview` |
| Tool descriptions | Tool listing | `vault_scope` string for routing |

### Modes

| Mode | How overview is managed |
|---|---|
| `assisted` *(default)* | Host agent calls `system.prepare_overview` β†’ generates prose β†’ calls `system.save_overview` |
| `manual` | You author `meta/overview.md` yourself; server creates stub but never overwrites |

To rebuild in assisted mode: invoke the **`rebuild-overview`** MCP prompt, or ask your agent to call `system.prepare_overview` then `system.save_overview`.

### Vault meta files

On first startup, the server creates two files in `<VAULT_PATH>/meta/`:

| File | Purpose | Managed by |
|---|---|---|
| `meta/overview.md` | Vault description + `vault_scope` routing hint | Host agent (assisted) or you (manual) |
| `meta/contract.md` | Tool usage conventions (frontmatter schema, search hints, naming) | Created once, never overwritten |

> **Tip:** Keep `vault_scope` short and specific β€” it tells MCP hosts what information this vault can answer.


---

## 🚒 CI/CD & Release

Fully automated via GitHub Actions and [Semantic Release](https://semantic-release.gitbook.io/):

| Workflow | Trigger | What it does |
|---|---|---|
| **PR Check** | Pull request to `main` | Lint β†’ Build β†’ Test |
| **Release** | Push to `main` | Lint β†’ Test β†’ Semantic Release (NPM + GitHub Release) β†’ Docker build & push to `ghcr.io` |

- Versioning follows [Conventional Commits](https://www.conventionalcommits.org/) β€” `feat:` = minor, `fix:` = patch, `feat!:` / `BREAKING CHANGE:` = major
- Docker images are built for `linux/amd64` and `linux/arm64` via QEMU
- NPM package published as [`@wirux/mcp-markdown-vault`](https://www.npmjs.com/package/@wirux/mcp-markdown-vault)
- Docker image available at [`ghcr.io/wirux/mcp-markdown-vault`](https://github.com/wirux/mcp-markdown-vault/pkgs/container/mcp-markdown-vault)

---

## πŸ§ͺ Testing

**568 tests** across 49 files, written test-first (TDD).

```bash
npm test                                          # Run all tests
npx vitest run src/use-cases/ast-patcher.test.ts  # Single file
npm run test:watch                                # Watch mode
npm run test:coverage                             # Coverage report
```

> Tests use real temp directories for file system operations and in-memory MCP transport for integration tests. No external services required.

---

## πŸ”’ Security

- πŸ›‘οΈ All file paths validated through `SafePath` value object before any I/O
- 🚫 Blocks path traversal: `../`, URL-encoded (`%2e%2e`), double-encoded (`%252e`), backslash, null bytes
- ✍️ Atomic file writes (temp file + rename) prevent partial writes
- πŸ‘€ Docker container runs as non-root user

---

## πŸ“„ License

[MIT](LICENSE)

TDQS

B3.4/5.0

Scored across 5 tools

Disambiguation2/5

The view and vault tools both expose read operations (view.read and vault.read) with overlapping semantics, and the edit tool overlaps with vault.update in that both can write content to notes. The system and workflow tools also have a 'status' action each, which could be confused despite different scopes. An agent must read descriptions carefully to choose correctly, and the boundaries remain fuzzy for common tasks like reading or updating a note.

Naming Consistency4/5

All tool names use a consistent snake_case pattern, and the internal actions are consistently named with single lowercase words. The only slight inconsistency is that there is no shared verb_noun pattern across toolsβ€”they are domain-oriented nouns with action parametersβ€”but within this design, naming is predictable.

Tool Count5/5

Five tools for a markdown vault server is a well-scoped, compact surface. Each tool covers a distinct domain (system, view, edit, vault, workflow), and no tool appears redundant in count.

Completeness5/5

The toolset provides full CRUD (vault.create/read/update/delete), search, backlinks, frontmatter handling, templates, system administration, and workflow management. This covers the lifecycle of markdown notes and vault operations without obvious gaps.

Maintenance

ActivityMaintained
ResponsivenessWithin a week