Skip to main content
Glama
README.md
# brain-mcp

[![MIT License](https://img.shields.io/badge/license-MIT-blue.svg)](LICENSE)
[![Python 3.12+](https://img.shields.io/badge/python-3.12%2B-blue.svg)](pyproject.toml)

A "personal brain" MCP server backed by an Obsidian vault: the agent can search,
read, create, and update notes, navigate the backlink/tag graph, and consult
daily notes. The vault is the only source of truth; the server maintains only
derived indexes (FTS5 + vector) outside the vault.

## Quickstart

Requires Python 3.12 or later and uv. From a repository checkout:

```bash
git clone https://github.com/alefiengo/brain-mcp.git
cd brain-mcp
uv sync --locked
```

To install independently of the checkout, build and install the wheel:

```bash
uv build --wheel
uv tool install dist/brain_mcp-0.1.0-py3-none-any.whl
```

With that installation, use `brain-mcp` directly in place of `uv run --locked brain-mcp`.
In each project, create `brain.toml` from [brain.example.toml](brain.example.toml):

```toml
[brain]
vault = "/path/to/this-project-vault"
```

Check the configuration before connecting the client:

```bash
uv run --locked brain-mcp --project /path/to/project --check
uv run --locked brain-mcp --project /path/to/project
```

The second command serves stdio and waits for an MCP client. `--check` prints
resolved paths without opening indexes or starting a watcher.

## Tools

| Tool | Description |
| --- | --- |
| `search(query)` | Hybrid search: FTS5 (literal) + local embeddings (semantic), with a score and snippet. |
| `read_note(path)` | Reads complete frontmatter, body, and file revision. |
| `create_note(path, content, tags)` | Creates a note with automatic frontmatter (`title`, `created`, `updated`, `tags`). |
| `update_note(path, text, heading, expected_revision)` | Appends text or replaces a section with a backup; accepts the revision read earlier to reject stale edits. |
| `list_recent(limit)` | Most recently modified notes. |
| `get_backlinks(path)` | Notes linking to a note (wiki-links `[[...]]` and links `[text](x.md)`). |
| `get_tags()` | Tags aggregated by frequency. |
| `daily_summary(date)` | Daily summary: daily note + tasks `- [ ]` / `- [x]`. |

Tool descriptions, errors, confirmations, status labels, and daily task headings
are in English. Tool names, arguments, and resource URIs remain unchanged.

## Project configuration

`--project DIR` loads **only** `DIR/brain.toml`. It does not search parent
directories or use inherited `VAULT_PATH` or `BRAIN_INDEX_DIR` values. If the
declaration is missing or invalid, startup fails: it never silently selects
another vault.

Relative `vault` and `index_dir` paths are resolved from the project, regardless
of the server's cwd. `vault` is required. `index_dir` is optional and must stay
outside the vault; the default is
`<project>/.brain/projects/<SHA-256 of project and vault>`.
For two simultaneous clients of the same project, pass a different explicit
`--index-dir` in each client's command, for example
`--index-dir /path/to/project/.brain/my-project-codex` and another ending in
`-opencode`. This argument overrides only the index directory, never the vault.
The lease prevents active processes from sharing an index.

`brain://status` shows the effective project, vault, and indexes. Check it before
writing. The association is fixed for the lifetime of the process: changing
`brain.toml` requires restarting the server. `--project` is an explicit selection;
it does not automatically detect which repository is open in the client's interface.

Without `--project`, legacy mode remains available: `VAULT_PATH` is required,
and `BRAIN_INDEX_DIR` is optional (default `<cwd>/.brain`). Do not install this
mode as a global MCP server when working with different projects.

## MCP client integration

Register a stdio server in the client's project configuration. Adjust the
absolute paths and project name in this command and arguments example for
an installation from a wheel:

```json
{
  "command": "/path/to/bin/brain-mcp",
  "args": ["--project", "/path/to/project"]
}
```

When running from the checkout, the equivalent command is:

```bash
uv run --locked --directory /path/to/brain-mcp brain-mcp --project /path/to/project
```

The registration schema depends on the MCP client. Use an absolute path for
`--project`, restart the server after configuration changes, and check
`brain://status` before writing. Each concurrent client needs its own
`--index-dir`. The checkout directory does not identify the user's project.

## Reproducible demo

The included fixtures are synthetic. Copy the vault to test writes without
modifying test files; run from the checkout:

```bash
mkdir -p /tmp/brain-mcp-demo
cp -R tests/fixtures/vault /tmp/brain-mcp-demo/vault
VAULT_PATH=/tmp/brain-mcp-demo/vault BRAIN_INDEX_DIR=/tmp/brain-mcp-demo/.brain uv run --locked brain-mcp --check
```

Remove `--check` to serve stdio. Example notes include frontmatter, wiki-links,
tags, and a daily note. You can organize your own vault freely;
`daily_summary` looks for `daily/YYYY-MM-DD.md`.

## Development

```bash
uv run ruff check --fix   # lint (E,W,F,I,UP,B,SIM, line-length 100)
uv run ruff format        # format
uv run pytest             # tests + coverage (minimum 90%)
```

Enable the hook in each checkout with `git config core.hooksPath .githooks`.
The hook runs all three steps on each commit. See [CONTRIBUTING.md](CONTRIBUTING.md),
the [architecture](docs/architecture.md), and the [roadmap](docs/roadmap.md).

## Repository structure

- `src/brain_mcp/`: server, tools, vault access, and indexes.
- `tests/`: regression tests and synthetic vaults in `tests/fixtures/`.
- `docs/`: architecture, roadmap, and [publishing guide](docs/publishing.md).
- `.github/`: CI, Dependabot, issue forms, and pull request template.
- `.githooks/`: optional development hook.
- `brain.example.toml`: configuration to copy into each project.
- `AGENTS.md`: instructions for agents and contributors.
- `CLAUDE.md`: Claude entry point for the shared instructions in `AGENTS.md`.
- `GEMINI.md`: Gemini entry point for the shared instructions in `AGENTS.md`.
- `pyproject.toml` and `uv.lock`: metadata, dependencies, and quality configuration.
- `LICENSE`, `SECURITY.md`, and `CONTRIBUTING.md`: license and project policies.

`.brain/` contains local indexes and backups and is excluded from Git.

## Architecture

- Stdio transport (FastMCP). No HTTP.
- Indexes (FTS5 in `<index_dir>/fts.db`, vector with sqlite-vec in
  `<index_dir>/index.db`) are derived and can be reindexed on demand
  (incrementally using a SHA-256 content digest). Deleting them does not lose vault data.
- **Watcher**: at startup, the server watches the vault and reindexes in the
  background whenever `.md` files change, so searches reflect modifications
  without manual reindexing. A periodic scan recovers missed events and retries
  transient errors.
- Local embeddings with `fastembed`
  (`paraphrase-multilingual-MiniLM-L12-v2`, 384 dimensions); no network calls at
  runtime (except the initial model download).
- Exact backups of destructive writes: `<BRAIN_INDEX_DIR>/backup/`.
- `Config` validates paths before opening resources; `BrainRuntime` belongs to
  one MCP instance and closes the watcher/connections even on failure.
- `brain://status` reports whether the watcher is active and whether semantic
  search is ready, unloaded, or degraded.

## Operation and recovery

The server starts with FTS5 without loading the embedding model. The first
search attempts to load it and may download it if it is not cached. If the
model fails, the server retains available literal results and logs the cause
to stderr; if there are no literal results either, it returns a ToolError
explaining the limitation. A later search retries semantic search.
Logs go to stderr to preserve the MCP protocol on stdout.

Each process leases `BRAIN_INDEX_DIR` through a SQLite transaction in
`session.db`. A second process using that directory fails at startup; use a
different directory for another vault. The lease is released on shutdown or
process termination. Indexes store the vault's identity and reject access
from old instances after a database is reassigned.

To rebuild indexes, stop the server, move `fts.db` and `index.db` to another
directory, and restart it. FTS5 rebuilds at startup; the vector index rebuilds
on the first search. Preserve `backup/`: it contains the original files,
including YAML comments and line endings. To restore a note, stop the server,
identify its backup, and copy that file to the corresponding vault path;
the indexes will be updated after restarting.

Notes are published in full using a temporary file in the same directory,
fsync, and atomic replacement. Creation uses exclusive publication and does
not overwrite a path occupied by another creator. Updates with invalid YAML
are rejected. Headings inside code blocks are not treated as editable sections.
YAML in updated notes is serialized again: it preserves keys and values but
normalizes formatting and comments.

To edit content you have just read, pass the `expected_revision` included in
`read_note` to `update_note`. If someone has modified the note since that read,
the server rejects the edit and asks you to read it again. It also compares
the bytes immediately before publishing an update.

## Explicit limits

- Tool and watcher concurrency is serialized per instance. Comparison with
  external editors is optimistic: no filesystem operation atomically compares
  and replaces bytes. An editor writing between the final comparison and
  replacement can still race. Keep a single writer process per vault and read
  again after a conflict.
- Durability and exclusive publication require a local filesystem supporting
  hard links and atomic replacement; directory fsync is used on POSIX. Failure
  guarantees are tested on Linux. Behavior on network drives or with external
  synchronization tools is not guaranteed.
- Semantic search represents the title, tags, and first 1000 characters of the
  body. FTS5 indexes the entire body. There is no chunking of long notes or
  evaluation based on a real vault yet.
- Indexing compares snapshots of each note; the two indexes are updated in
  independent transactions. During external changes, they do not necessarily
  represent the same instant in the vault. The next search or scan converges
  to the current content.
- Each scan reads the content to detect changes even when mtime does not change.
  Cost grows with vault size; there is no scale benchmark.
- Backups are not automatically deleted. Monitor their disk usage.

## Quality verification

CI installs dependencies from `uv.lock`, runs lint, formatting, and tests with
minimum coverage of 90%, and builds the package. It includes tests for write
failures, index rollback, external conflicts, MCP isolation, and resource cleanup.
Semantic tests use the real model; a new environment needs network access to
download it. Resilience tests use controlled embeddings to make injected
failures reproducible.

The retrieval evaluation uses 14 synthetic notes, six paraphrased queries,
and relevance labels. It requires recall@3 of 1.0 and MRR >=0.8.
To view the metrics:

```bash
uv run pytest tests/test_retrieval.py --no-cov -s
```

The corpus is a small regression check, not a measure of general quality.
Add distractor notes and labeled queries before changing the model,
threshold, or ranking fusion.

Vulnerability reports follow [SECURITY.md](SECURITY.md).

## License

[MIT](LICENSE). The license covers this code; dependencies and embedding models
retain their own licenses.

Maintenance

ActivityMaintained
ResponsivenessNo issues