Skip to main content
Glama
README.md
```
 ██████╗ ██████╗ ███████╗██╗██████╗ ██╗ █████╗ ███╗   ██╗
██╔═══██╗██╔══██╗██╔════╝██║██╔══██╗██║██╔══██╗████╗  ██║
██║   ██║██████╔╝███████╗██║██║  ██║██║███████║██╔██╗ ██║
██║   ██║██╔══██╗╚════██║██║██║  ██║██║██╔══██║██║╚██╗██║
╚██████╔╝██████╔╝███████║██║██████╔╝██║██║  ██║██║ ╚████║
 ╚═════╝ ╚═════╝ ╚══════╝╚═╝╚═════╝ ╚═╝╚═╝  ╚═╝╚═╝  ╚═══╝

██╗      ██████╗  ██████╗ █████╗ ██╗         ███╗   ███╗ ██████╗██████╗
██║     ██╔═══██╗██╔════╝██╔══██╗██║         ████╗ ████║██╔════╝██╔══██╗
██║     ██║   ██║██║     ███████║██║         ██╔████╔██║██║     ██████╔╝
██║     ██║   ██║██║     ██╔══██║██║         ██║╚██╔╝██║██║     ██╔═══╝
███████╗╚██████╔╝╚██████╗██║  ██║███████╗    ██║ ╚═╝ ██║╚██████╗██║
╚══════╝ ╚═════╝  ╚═════╝╚═╝  ╚═╝╚══════╝    ╚═╝     ╚═╝ ╚═════╝╚═╝
```

Gives your AI coding assistant a **memory of your codebase**. Instead of
grepping around blindly every session, it asks a local graph "what does this
touch, what breaks if I change it, and what did we already learn about it?"

No API keys, no network calls, no embeddings. Everything runs on your machine.

![Architecture](img/diagram.png)

---

## What it does, in plain words

You ask your assistant to change something. Without this server, it searches
for text and hopes for the best. With it:

```
You:  "the daily podcast is generating twice, fix it"

Assistant calls retrieve_context("daily podcast cron") and gets back:
  MUST_READ    the files that actually implement it
  MAY_CHANGE   what depends on them and would break
  MUST_TEST    the tests that cover it
  KNOWLEDGE    your note from three weeks ago explaining why the cron
               runs at 540s and what already went wrong once
  DANGER       anything flagged as risky to touch
```

Two things make that possible:

1. **A graph of your code**, built mechanically by reading your files
   (Python `ast`, TypeScript tree-sitter) and your git history. Files,
   symbols, who imports whom, who calls whom, what gets edited together.
2. **An Obsidian vault of notes** — plain `.md` files where you (or the
   assistant) write down the *why*: decisions, incidents, gotchas. Things
   that are nowhere in the code.

The code half gets rebuilt whenever you reindex. The notes half is yours and
never gets overwritten.

---

## What you need

- **[uv](https://docs.astral.sh/uv/)** — handles Python for you (it downloads
  Python 3.13 automatically; that version is pinned because
  `tree-sitter-typescript` wheels lag newer ones).
- **A folder for your vault.** That is all a vault is: a folder of `.md`
  files. Installing [Obsidian](https://obsidian.md) is optional and only
  needed if you want to *browse* your notes and see the graph visually.
- **Git repos** you want indexed. Only files tracked by git are scanned.

Optional: a GPU or Apple Silicon, if you want the
[decision model](#the-decision-model-optional). It is too slow on a CPU.

---

## Setup

### 1. Install

```bash
git clone <this repo>
cd obsidian-local-mcp
uv sync
```

### 2. Set up your vault

```bash
uv run obsidian-local-mcp-setup
```

It asks two questions — whether to install the decision model (say no for
now) and where your vault is — then writes
`~/.obsidian-local-mcp/vaults.yaml` and prints the block for step 3.

<details>
<summary>Or write the config yourself</summary>

Copy `vaults.example.yaml` to somewhere **outside this repo** — it contains
real paths from your machine, so it must never be committed:

```yaml
default_vault: main

vaults:
  main:
    path: /path/to/your/vault-folder
    conventions: generic          # "generic" = no rules about note format
    repos:
      frontend:
        path: /path/to/your/frontend
        include: ["src/**"]
        exclude: ["**/node_modules/**"]
        languages: [ts, tsx]
      backend:
        path: /path/to/your/backend
        include: ["app/**", "tests/**"]
        languages: [py]
```

`repos` names (`frontend`, `backend`) are yours to pick — you will use them
when indexing. Use `conventions: englora` instead of `generic` if you want
strict note contracts enforced (see [Vault profiles](#vault-profiles)).

</details>

### 3. Connect it to your assistant

Create a `.mcp.json` in each repo you work from:

```json
{
  "mcpServers": {
    "obsidian": {
      "command": "uv",
      "args": ["run", "--directory", "/path/to/obsidian-local-mcp", "obsidian-local-mcp"],
      "env": {
        "OBSIDIAN_MCP_CONFIG": "/path/to/your/vaults.yaml",
        "OBSIDIAN_MCP_VAULT": "main"
      }
    }
  }
}
```

Restart your assistant. In Claude Code, `/mcp` should now list `obsidian`.

### 4. Build the graph (once)

Just ask your assistant, in plain language:

> "index the frontend and backend repos, then index the git co-edits"

Or, if you prefer to know what it runs under the hood:

```
index_repo(repo="frontend")
index_repo(repo="backend")
index_git_coedits(repo="frontend")
index_git_coedits(repo="backend")
```

Then ask for `graph_stats()` to confirm. A few thousand nodes for a
mid-sized repo is normal. This takes seconds, not minutes.

---

## The CLI

One command, three subcommands. Run them from this repo with `uv run`.

```bash
uv run obsidian-local-mcp-setup           # set up your vault (and optionally the model)
uv run obsidian-local-mcp-setup serve     # run the decision model, if you installed one
uv run obsidian-local-mcp-setup doctor    # check everything and print the triage latency
```

`doctor` is the one to run when something looks wrong — it checks the
config, your vaults, whether the graph is built and whether the model
answers, and says what to do about anything it finds.

To set up without the questions:

```bash
uv run obsidian-local-mcp-setup install --vault ~/notes --no-model --force
```

| `install` flag | |
| --- | --- |
| `--vault PATH` | use this vault, skip all prompts |
| `--model {0.8b,4b,9b}` | install this model without asking |
| `--no-model` | skip the model |
| `--force` | overwrite an existing config |
| `--port N` | port for the model (default 8009) |

Config lives at `~/.obsidian-local-mcp/vaults.yaml` and the model in
`~/.obsidian-local-mcp/kev-env/`. `OBSIDIAN_MCP_CONFIG` and
`OBSIDIAN_MCP_HOME` move them.

---

## The decision model (optional)

Skip this unless you have a GPU or Apple Silicon.

`retrieve_context` ranks candidates by how your code is wired together,
which cannot tell that a well-connected file is irrelevant *to the question
you asked*. Your assistant finds that out by opening them — and every one
that did not matter is now in its context for good.

`triage_candidates(query)` runs the same search, then asks
[Kev](https://github.com/jaredpalmer/kev), a small model running on your
machine, which candidates are worth reading. You get those in full and the
rest as one-line stubs.

```bash
uv run obsidian-local-mcp-setup install --model 0.8b   # once
uv run obsidian-local-mcp-setup serve                  # leave running
uv run obsidian-local-mcp-setup doctor                 # check the latency
```

Tune it in `vaults.yaml` under `kev:` (`enabled`, `base_url`, `timeout_ms`,
`threshold`, `max_candidates` — see `vaults.example.yaml`).

**It cannot break anything.** If the model is down, slow or disabled,
`triage_candidates` returns every candidate with `triaged: false` and a
`reason` — exactly what `retrieve_context` would have given you. Never an
error. So it is safe to leave configured, and safe to kill `serve`.

**But check the latency.** Triage has to cost less than the reading it
saves. We measured kev-4b scoring 15 candidates on a CPU with no GPU:
**~73 seconds**. Unusable — it would just time out and fail open every
time. `0.8b` is ~5x smaller and still probably too slow. Without a GPU,
leave it off and use `retrieve_context`.

The judgement itself is good, which was the part in doubt: on a Spanish
query against Spanish notes, kev-4b kept exactly the five candidates that
mattered and dropped eight structural neighbours and an unrelated billing
module at 0.00–0.07.

---

## Using it day to day

**You do not call these tools yourself.** You work exactly as you always
did — the assistant decides when to reach for the graph, because every tool
describes when it should be used.

```
You:  "why does the newsletter send the wrong level sometimes?"
      → assistant calls retrieve_context, reads your notes, then the code
```

Three things worth knowing:

**Ask it to use the graph when it forgets.** A nudge works ("check the graph
first"). To make it permanent, put one line in your `CLAUDE.md`:

> Before using Grep/Glob to understand code or relationships, call
> `retrieve_context` first. Grep is only for locating a known literal string.

**Reindex after you move things.** Adding, deleting, renaming files or
changing imports makes the graph drift. A stale map is worse than no map:

> "reindex the files I just changed"  →  `reindex_paths(repo, paths=[...])`

**Let it write notes.** When you explain something non-obvious ("we cap it
at 540s because Cloud Run was OOMing"), tell it to save that:

> "write that down in the vault"  →  `write_note(...)`

That note is what shows up in `KNOWLEDGE` next month, long after you and the
assistant have both forgotten the conversation.

---

## Searching well

Seeding is lexical (SQLite FTS5 + bm25), **not** semantic. So queries work
best with few, distinctive words:

| Query                                        | Result                        |
| -------------------------------------------- | ----------------------------- |
| `"podcast"`                                  | excellent                     |
| `"daily podcast cron"`                       | good                          |
| `"how does the daily podcast work and what generates it"` | poor — common words match everything |

If the buckets come back unrelated to what you asked, retry with fewer and
rarer words (a file name, a feature name, a symbol).

---

## The two kinds of data

This distinction matters, and it is the one thing to get right:

|                | **Code nodes** (mechanical)                       | **Curated notes** (yours)                    |
| -------------- | ------------------------------------------------- | -------------------------------------------- |
| Looks like     | `file:back/app/cron.py`, `fn:back/app/cron.py#run` | `feature:daily-podcast`, `note:...`          |
| Written by     | `index_repo`, `index_git_coedits`                  | `write_note`, `upsert_node`, `upsert_edge`   |
| To change it   | re-run the indexer                                 | edit the note                                |
| Never          | hand-edit it — the indexer will overwrite you       | gets touched by an indexer                   |

For a note to surface as `KNOWLEDGE` when someone touches a file, it needs a
`documents` edge pointing at that file:

```
upsert_edge(src="feature:daily-podcast", dst="file:back/app/podcast.py", type="documents")
```

Without that edge the note only appears when your query happens to match its
text. With it, the note appears whenever that file is in play — which is the
whole point.

---

## Troubleshooting

**Changed the server's code and nothing happened.** The running process still
has the old code loaded. Restart the MCP server (in Claude Code: restart the
session, or reconnect via `/mcp`).

**`KNOWLEDGE` is always empty.** Your notes probably have no `documents`
edges yet. See [The two kinds of data](#the-two-kinds-of-data).

**Results point at files that no longer exist.** The graph drifted. Run
`detect_divergence()` to see what broke, then `reindex_paths` or a full
`index_repo`.

**A note says something that is no longer true.** Notes carry `status` and
`confidence` in their frontmatter. `status: "assumed"` means a past session
inferred it without verifying — always confirm against the code before
trusting it. `consolidation_report()` lists stale and orphaned notes.

**Nothing gets indexed.** Only git-tracked files are scanned. Check your
`include`/`exclude` globs and that the files are actually committed.

**Anything looks wrong.** Run `obsidian-local-mcp-setup doctor` first. It
checks the config, your vaults, the graph and the model, and says what to
do about whatever it finds.

**`triage_candidates` always says `triaged: false`.** The model is not
answering; the `reason` field says why. Your results are unaffected —
`kept` holds every candidate.

---

## Vault profiles

Set per-vault with `conventions:`:

- **`generic`** — no rules. Notes are plain markdown; write whatever you
  want. Good for a personal wiki.
- **`englora`** — enforces a note contract (`id`, `type`, `area`, `status`,
  `confidence`, `source`, `created`, `updated`, `decay`, `supersedes`,
  `links`) and refuses to mark inferred claims as `verified`. Good when you
  want the assistant held to a standard.

Multiple vaults can be configured at once; `OBSIDIAN_MCP_VAULT` picks the
default per repo, and every tool takes an optional `vault=` argument.

---

# Internals

Everything below is for working *on* this server rather than *with* it.

## Architecture

```
src/obsidian_local_mcp/
├── server.py            FastMCP instance + all tools (thin: validate → call → model)
├── instructions.py       server-level instructions handed to the connecting assistant
├── config.py              vaults.yaml → Config/VaultConfig/RepoConfig/KevConfig
├── models.py               pydantic request/response types for every tool
├── setup_cli.py             the obsidian-local-mcp-setup command (install/serve/doctor)
├── triage/
│   ├── client.py            stdlib-only HTTP client for the Kev sidecar
│   └── candidates.py         RetrievalResult → one Kev request → keep/drop verdicts
├── graph/
│   ├── store.py           SQLite (<vault>/.graph/graph.db): nodes, edges, FTS5
│   ├── ids.py               file:<repo>/<path>, fn:<repo>/<path>#<symbol>, feature:, ext:, note:
│   ├── weights.py            edge-type weight / reverse-factor / origin table
│   └── retrieve.py            seed_fts() + personalized_pagerank() + classify()
├── index/
│   ├── scan.py             repo file listing (git ls-files) + parser dispatch + test-edge derivation
│   ├── python_parser.py      ast: contains/call/import, two-pass (forward refs resolve)
│   ├── ts_parser.py           tree-sitter: contains/call/import + JSX-usage-as-call
│   └── git_coedit.py         co_edit edges from `git log --name-only`
└── vault/
    ├── conventions.py       vault "profiles": englora (strict) vs generic (unenforced)
    ├── notes.py               frontmatter read/write, wikilink-tolerant YAML, secret guard
    └── maintenance.py         decay/orphan/divergence reports, additive sync_note_links
```

### The origin rule

Every node/edge carries `origin: scanner | git | manual`. Reindexing only
ever rewrites rows it owns:

- `index_repo`/`reindex_paths` delete+rewrite `scanner`-origin rows for the
  paths touched, then purge any `scanner`-origin edge left dangling.
- `index_git_coedits` delete+rewrite `git`-origin edges for the repo.
- `manual`-origin rows (written by `upsert_node`/`upsert_edge` — curated
  `feature:`/`ext:` nodes, `depends_on`/`documents`/`decided_by`/`dangerous`
  edges) are never touched by either indexer.

This is what lets an assistant curate knowledge on top of a graph that gets
mechanically rebuilt.

### Retrieval

`retrieve_context(query)`: FTS5 bm25 over node id/label/path/curated-note
body seeds a personalized-PageRank walk (restart α=0.15, both edge
directions at their type's `weight`/`reverse_factor`). The ranked result is
then classified into MUST_READ / MAY_CHANGE / MUST_TEST / KNOWLEDGE /
DANGER by walking the same edges. `impact_of(node_id)` runs the same
classification anchored on one explicit node instead of a query.

`documents`/`decided_by` edges run note → code, so the note explaining an
anchor is reached by walking *into* the anchor (`by_dst`), not out of it.
Those notes are surfaced even with no lexical hit of their own — curated
prose exists precisely for the queries whose wording does not match it.

No embeddings by default — `seed_fts()` is the one seam where a vector
search could be substituted later without touching the ranking/
classification code.

### Triage

`retrieve_context` ranks by topology. Topology cannot tell that a
well-connected file is irrelevant *to this particular question*, so the
assistant settles that by opening the candidates — and every one that turns
out not to matter has already been paid for and is in its context for good.

`triage_candidates(query)` runs the same search, then asks
[Kev](https://github.com/jaredpalmer/kev) — a small local decision model —
which candidates are worth reading. Kev is a *pointer* model: one `state`
plus N typed questions go in, one prefill pass runs, and the answers are
read off the logits. Nothing is generated. Questions share the state's
activations and cannot see each other, so 25 candidates cost about what one
costs and no candidate's verdict is dragged around by its neighbours'.

Consequences that shaped the code, all in `triage/candidates.py`:

- **One `Noul` per candidate, never one big `Choice`.** Choice options are
  read in order and the ranking moves when you permute them; independent
  yes/no questions have no order to be biased by, and they let several
  candidates be essential at once.
- **`must_read` is never dropped.** The graph put it there off the query's
  own seeds; a model wrong on that one would hide the file you came for.
- **A missing or malformed answer keeps the candidate.** Silence must not
  be indistinguishable from rejection.
- **Fail-open, always.** Kev down, slow, disabled or throwing returns every
  candidate with `triaged: false` and a `reason`. A filter that can hide
  results is never allowed to be the reason a lookup fails.

Kev is reached over HTTP and never imported: it needs torch and a multi-GB
checkpoint, and this server is loaded by a stdio client that would
otherwise wait on that at every start. That is why the sidecar lives in its
own virtualenv and why `setup_cli.py` exists at all.

The user-facing half — installing it, running it, configuring it, and the
measured latency that decides whether it is worth running — is in
[The decision model](#the-decision-model-optional).

## Tools

Graph write: `upsert_node[s]`, `upsert_edge[s]`, `delete_node`,
`delete_edge`, `rename_node`, `set_node_status`.

Graph read: `get_node`, `neighbors`, `search_nodes`, `retrieve_context`,
`triage_candidates`, `impact_of`, `path_between`, `graph_stats`.

Indexing: `index_repo`, `reindex_paths`, `index_git_coedits`,
`detect_divergence`, `clear_graph` (requires `confirm=True`).

Vault notes: `read_note`, `write_note`, `update_note_frontmatter`,
`promote_node_to_note` (scaffold graph nodes into vault notes),
`list_notes`, `search_vault`, `archive_note`.

Maintenance: `consolidation_report`, `sync_note_links`, `get_status`.

## Tests

```bash
uv run pytest
```

All 116 tests pass (one skips unless kev is importable). Key fixtures:

- `tests/fixtures/mini_repo/` — a small backend (5 Python files) that
  reproduces a complete indexing workflow: `index_repo` →
  `index_git_coedits` → `retrieve_context`.
- `tests/fixtures/mini_vault/` — a minimal vault used to test note
  read/write, frontmatter validation, and `promote_node_to_note`.
- Dynamic `mini_repo_git` fixture — builds a throwaway git repo with two
  commits to test co-edit detection.

TDQS

B3.1/5.0

Scored across 30 tools

Disambiguation4/5

The set is mostly well-differentiated, with clear splits between node/edge CRUD, note operations, indexing, and context queries. Minor overlaps remain, especially between search_vault and search_nodes, and between consolidation_report and detect_divergence, but descriptions usually clarify the intended choice.

Naming Consistency4/5

All tool names use snake_case consistently, and most follow recognizable verb_noun or verb_object patterns. A few tools are noun-only or prepositional phrases (neighbors, graph_stats, path_between, impact_of), which is a minor deviation from a strict verb_noun convention.

Tool Count2/5

With 30 tools, the surface is above the 25+ threshold that typically indicates an overloaded set. The domain is complex, but the many granular single/batch and lifecycle operations could likely be consolidated or grouped more tightly.

Completeness4/5

The set covers node/edge lifecycle, note lifecycle, indexing, retrieval, impact analysis, pathfinding, stats, divergence detection, and consolidation. Gaps are minor: no batch delete, no explicit edge listing/search, no note rename, and no wikilink removal beyond additive sync.

Maintenance

ActivityMaintained
ResponsivenessNo issues