localmind
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@localmindremember that I prefer dark mode in every editor"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
localmind
Local-first semantic memory for AI agents — now at codebase scale.
localmind is a Model Context Protocol (MCP) server that gives your AI agent persistent memory that never leaves your machine. Memories live in a single SQLite file and are searched with embeddings produced entirely on-device — no API keys, no accounts, no cloud, no telemetry.
v2 adds what you need for whole-codebase memory: an approximate-nearest-neighbour index instead of a linear scan, named project namespaces, an ingest_codebase tool that splits source into functions and classes, incremental re-ingestion that skips unchanged code, and token-budgeted recall so an agent pulls in only what fits.
What it does
Your agent gets 6 MCP tools:
Tool | What it does |
| Store a note, fact or decision (optionally tagged) in a project. |
| Semantic search over one project via HNSW. Optional |
| Delete one memory by id, inside one project. |
| Browse a project, filterable by tag and/or kind ( |
| Index a source tree: walk → chunk at function/class boundaries → embed. Incremental. |
| List namespaces with counts and index sizes. |
Related MCP server: Memory MCP
Project namespacing
Every memory belongs to exactly one project (e.g. pixora, localmind). Isolation is structural, not a filter: each project gets its own vector index file, so a recall in project A physically cannot reach project B's vectors.
{ "project": "pixora", "query": "why did we pick Kafka here?" } // only pixora
{ "query": "what must I avoid cooking?" } // implicit project "default"Omitting project means the default namespace, which is where v1 memories land after migration — so existing behaviour is preserved exactly.
Ingesting a codebase
ingest_codebase(path=<absolute path>, project=<name>, …)
Walks the directory recursively, honouring
.gitignore(comments, negations, nested files,**globs) and pruning excluded directories instead of descending into them.node_modulesand.gitare always skipped.Chunks by logical unit: Go (
func, receiver methods,type … struct/interface), TypeScript/JavaScript (functions, classes, arrow-function consts, interfaces, enums), and Python (def,async def,class, with decorators attached). A brace/indentation scanner that strips string literals and comments keeps a{inside a string from breaking a function. Anything else — Rust, Markdown, YAML — falls back to an overlapping line window, so no file type is skipped and no content is lost.Stores per chunk: the source text, project, file path, start/end line, and a SHA-256 of the text. Chunks are embedded as
<path>:<start>-<end> <kind> <name>plus the source, so queries like "the upload handler in the payments service" match.Re-ingestion is incremental. Chunks are keyed by content hash, so a second run only re-embeds what actually changed:
Re-run scenario | Re-embedded |
Nothing changed | 0 chunks |
One function edited in one file | only that function's chunks |
Function moved to another line (same text) | 0 — metadata-only line-range update |
File deleted | its chunks are tombstoned out of the index |
Batches embedding calls and buckets them by length, so padding stays minimal (padding measurably perturbs mean-pooled output) and each model call stays cheap.
Returns a summary: files scanned/skipped, chunks created/updated/unchanged/deleted, embedding time.
Why local-first matters
Privacy by construction — your notes, decisions, and personal facts are embedded and stored on your disk. There is no server to leak them from, because there is no server.
Zero network dependency — the embedding model (all-MiniLM-L6-v2, 384-dim) ships inside the npm package and runs via ONNX in-process, and the vector index is a local native addon.
npx localmindworks offline.No keys, no quota, no cost — nothing to sign up for, nothing to rate-limit you, nothing to expire.
You own the data — one SQLite file plus one index file per project. Back it up, sync it, grep it, or delete it whenever you like.
Install & run
npx localmindThat's it. First launch embeds and stores straight from the bundled model — no manual setup steps.
npx localmind remember "I prefer dark mode in every editor" --tags prefs
npx localmind recall "what theme should the UI use?" # → [1] (score=0.44…) I prefer dark mode…
npx localmind list
npx localmind forget 1A realistic example: a large multi-service repo
Take PIXORA — 41 microservices in Go, Node and Python, tens of thousands of functions. recall is only useful here if it returns functions, not files, and if ingesting the tree doesn't cost an hour of CPU every time a branch is merged.
Preview —
ingest_codebase { path: "/work/pixora", project: "pixora", dry_run: true }shows how many chunks you are about to create before spending any embedding time.Ingest service by service rather than in one shot, so the agent can start using the index before the whole tree is done:
{ "path": "/work/pixora/services/payments", "project": "pixora" }→{ "path": "/work/pixora/services/search", "project": "pixora" }→{ "path": "/work/pixora/web", "project": "pixora" }Ask with a budget.
{ "project": "pixora", "query": "where do we validate webhook signatures?", "max_tokens": 2000 }returns a handful of function chunks withservices/x/webhook.go:88-134-style references — open those ranges instead of whole files.Re-run after each merge.
{ "path": "/work/pixora", "project": "pixora" }— unchanged functions are skipped by hash, so a 20k-chunk tree re-syncs in seconds (see the benchmark below).Keep decisions in the same namespace:
{ "project": "pixora", "text": "We chose Redis for idempotency keys because the billing exporter assumes at-most-once.", "tags": "decision" }. Notes and code chunks coexist andrecallranks across both.Isolate clients with a second namespace —
{ "project": "pixora-acme" }for a client-specific fork — and be certain a query in one can never surface the other's code.
Benchmark: measured at 21,727 chunks across 2 projects
npm run benchmark generates a synthetic Go/TypeScript/Python monorepo (2,507 files / 9.8 MiB, 23 "services" split into two project namespaces), ingests it, and gates the run on five scale assertions. On an M-series MacBook (10 cores, 16 GiB, local quantized MiniLM):
What | Measured |
Cold ingestion (embed + index) | 21,727 chunks in ~10 min (~36–40 chunks/s, ONNX CPU-bound) |
Recall latency p50 / p95 / max (60 queries, 21.7k vectors) | 6 ms / 11 ms / 12 ms |
Recall with | p50 74 ms / p95 115 ms |
HNSW vs the brute-force cosine scan it replaces | 0.6 ms vs 5.5 ms per query (~9×, plus ~80 ms blob decode the scan also needed) |
Re-ingest after editing one function | 1 of 13,226 chunks re-embedded (0.01%), 359 ms total |
No-op re-ingest (nothing changed) | 195 ms, 0 embedded |
Cold reload of 21,728 vectors from disk | 53 ms |
Rebuild both indexes from SQLite BLOBs | 5.3 s |
Disk | 86 MiB SQLite + 35 MiB index files (5.7 KiB/chunk) |
Peak RSS | ~104 MiB during ingestion; ~409 MiB with both indexes resident |
Cross-project leaks across 60 queries | 0 |
The run fails non-zero unless: corpus ≥ target, recall p95 < 1 s, zero cross-project leaks, incremental re-ingestion < 5% of the corpus, and a no-op re-ingest embeds nothing. Raw numbers land in test/benchmark-results.json.
Cold ingestion is dominated by the local embedding model, not by localmind — and it is paid once per chunk, ever. The steady-state cost an agent actually feels is the re-ingest row and the recall rows, and those stay in milliseconds at PIXORA scale.
Token-budget recall
recall accepts max_tokens. When set, results are ranked, de-duplicated, then packed greedily in rank order until the budget is exhausted — the budget, not a fixed count, decides how much comes back.
Dedup catches three kinds of redundancy before packing: byte-identical chunks (same content hash — the copy-pasted helper that exists in 6 services), overlapping line ranges in the same file, and near-identical text (word-shingle Jaccard ≥ 0.85).
Every hit carries path:start-end, so the caller can open exactly that range and read more on demand. Tokens are estimated at ~4 characters each plus a small per-word allowance, biased high so the limit is a ceiling you can rely on.
# 3 result(s) from project 'pixora' using 203/220 tokens (3 lower-ranked dropped by budget/dedup)
[4] (score=0.5241, services/ledger/ledger.go:3-35) func FlushLedger(…) { … }
[7] (score=0.2033, services/us/tax.go:3-19) func ComputeTax(…) { … }Configure in Claude Code / Claude Desktop
Claude Code
claude mcp add localmind -- npx -y localmindOr in .mcp.json / ~/.claude.json:
{
"mcpServers": {
"localmind": {
"command": "npx",
"args": ["-y", "localmind"]
}
}
}Claude Desktop
Edit claude_desktop_config.json (macOS: ~/Library/Application Support/Claude/claude_desktop_config.json, Windows: %APPDATA%\Claude\claude_desktop_config.json):
{
"mcpServers": {
"localmind": {
"command": "npx",
"args": ["-y", "localmind"]
}
}
}Restart Claude and the remember / recall / forget / list_memories / ingest_codebase tools appear.
MCP client config reference
Field | Value |
command |
|
args |
|
Configuration (optional)
Env var | Default | Purpose |
|
| Where |
| bundled | Directory containing the |
CLI
The same storage and ANN search the MCP server uses, from a terminal:
npx localmind ingest /work/pixora/services/payments --project pixora
npx localmind recall "where do we validate webhook signatures" --project pixora --max-tokens 2000
npx localmind recall "why Kafka and Redis" --project pixora --limit 5
npx localmind list --project pixora --kind code --limit 20
npx localmind forget 42 --project pixora
npx localmind projects
npx localmind dbHow it works
you / your agent
│ MCP (stdio JSON-RPC)
▼
localmind server ─┬─► SQLite (better-sqlite3, WAL) ~/.localmind/memory.db
│ text, tags, project, file path, line range,
│ content hash, timestamps — and the embedding BLOBs,
│ which are the source of truth for every index
│
├─► HNSW index per project (hnswlib-node) ~/.localmind/indexes/<project>.hnsw
│ cosine space; labels are SQLite memory ids; deletes are
│ tombstones; persisted atomically and verified against
│ SQLite on open (rebuilt from the BLOBs if they disagree)
│
└─► all-MiniLM-L6-v2 (ONNX, in-process) 384-dim embeddings
bundled under models/; remote fetching hard-disabledVector search — one hnswlib graph per project via hnswlib-node (Apache-2.0, no network). Recall is approximate-nearest-neighbour: sub-second at six-figure vector counts instead of a full cosine scan per query. The benchmark measures the speedup rather than claiming it.
Incremental ingestion — the SHA-256 of a chunk's text is its identity. Unchanged chunks keep their row, embedding and index slot; moved chunks get a metadata-only line-range update; vanished chunks are tombstoned (
markDelete) and their slots reused for new vectors, so long-lived indexes never need a full rebuild.Self-healing index — SQLite owns the data, the index is a cache. If an index file is missing, corrupted, or disagrees with the database's id set (e.g. a crash between a write and a flush), it is rebuilt from the stored embedding BLOBs on next open.
Data & privacy
All data lives under one directory:
~/.localmind/—memory.dbplus oneindexes/<project>.hnswfile per project (override the directory withLOCALMIND_HOME).The embedding model is loaded from the package's own
models/directory; remote model fetching is disabled when the bundled model is present.No API keys, no telemetry, no network calls at runtime.
Deleting a project is a file delete plus a SQL delete — nothing is synced anywhere.
Platform note
better-sqlite3 and hnswlib-node ship prebuilt native binaries for the common platforms (macOS Intel/ARM, Linux x64/ARM, Windows x64); npm install downloads the right ones automatically. If a platform has no prebuilt binary, node-gyp builds from source (needs a C++ toolchain).
Security notes
No network at runtime. The embedding path loads the bundled ONNX model from disk; remote model fetching is disabled in code. Verify with:
grep -RIn "fetch(\|axios\|http" src/npm auditadvisory —@xenova/transformerspinssharp@0.32.6(image codec) which carries advisories. localmind's text-embedding path never invokes sharp (it's used only by image/audio pipelines), but if you treat supply-chain alerts as blockers, wait for upstream transformers.js ≥3 or vendor the model with your own loader.
Development
git clone https://github.com/sahinsamrat75/localmind
cd localmind
npm install
npm run build
npm test # smoke test over a real MCP stdio client
npm run test:isolation # project namespaces cannot leak into each other
npm run test:migration # v1 db migrates with zero data loss
npm run test:chunker # Go/TS/Python boundaries + fallback chunking
npm run test:incremental # re-ingestion skips unchanged chunks by hash
npm run test:budget # max_tokens bounds output, dedups, keeps rank order
npm run benchmark # 20k+ synthetic chunks: latency, memory, diskEvery suite drives the real server as a child process over stdio with the official MCP client — the same path an agent takes — except the chunker suite, which is pure unit-level. The benchmark generates a synthetic Go/TS/Python monorepo, ingests it into two projects, and reports recall latency percentiles, ANN-vs-brute-force speedup, incremental re-ingestion cost, index reload/rebuild, and memory/disk footprint; it fails non-zero if any scale gate is missed. Results land in test/benchmark-results.json.
License
This server cannot be deployed
Maintenance
Related MCP Connectors
Persistent memory for AI agents. Search and store durable facts, preferences and decisions.
Memory system for AI agents with semantic search. Store and recall memories with ease.
Persistent memory for AI agents. Semantic recall by meaning, not just keywords. No signup needed.
Persistent memory for AI agents. Search, store, and recall across sessions.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceProvides AI agents with persistent, searchable memory that survives across conversations using semantic search, temporal versioning, and smart organization. Enables long-term context retention and cross-session continuity for AI assistants.14-
- FlicenseNot gradedqualityDmaintenanceProvides persistent AI agent memory using a local vector database for long-term semantic storage and short-term session scratchpads. It enables low-latency memory operations including search, storage, and bulk management without external cloud dependencies.-
- AlicenseNot gradedqualityDmaintenanceProvides persistent memory storage for AI agents with full-text search, tagging, and importance levels, enabling agents to store and retrieve memories efficiently.MIT
- AlicenseNot gradedqualityDmaintenanceEnables AI agents to maintain persistent, local memory with retrieval-augmented search, knowledge graphs, and context surfacing, without any cloud dependencies.52 npmMIT