local-memory
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@local-memorysearch my memory for the postgres connection pool fix"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
local-memory
Local-first long-term memory for AI agents. SQLite, offline, zero cloud.
Your agent's memory stays on your machine. SQLite. No cloud. No API keys. Works offline.
local-memory is a standalone MCP server (stdio transport) that gives any AI agent durable long-term memory: search over past sessions (keyword, semantic, exact), a catalog of sessions, full session reads, and a write channel for new memories. It works with Claude Code, opencode, Cursor, Cline, Codex, or any other MCP client.
The engine is small and boring on purpose: one SQLite file, the Python
standard library, and numpy. The only other dependency is the mcp package,
which powers the stdio server.
Why
AI agents are stateless by default. Every new session starts from zero: the decisions made yesterday, the bugs already chased away, the project's conventions, the exact error message you fixed last week — all of it is gone unless you copy-paste it back in by hand.
The usual fix is a cloud memory API. But that means your conversations — often your most sensitive work context — are stored, indexed, and processed on someone else's servers. For local work that is a bad trade: the "memory" you paid for is also the leak you did not want.
local-memory flips the trade. The memory is a plain SQLite database on your machine, readable with any SQL client, back-upable with any backup tool, and searchable with your own embeddings model (or without any model at all). The agent gets the memory interface it wants; you keep the data you already had.
Related MCP server: engram-mcp
How it works
+------------------------------------------------+
| your machine (offline) |
| |
MCP client | +----------------+ +-------------+ |
(Claude Code, | | local-memory | stdio | SQLite | |
opencode, +-->| MCP server |<-------->| memory.db | |
Cursor, ...) | | | | + FTS5 | |
via stdio | +-------+-------+ +------+------+ |
| | ^ |
| | embeddings (optional, | |
| | only if YOU configure it) | |
| v | |
| +----------------+ +-------------+ |
| | embeddings API | | numpy matrix | |
| | (OpenAI-style) | | cache (.npy) | |
| +----------------+ +-------------+ |
+-------------------------------------------------+One database, three search engines, chosen automatically:
Mode | When used | How |
exact | query looks like code ( | literal LIKE, then FTS5 |
keyword | default for | FTS5 with prefix expansion, LIKE safe |
semantic |
| embeddings -> TF-IDF -> keyword |
Every result carries a context window: the neighboring chunks around the hit, so the agent sees what came before and after without an extra round trip.
Tools
Tool | Args | Returns |
|
| recent sessions: id, title, counts, times |
|
| keyword hits + context, engine used |
|
| semantic hits + context, engine used |
|
| consecutive chunks of one session |
| — | db statistics (counts, size, cache) |
|
| write result (position, ok) |
Result rows (a public contract — stable across versions):
{
"session_id": "alpha",
"position": 3,
"content": "the postgres connection pool timed out under load",
"score": 0.83,
"context": [
{"position": 2, "content": "..."},
{"position": 3, "content": "the postgres connection pool timed out under load"},
{"position": 4, "content": "..."}
]
}Quickstart
1. Install
From source (recommended until the package is published on PyPI):
git clone https://github.com/Sergey-Vladimirovich-Hankok/local-memory
cd local-memory
pip install -e . # core (keyword + FTS search)
pip install -e ".[semantic]" # + TF-IDF semantic search fallback (scikit-learn)Once available on PyPI:
pip install local-memory
# optional: TF-IDF semantic search fallback
pip install local-memory[semantic]2. Initialize the database
local-memory init # creates ~/.local/share/local-memory/memory.db
local-memory status # {"exists": true, "sessions": 0, ...}3. Write a memory
local-memory ingest --session-id demo --project myproject \
--text "the postgres connection pool timed out under load, raised max size"4. Search it
local-memory search "postgres connection pool" -k 5
local-memory search "pool timeout" --semantic5. Connect an MCP client
Claude Code:
claude mcp add local-memory -- local-memory serveopencode (opencode.json):
{
"mcp": {
"local-memory": {
"command": "uvx",
"args": ["local-memory", "serve"]
}
}
}Cursor / any generic MCP client (mcpServers snippet):
{
"mcpServers": {
"local-memory": {
"command": "uvx",
"args": ["local-memory", "serve"]
}
}
}If you installed from source (step 1) instead of PyPI,
uvxcannot see the package — pointcommandat the installed binary directly:"command": "local-memory"(seeexamples/generic_mcp.json).
Ready-made snippets live in examples/.
Configuration
Everything is configured via environment variables. There is no config file — one process, one database, zero moving parts.
Variable | Default | Meaning |
|
| path to the SQLite database |
| (empty = embeddings disabled) | OpenAI-compatible |
|
| model name sent to the embeddings API |
| (unset) |
|
|
| context window width (neighbors per side) |
|
| cap on rows scanned per search |
|
| max new chunks embedded per call |
Ingesting memory
memory_ingest / local-memory ingest is the write channel. A client hook or
plugin calls it once per message (or chunk) it wants the agent to remember.
Positions auto-increment per session; the FTS index is updated by a database
trigger, so ingested text is searchable immediately.
CLI:
local-memory ingest --session-id my-project --project my-project \
--text "decided: use SQLite WAL mode for the memory db" \
--metadata '{"source": "standup", "date": "2026-10-06"}'
local-memory ingest --session-id my-project --file notes.mdMCP tool:
{
"session_id": "my-project",
"content": "decided: use SQLite WAL mode for the memory db",
"project": "my-project",
"metadata": {"source": "standup"}
}Example hook script (call from your agent's post-message hook):
#!/usr/bin/env bash
# remember.sh — feed one message into local-memory
local-memory ingest \
--session-id "project-${1}" \
--text "${2}" \
--project "hooks" >/dev/nullSemantic search (optional)
memory_search_semantic tries, in order:
Embeddings — any OpenAI-compatible
/v1/embeddingsendpoint (llama.cppllama-server, Ollama, vLLM, a local transformer service...). PointEMBED_API_URLat it; vectors are cached on disk next to the database (embed_cache_*.npy) and reused, so repeated queries are fast.TF-IDF — character n-gram cosine similarity, offline, needs the optional
scikit-learnextra.Keyword fallback — the exact/FTS5/LIKE engines, with
"engine": "keyword-fallback"in the response.
Prebuilding the cache (only needed for the embeddings path):
export EMBED_API_URL=http://localhost:8080/v1/embeddings
export EMBED_MODEL=my-local-embedding-model
local-memory build # embed all chunks onceIf no embeddings endpoint is configured, semantic search silently degrades to
TF-IDF/keyword. Nothing is sent over the network unless you configure a
EMBED_API_URL.
Privacy
The database, FTS index, and embedding caches are plain files under
MEMORY_DB_PATH(default~/.local/share/local-memory/). Back them up or delete them like any other file.No telemetry, no update checks, no phone-home of any kind.
The only network access in the entire codebase is the embeddings call, and it only happens if you set
EMBED_API_URLyourself.memory_ingeststores content as-is, unencrypted, in SQLite. If your data is sensitive, protect the file with filesystem permissions.Nothing is ever sent to the authors of this project.
Testing
The test suite runs fully offline against a synthetic in-memory-style fixture database (three fake sessions, ~30 chunks). No real conversation data is read, no network is touched.
python3 -m pytest tests/ -qContributing
See CONTRIBUTING.md. Short version: local-first is the product promise; keep dependencies light; keep fixtures synthetic; add tests.
Security
See SECURITY.md. Report vulnerabilities privately via the repository's Security tab — not in public issues.
License
MIT — do what you want, just keep the notice.
Attribution — if you use or copy this code, keep the author notice in the source headers, the repository link (https://github.com/Sergey-Vladimirovich-Hankok/local-memory), and the contact email (kokgfnu@gmail.com). MIT license covers permissions; attribution keeps the work traceable. See NOTICE.
Roadmap
Qdrant / sqlite-vec backend for the embeddings matrix (still local).
REST mode (HTTP transport alongside stdio) for non-MCP clients.
Plugin packages for specific agents (Claude Code, opencode, Cursor).
Session summaries via any local LLM (opt-in, offline).
Not on the roadmap: cloud sync, web UI, Rust rewrites. If it must be local, it stays local.
Disclaimer
local-memory is provided for educational and personal use. You are responsible for what you ingest into it and for protecting the resulting database file. The authors are not liable for lost, leaked, or misused data — the database is yours, and so is the responsibility.
This server cannot be deployed
Maintenance
Related MCP Connectors
Persistent memory for AI agents. Search, store, and recall across sessions.
Persistent memory, hybrid search and a goal graph for AI agents, over stdio or remote HTTP.
Persistent memory for AI agents. Search and store durable facts, preferences and decisions.
Universal memory for AI agents and tools. Save, organize and search context anywhere.
Related MCP Servers
- -licenseNot gradedqualityNot gradedmaintenanceProvides persistent local memory functionality for AI assistants, enabling them to store, retrieve, and search contextual information across conversations with SQLite-based full-text search. All data stays private on your machine while dramatically improving context retention and personalized assistance.3-
- AlicenseAqualityCmaintenancePersistent semantic memory for AI agents. SQLite-backed, local-first, zero config. Semantic search via Ollama embeddings with keyword fallback. Tools: remember, recall, history, forget, stats.1737 npm1MIT
- AlicenseNot gradedqualityBmaintenanceProvides long-term memory for LLMs via local SQLite storage with hybrid search (BM25, vectors, recency decay), enabling AI coding agents to persist and recall memories across sessions without cloud or API keys.53MIT
- AlicenseNot gradedqualityDmaintenanceProvides persistent long-term memory for LLMs via local SQLite storage and semantic search, enabling recall across sessions without external APIs.15 npm5MIT