mem-persistence
README.md
# mem-persistence
> π§ Persistent memory MCP server for AI agents β one memory, every agent, your files.
mem-persistence lets Claude Desktop, Claude Code, Cursor, Zed, and any MCP-compatible client share the same persistent memory, backed by plain Markdown files you own and can edit by hand.
## Why?
AI agents have amnesia. Each tool keeps its own silo β Claude Code forgets what OpenClaw knows, Cursor can't recall what you told Claude yesterday. Your context is scattered across sessions that evaporate.
mem-persistence fixes this:
- **Markdown is the source of truth** β not a database, not a binary blob. Files you can read, edit, and version with git.
- **Hybrid search** β token matching + semantic embeddings for accurate recall.
- **Embedding providers** β Ollama (local & private, recommended), Gemini, OpenAI, or none (token-only). Cached to disk.
- **Deduplication** β prevents writing the same fact twice (token + entity overlap detection).
- **Works offline** β no cloud dependency. Embeddings are optional.
## MCP Tools
| Tool | Description |
|---|---|
| `memory_search(query, maxResults?)` | Hybrid search across all `.md` files |
| `memory_write(content, file?, section?)` | Write with automatic deduplication |
| `memory_read(path, from?, lines?)` | Read a specific file or section |
| `memory_checkpoint(summary)` | Save a session checkpoint to a daily note |
| `memory_entities(query?)` | Query the knowledge graph (if `entities.md` exists) |
| `memory_status()` | Index stats: files, chunks, last sync |
| `fact_save(entity, attribute, value, confidence?, source?)` | Save an exact (entity, attribute, value) fact with a `verifiedβhigh_probabilityβfalse` confidence ladder |
| `fact_get(entity, attribute)` | Exact key-value lookup |
| `fact_query(entity?, attribute?, confidence?)` | List facts by filter |
| `fact_demote(entity, attribute, source?)` | Mark a fact false without a replacement |
Facts are stored in `data/facts.db` (built-in `node:sqlite`), separate from the markdown corpus β exact key-value recall for names, settings, and IDs that semantic search handles badly.
---
## Quick Start
### 1. Install and build
```bash
git clone https://github.com/emiliotorrens/mem-persistence.git
cd mem-persistence
npm install
npm run build
```
### 2. Start the server
```bash
node dist/index.js --workspace /path/to/your/workspace --port 3456
```
### 3. Connect a client
Pick the setup that matches your client β see [Client Setup](#client-setup) below.
### 4. Add agent instructions
Copy [`AGENT_INSTRUCTIONS.md`](AGENT_INSTRUCTIONS.md) into your agent's instruction file:
| Editor | Where to paste |
|---|---|
| Claude Desktop | Settings β Personal Preferences |
| Claude Code | `CLAUDE.md` in project root |
| Cursor | `.cursorrules` in project root |
| Windsurf | `.windsurfrules` in project root |
---
## Client Setup
There are two MCP transports. Which one you need depends on the client:
| Transport | Clients | Where server runs | Remote access |
|---|---|---|---|
| **HTTP** | Claude Code, Cursor, Zed | Anywhere (local or remote) | β
via Tailscale/VPN |
| **stdio** | Claude Desktop | Same machine as Desktop | β (see [proxy workaround](#remote-claude-desktop-via-proxy)) |
### HTTP clients (Claude Code, Cursor, Zed)
Point to the running server URL:
```json
{
"mcpServers": {
"memory": {
"url": "http://127.0.0.1:3456/mem-persistence/mcp"
}
}
}
```
For remote access over Tailscale, replace `127.0.0.1` with the server's Tailscale hostname:
```json
{
"mcpServers": {
"memory": {
"url": "http://my-machine.tail1234.ts.net:3456/mem-persistence/mcp"
}
}
}
```
### Claude Desktop (stdio, same machine)
Claude Desktop only supports stdio β it spawns mem-persistence as a child process.
- macOS: `~/Library/Application Support/Claude/claude_desktop_config.json`
- Windows: `%APPDATA%\Claude\claude_desktop_config.json`
```json
{
"mcpServers": {
"memory": {
"command": "node",
"args": [
"/path/to/mem-persistence/dist/index.js",
"--workspace", "/path/to/your/workspace"
],
"env": {
"MEM_PERSISTENCE_EMBEDDINGS": "ollama",
"MEM_PERSISTENCE_EMBEDDINGS_MODEL": "embeddinggemma"
}
}
}
}
```
See [Embeddings](#embeddings) for the other providers. If you omit `env` entirely, the server falls back to token-only search (no semantic matching).
> **WSL users (Windows):** replace `"command": "node"` with `"command": "wsl"` and add `"node"` as the first element of `args`.
>
> β οΈ Claude Desktop's `env` block is set on the **Windows** process and is **not** inherited by the WSL child. Pass the variables inside the command instead:
>
> ```json
> {
> "mcpServers": {
> "memory": {
> "command": "wsl",
> "args": [
> "-e", "env",
> "MEM_PERSISTENCE_EMBEDDINGS=ollama",
> "MEM_PERSISTENCE_EMBEDDINGS_MODEL=embeddinggemma",
> "node", "/path/to/mem-persistence/dist/index.js",
> "--workspace", "/path/to/your/workspace"
> ]
> }
> }
> }
> ```
> β οΈ **Do not pass `--port` in stdio mode.** It causes an `EADDRINUSE` conflict if an HTTP instance is already running.
> π‘ **Already running an HTTP instance?** Prefer the [proxy setup](#remote-claude-desktop-via-proxy) over a second stdio process β even on the same machine. One server means one embedding config and one cache instead of two.
### Remote Claude Desktop via proxy
If Claude Desktop runs on a **different machine** (e.g., a laptop) where mem-persistence isn't installed, use the bundled `mcp-proxy.js` to bridge stdio to the remote HTTP server.
**Requirements on the client machine:** Node.js + Tailscale. That's it β no cloning, no `npm install`.
1. Copy `mcp-proxy.js` to the laptop (one file, zero dependencies).
2. Add to Claude Desktop config:
```json
{
"mcpServers": {
"memory": {
"command": "node",
"args": ["/path/to/mcp-proxy.js"],
"env": {
"MCP_REMOTE_URL": "http://my-machine.tail1234.ts.net:3456/mem-persistence/mcp"
}
}
}
}
```
Desktop thinks it's talking to a local stdio server; the proxy forwards everything over HTTP.
Set `MCP_DEBUG=1` to log proxy traffic to stderr for troubleshooting.
---
## Running as a Service (PM2)
For production use, run the server as a persistent background service with PM2:
```bash
# 1. Install pm2
npm install -g pm2
# 2. Copy and edit the config
cp ecosystem.config.cjs.example ecosystem.config.cjs
# β Set workspace path and optional API keys
# 3. Start and persist
pm2 start ecosystem.config.cjs
pm2 save
pm2 startup # autostart on reboot (follow the printed instructions)
```
Health check: `curl http://127.0.0.1:3456/health`
---
## Network Binding
By default, the server listens on `127.0.0.1` only. Use `--bind` to control which interfaces it binds to:
```bash
# Localhost + Tailscale (recommended for remote access)
node dist/index.js --workspace /path --port 3456 --bind 127.0.0.1,tailscale
# Localhost + explicit VPN IP
node dist/index.js --workspace /path --port 3456 --bind 127.0.0.1,10.0.0.5
# All interfaces (β οΈ only behind a firewall)
node dist/index.js --workspace /path --port 3456 --bind all
```
| `--bind` value | Resolves to |
|---|---|
| `localhost` | `127.0.0.1` |
| `tailscale` | Auto-detected via `tailscale ip -4` (100.x.x.x) |
| `all` / `0.0.0.0` | All network interfaces |
| Any IP | Used as-is |
In `ecosystem.config.cjs`:
```js
args: '--workspace /path --port 3456 --bind 127.0.0.1,tailscale',
```
Or via environment variable: `MEM_PERSISTENCE_BIND=127.0.0.1,tailscale`
> β οΈ **Security:** mem-persistence has no built-in authentication. **Never expose the port to the public internet.** Use `--bind 127.0.0.1,tailscale` to limit access to localhost + your private network.
---
## Workspace
The `--workspace` flag points to the directory containing your memory files. mem-persistence indexes all `.md` files recursively.
Any directory with `.md` files works. Search quality improves with a layered layout:
| Layer | Path | Purpose |
|---|---|---|
| L1 | `MEMORY.md` | Long-term curated memory β highest search priority |
| L2 | `memory/*.md` | Daily notes, recent context |
| L3 | `reference/*.md` | Detailed data, historical records |
For automatic setup of this structure (with crons, dedup, and knowledge graph), see [layered-memstack](https://github.com/emiliotorrens/layered-memstack).
You can also set the workspace via environment variable: `MEM_PERSISTENCE_WORKSPACE=/path/to/workspace`
---
## Embeddings
By default, search uses **token matching only** (Jaccard + containment + entity overlap). No API calls, works offline.
Enabling embeddings adds **semantic understanding**:
| Query | Token-only | With embeddings |
|---|---|---|
| `"where does Emilio work"` | β no keyword overlap | β
understands meaning |
| `"what trips are coming up?"` | β misses if phrased differently | β
matches semantically |
Configure via environment variables. Three providers are supported β **local (Ollama) is recommended**: embeddings never leave your machine and there are no API quotas.
**Option A β Ollama (local, private) β recommended**
Run [Ollama](https://ollama.com) and pull a small embedding model:
```bash
ollama pull embeddinggemma # 621 MB Β· 768 dims Β· multilingual
MEM_PERSISTENCE_EMBEDDINGS=ollama
MEM_PERSISTENCE_EMBEDDINGS_MODEL=embeddinggemma # optional (default)
MEM_PERSISTENCE_EMBEDDINGS_BASE_URL=http://127.0.0.1:11434 # optional (default)
```
**Option B β Cloud (Gemini / OpenAI)**
Zero local footprint, but every chunk is sent to the provider's API and is subject to quotas. Get a free Gemini API key β [aistudio.google.com](https://aistudio.google.com).
```bash
MEM_PERSISTENCE_EMBEDDINGS=gemini # "gemini" or "openai"
GOOGLE_API_KEY=your-key # Gemini β free
OPENAI_API_KEY=your-key # OpenAI β $0.02/M tokens
```
Details:
- **Hybrid scoring**: 0.4 Γ token + 0.6 Γ vector
- **Disk cache**: `.mem-persistence/embeddings/` β keyed by model, no repeated calls
- **Silent fallback**: if the provider is unavailable, falls back to token-only automatically
---
## Deduplication
Before writing, mem-persistence checks if similar content already exists:
```
Input: "GitHub configured with gh auth login, user emiliotorrens"
Match: "gh auth login hecho β cuenta emiliotorrens, protocolo HTTPS"
Result: DUPLICATE (score: 0.90) β not written
```
Uses token similarity (Jaccard + containment) and entity overlap (IDs, dates, versions, URLs).
Adjust the threshold: `MEM_PERSISTENCE_DEDUP_THRESHOLD=0.65` (default β lower = stricter).
---
## OpenClaw Integration
If you use [OpenClaw](https://github.com/openclaw/openclaw), mem-persistence coexists with OpenClaw's native memory:
- **External clients** (Claude Desktop, Code, Cursor) β connect via mem-persistence (stdio or HTTP)
- **OpenClaw agent** β uses its native `memory-core` plugin with hybrid search + embeddings
Both systems index the same Markdown files. mem-persistence is the MCP bridge for external clients; OpenClaw handles its own recall, wiki compilation, and dreaming.
---
## Roadmap
- [x] Deduplication engine
- [x] Hybrid search (token + vector + MMR + temporal decay)
- [x] MCP server β stdio and HTTP transports, 6 tools, TypeScript + ESM
- [x] Embedding providers: Gemini (free) and OpenAI, with disk cache
- [x] Request/response logging (`.mem-persistence/logs/`)
- [x] HTTP mode β Tailscale-friendly, pm2-ready
- [x] stdioβHTTP proxy for remote Claude Desktop
- [ ] CLI (`mem-persistence search "query"`)
- [ ] Local embeddings via transformers.js (offline, no API key)
- [ ] npm publish
---
## Related
- **[layered-memstack](https://github.com/emiliotorrens/layered-memstack)** β OpenClaw skill that sets up a 3-layer memory system with automated maintenance. Uses mem-persistence as the MCP bridge for external clients.
## Credits
- **[OpenClaw](https://github.com/openclaw/openclaw)** β the agent framework where this was born and battle-tested
- **[MCP](https://modelcontextprotocol.io)** β the protocol that makes cross-agent memory possible
## License
MIT
---
Built with πΎ by [Emilio Torrens](https://github.com/emiliotorrens) and [Claw](https://github.com/openclaw/openclaw).
This server cannot be deployed
Maintenance
ActivitySlowing
ResponsivenessNo issues