mem0-local-mcp
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@mem0-local-mcpwhat do you remember about my coding preferences?"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
mem0-local-mcp
Long-term, semantic memory for Claude Code (or any MCP client), powered by mem0, and running entirely on your machine.
No API key. No OpenAI, no mem0 cloud account.
No local LLM needed. Your agent (e.g. Claude) decides what's worth remembering. mem0 handles storage, embeddings, and semantic search.
Light. Embeddings run through FastEmbed (ONNX, no PyTorch). The default model is ~67 MB and is downloaded once.
Private. Memories are stored in a local Chroma DB. mem0 telemetry is disabled.
How it works
Claude Code ──MCP (stdio)──▶ mem0-local-mcp ──▶ mem0 ──▶ FastEmbed (local embeddings)
└─▶ Chroma (~/.mem0-local-mcp)The server exposes three tools:
Tool | What it does |
| Stores one fact verbatim and returns its id |
| Runs a semantic search and returns |
| Removes a stale or wrong memory |
Related MCP server: mcp-memory
Requirements
Python 3.10+ (uv can install it for you)
Install in Claude Code
claude mcp add mem0 -s user -- \
uvx --from git+https://github.com/ShmuelOps/mem0-local-mcp mem0-local-mcp-s user makes the memory available in every project. Use -s project to share the config through
.mcp.json, or -s local for the current project only.
Verify:
claude mcp list # mem0: ... ✔ ConnectedThe first start downloads the dependencies and the embedding model, so it can take a minute. Later starts take a few seconds.
Tell Claude when to use it
Add this to ~/.claude/CLAUDE.md (global) or to a project's CLAUDE.md:
## Long-term memory (mem0 MCP)
- At the start of a non-trivial task, call `search_memory` with the task topic.
- When you learn a durable fact (user preference, project decision, gotcha), `search_memory`
first, then `add_memory` one concise sentence if it's new.
- `delete_memory` entries that turn out wrong or stale.Other MCP clients
Any client that speaks MCP over stdio works. For example, Claude Desktop's
claude_desktop_config.json:
{
"mcpServers": {
"mem0": {
"command": "uvx",
"args": ["--from", "git+https://github.com/ShmuelOps/mem0-local-mcp", "mem0-local-mcp"]
}
}
}Configuration
All configuration is optional and set through environment variables. With claude mcp add, pass
them as -e KEY=value.
Variable | Default | Purpose |
|
| Memory namespace. Use different values to keep separate memory sets |
|
| Where the Chroma DB and history are stored |
|
|
Example with a per-project namespace:
claude mcp add mem0 -s project -e MEM0_USER=my-project -- \
uvx --from git+https://github.com/ShmuelOps/mem0-local-mcp mem0-local-mcpChanging
MEM0_EMBED_MODELafter you've stored memories requires a freshMEM0_DATA_DIR, because vectors from different models are not compatible.
Data & privacy
Everything is stored in
MEM0_DATA_DIR. Delete that directory to wipe all memories.The only network access is the one-time download of packages and the embedding model.
mem0 telemetry is turned off (
MEM0_TELEMETRY=False).
Troubleshooting
Symptom | Fix |
| The first launch is still downloading. Run the |
Duplicate memories | This happens because |
Want to wipe everything |
|
Development
git clone https://github.com/ShmuelOps/mem0-local-mcp && cd mem0-local-mcp
uv sync
uv run pytest # real end-to-end tests: mem0 + Chroma + FastEmbed, no mocks
uv run ruff check . && uv run ruff format --check .Run the server from source:
claude mcp add mem0-dev -- uv run --directory "$PWD" mem0-local-mcpSee CONTRIBUTING.md.
License
Available Tools
3 toolsadd_memoryA
Store one concise, durable fact worth remembering across sessions.
Search first to avoid duplicates. Returns the new memory id.
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It discloses the return value (the new memory id) and constrains input to one concise durable fact, but says nothing about permissions, idempotency, overwrite/duplicate behavior, or limits on how many memories can be stored.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short, front-loaded sentences with zero waste: the action and its scope first, then the deduplication guidance, then the return value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter mutation tool with an output schema present, the description covers purpose, precondition, and return, which is nearly sufficient. Only the handling of duplicates and any storage limits remain unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% for the single 'text' parameter, so the description must compensate. It does add meaning by specifying the input should be 'one concise, durable fact worth remembering across sessions,' which constrains what to pass, but gives no format, length, or example guidance.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (store) and resource (a durable fact/memory) with a clear qualifier about scope across sessions. This cleanly distinguishes it from the sibling search_memory and delete_memory without needing to name them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly instructs the agent to 'Search first to avoid duplicates,' which routes to the sibling search_memory and gives the precondition for calling add_memory. It lacks an explicit when-not clause (e.g., what to do on a near-duplicate), but the routing guidance is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_memoryB
Delete a stale or wrong memory by id.
| Name | Required | Description | Default |
|---|---|---|---|
| memory_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden, and it discloses almost nothing. It does not state whether deletion is permanent/irreversible, whether permissions or confirmation are required, or what happens on a missing id. For a destructive single-target operation this is a significant gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence with the action and target front-loaded and no wasted words. Nothing is padded or redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be explained, but for a destructive tool with no annotations the description omits the essential behavioral context (permanence, permissions, error behavior). It is not complete enough for an agent to call it safely.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and 'by id' at least confirms that the single parameter is the identifier of the target memory. It adds no format, source, or validation detail beyond that, leaving the meaning of memory_id largely to the schema's type alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb ('Delete') and resource ('a memory') and identifies the target via id, which cleanly separates it from the add_memory and search_memory siblings. It stops short of explicitly naming those siblings or stating what makes a memory 'stale or wrong', so it earns a 4 rather than a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'stale or wrong memory' implies the condition under which deletion is appropriate, so usage is implied rather than spelled out. There is no explicit guidance on when not to delete, no mention of an alternative (e.g., updating instead), and no prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_memoryB
Semantic search over long-term memory. Returns one 'id: memory' per line.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| query | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It usefully discloses the output shape ('one id: memory per line'), but says nothing about how results are ranked, whether the search is scoped, or any auth/rate considerations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with purpose followed by return format. Nothing wasteful, though the return-format sentence partially duplicates the output schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter tool with an output schema, purpose and response format are covered. However, with 0% schema coverage the meaning and behavior of 'query' and 'limit' remain unexplained, which is a real gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the schema documents neither parameter. The description does not explain what 'query' accepts (natural language? keywords?) or how 'limit' affects results, leaving both parameters semantically bare.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (semantic search) and resource (long-term memory), which cleanly separates it from the add_memory and delete_memory siblings. It does not explicitly name those siblings, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No statement of when to use this vs. alternatives, no prerequisites, no mention of when semantic search is preferable to other retrieval approaches. Usage is only implied by the verb.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
v0.1.0- First observed
add_memory - First observed
delete_memory - First observed
search_memory
TDQS
Scored across 3 tools
The three tools map to clearly distinct operations: add_memory (create), search_memory (read/query), and delete_memory (remove). Descriptions reinforce the boundaries, e.g. add_memory explicitly tells the agent to search first, eliminating overlap.
All three tools follow the identical verb_noun pattern (add_memory, search_memory, delete_memory) with the same noun for the resource. This is a textbook consistent naming scheme.
Three tools is a tight, well-scoped surface for a simple local long-term memory store. Each tool earns its place with no redundancy or filler.
Create, query, and delete cover the core memory lifecycle, and semantic search serves as retrieval. An explicit update/edit tool and a get-by-id or list operation are missing, though delete+add can work around the gap.
Maintenance
Related MCP Connectors
Private persistent memory for Claude, ChatGPT & Gemini via MCP - semantic search, zero-code setup.
Persistent personal memory for AI assistants — save, search, and recall across every MCP client.
Persistent memory for AI agents across Claude, ChatGPT and any MCP client.
- KogniteOAuthdev.kognite
Hosted agent memory: store, search, and recall facts across sessions from any MCP client.
Related MCP Servers
- AlicenseNot gradedqualityAmaintenanceProvides persistent memory for AI coding agents via MCP, enabling agents to store and semantically recall facts, events, and lessons across sessions, all running locally without cloud dependencies.Apache 2.0
- AlicenseAqualityDmaintenanceProvides persistent memory with semantic search for MCP-based AI agents, enabling them to store and recall information across sessions using vector embeddings.41MIT
- AlicenseNot gradedqualityBmaintenanceProvides persistent semantic memory for AI agents via MCP, enabling them to remember, recall, list, update, and forget memories with vector-based similarity search.ISC
- AlicenseAqualityBmaintenanceProvides persistent, searchable memory for AI agents across any MCP-compatible client, storing project context, user preferences, and session learnings locally in SQLite with tools to save, retrieve, search, and manage them.127 npmMIT