MCP Memory Service
This server provides persistent shared memory for AI agents, allowing storage and retrieval of context across sessions and runs.
Store new information with optional tags and metadata.
Retrieve relevant memories using natural-language queries, with a configurable number of results (default 5).
Search memories by tags.
Use it via MCP, REST API, CLI, or web dashboard; supports knowledge graphs, local embeddings, OAuth, agent identity tagging, conversation storage, and real-time SSE events.
Mentioned as a potential cloud storage option where users should ensure sync is complete before accessing from another device.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@MCP Memory Servicesearch for my notes about authentication setup"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
mcp-memory-service
Persistent Shared Memory for AI Agent Pipelines
Open-source memory backend for AI agents — REST API, MCP, OAuth, CLI, dashboard. One self-hosted service, every transport. Agents store decisions, share causal knowledge graphs, and retrieve context in 5ms — without cloud lock-in or API costs.
Works with LangGraph · CrewAI · AutoGen · any HTTP client · Claude Desktop · OpenCode
Related MCP server: Neuro MCP V2
Why Agents Need This
Your AI assistant forgets everything when you start a new chat. You spend 10 minutes re-explaining your architecture. Again. MCP Memory Service captures project context, architecture decisions, and code patterns automatically — new sessions start with everything already known.
Without mcp-memory-service | With mcp-memory-service |
Each agent run starts from zero | Agents retrieve prior decisions in 5ms |
Memory is local to one graph/run | Memory is shared across all agents and runs |
You manage Redis + Pinecone + glue code | One self-hosted service, zero cloud cost |
No causal relationships between facts | Knowledge graph with typed edges (causes, fixes, contradicts) |
Context window limits create amnesia | Autonomous consolidation compresses old memories |
Key capabilities for agent pipelines:
Framework-agnostic REST API — no MCP client library needed; the live endpoint list is at
/api/docsKnowledge graph — agents share causal chains, not just facts
X-Agent-IDheader — auto-tag memories by agent identity for scoped retrievalconversation_id— bypass deduplication for incremental conversation storageSSE events — real-time notifications when any agent stores or deletes a memory
Embeddings run locally via ONNX — memory never leaves your infrastructure
🚀 Get Started in 60 Seconds
Not sure which setup fits? The Setup Guide walks you to the right path in under a minute.
1. Install:
pip install mcp-memory-service2. Configure your AI client:
Add to your config file:
macOS:
~/Library/Application Support/Claude/claude_desktop_config.jsonWindows:
%APPDATA%\Claude\claude_desktop_config.jsonLinux:
~/.config/Claude/claude_desktop_config.json
{
"mcpServers": {
"memory": {
"command": "memory",
"args": ["server"]
}
}
}Restart Claude Desktop. Your AI now remembers everything across sessions.
claude mcp add memory -- memory serverRestart Claude Code. Memory tools will appear automatically.
MCP_ALLOW_ANONYMOUS_ACCESS=true memory server --http
# REST API running at http://localhost:8000Store a memory with POST /api/memories, search with POST /api/search, retrieve by tag
with POST /api/search/by-tag. Send X-Agent-ID: <id> on a store request and the server
tags the memory agent:<id>, which a tag search then scopes retrieval by.
Worked examples per framework, the tag conventions and the async patterns: docs/agents/
MCP_ALLOW_ANONYMOUS_ACCESS=true memory server --httpThe plugin ships as repository files for the local plugin directory, so clone the
repository once even if you installed from PyPI. Install steps, the /memory slash
command and the endpoint override:
opencode/README.md
Remote MCP puts persistent memory in the browser on any device, no desktop app required: OAuth 2.0 over HTTPS, self-hosted or cloud-hosted. Both claude.ai and ChatGPT (Developer Mode) connect to the same endpoint.
Cloudflare Tunnel quick start, Let's Encrypt, nginx, Caddy and Docker production setups: Remote MCP Setup · 5-minute tutorial
git clone https://github.com/doobidoo/mcp-memory-service.git
cd mcp-memory-service
python scripts/installation/install.pyChoose from SQLite (local, fast, single-user), Cloudflare (cloud, multi-device sync), Hybrid (5ms local reads with background cloud sync — recommended for production) or Milvus (dedicated vector DB: Lite file, self-hosted, or Zilliz Cloud).
For long-lived services, prefer Docker Milvus or Zilliz Cloud over Milvus Lite — why.
⚡ Works With Your Favorite AI Tools
Agent frameworks (REST): LangGraph · CrewAI · AutoGen · OpenClaw/Nanobot · any HTTP client
CLI and terminal (MCP): Claude Code · Gemini CLI · OpenCode · Codex CLI · Goose · Aider · Amp
Desktop and IDE (MCP): Claude Desktop · VS Code · Cursor · Windsurf · Raycast · JetBrains · Zed
Chat (MCP): ChatGPT (Developer Mode) · claude.ai (Remote MCP over HTTPS)
Full list, plus clients without OAuth such as Home Assistant: docs/integrations.md
✨ Features
🧠 Persistent Memory – Context survives across sessions with semantic search
🔍 Smart Retrieval – Finds relevant context automatically using AI embeddings
⚡ 5ms Speed – Instant context injection, no latency
☁️ Cloud Sync – Optional Cloudflare backend for team collaboration
🔒 Privacy-First – Local-first, you control your data
📊 Web Dashboard – Visualize and manage memories at http://localhost:8000
🧬 Knowledge Graph – Interactive D3.js visualization of memory relationships
🏠 Homelab Quality Scoring – Point scoring at any OpenAI-compatible endpoint (Ollama, LiteLLM, vLLM)
🔗 Entity Extraction – Auto-links @mentions, #tags, URLs, and file paths to a queryable entity graph
💡 Insight Cards – Consolidation surfaces patterns, trends, and knowledge gaps as structured insights
🏷️ Tag Match Filtering – tag_match=AND/OR on memory_search for precise multi-tag queries
The dashboard has eight tabs — Dashboard, Search, Browse, Documents, Manage, Analytics, Quality, API Docs. Two-minute walkthrough on YouTube · Web Dashboard Guide
How it compares to Mem0, Zep and the MCP-native alternatives, benchmark results, and deployments people run in production: mcpmemory.services
🛠️ Configuration Highlights
memory launch # Start HTTP server in background (127.0.0.1:8000)
memory launch --port 8192 # Custom port
memory info # Status and health
memory logs --lines 50 # Recent logs
memory stop # Stop serverThese commands are optimized for fast startup and avoid loading heavy ML dependencies unless needed.
⚠️ Security note: the server binds to
127.0.0.1(localhost only) by default.--host 0.0.0.0/MCP_HTTP_HOST=0.0.0.0exposes the API to your network — do that only in trusted environments with authentication and firewall rules, or behind TLS termination or a VPN overlay.
Backends, embedding models, quality scoring and every environment variable: Configuration Guide
📚 Documentation
Setup Guide – Decision tree and step-by-step paths
Agent Integration Guides – LangGraph, CrewAI, AutoGen, HTTP generic
Remote MCP Setup – claude.ai and ChatGPT via HTTPS + OAuth
Configuration Guide – Backend options and customization
Architecture Overview – How it works under the hood
Knowledge Graph Dashboard – Interactive graph visualization
Benchmarks – LongMemEval, DevBench, LoCoMo, with run commands
Migration Guide – Upgrading between major versions
Troubleshooting – Common issues and solutions
Wiki – Long-form guides and API reference
Full documentation index – Everything else
Also listed on Glama and Spark.
📦 Releases
Every release, with upgrade notes: CHANGELOG.md · GitHub Releases · archived history
🤝 Contributing
We welcome contributions! See CONTRIBUTING.md for guidelines and SECURITY.md for reporting a vulnerability.
Who authors this project, who holds copyright, and what every change passes before
it reaches main: AUTHORSHIP.md.
Quick Development Setup:
git clone https://github.com/doobidoo/mcp-memory-service.git
cd mcp-memory-service
pip install -e . # Editable install
pytest tests/ # Run test suiteSupporting the Project
MCP Memory Service is maintained by one person. If it saves you or your company time, you can support its development via Ko-fi, Buy Me a Coffee or PayPal: SPONSORS.md.
Available Tools
3 toolsretrieve_memoryC
Find relevant memories based on query
| Name | Required | Description | Default |
|---|---|---|---|
| n_results | No | ||
| query | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden but provides minimal behavioral context. It mentions 'find relevant memories' but doesn't disclose how relevance is scored, whether results are paginated, if there are rate limits, authentication needs, or what happens on failure. The description lacks details needed for safe and effective use.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence with no wasted words. It's front-loaded with the core action ('Find relevant memories'), though it could be more structured with additional context. For its brevity, it communicates the essence without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations, 0% schema coverage, no output schema, and two parameters, the description is incomplete. It doesn't explain what 'memories' are, how they're retrieved, the return format, or error handling. For a tool with query and result-limit parameters, more context is needed for effective use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate but adds no parameter-specific information. It mentions 'query' generally but doesn't explain its format, constraints, or how 'n_results' affects output. The description fails to clarify semantics beyond the bare schema, leaving parameters poorly understood.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description 'Find relevant memories based on query' states the general purpose (verb 'find' + resource 'memories') but lacks specificity about what 'memories' are or how relevance is determined. It distinguishes from 'store_memory' but not clearly from 'search_by_tag' (both involve finding memories). The purpose is understandable but vague.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives like 'search_by_tag'. The description implies usage for query-based retrieval, but there's no explicit mention of when-not-to-use, prerequisites, or comparison with siblings. Usage is implied from the name and description alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_by_tagC
Search memories by tags
| Name | Required | Description | Default |
|---|---|---|---|
| tags | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It states 'Search' which implies a read operation, but doesn't disclose behavioral traits like whether it's paginated, returns partial matches, requires authentication, or has rate limits. This is inadequate for a search tool with zero annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence with zero waste. It's appropriately sized and front-loaded, making it easy to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of a search operation, no annotations, no output schema, and low schema coverage, the description is incomplete. It lacks information on return values, error conditions, and behavioral context, making it insufficient for effective tool use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It mentions 'by tags' which hints at the 'tags' parameter, but doesn't add meaning beyond the schema's basic type information—no details on tag format, case sensitivity, or how multiple tags are combined (AND/OR). This partially compensates but leaves significant gaps.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description 'Search memories by tags' clearly states the verb ('Search') and resource ('memories'), but it's vague about scope and doesn't distinguish from sibling tools like 'retrieve_memory'. It doesn't specify whether this searches all memories or a subset, or how it differs from the retrieval sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives like 'retrieve_memory'. The description implies usage for tag-based searching but doesn't mention prerequisites, exclusions, or comparative contexts with siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
store_memoryC
Store new information with optional tags
| Name | Required | Description | Default |
|---|---|---|---|
| content | Yes | ||
| metadata | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden for behavioral disclosure. It states 'store new information' which implies a write/mutation operation, but doesn't specify permissions needed, whether storage is persistent, rate limits, or what happens on success/failure. This leaves significant gaps for a tool that appears to create data.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise at just 5 words, front-loading the core purpose without any wasted words. Every element ('store', 'new information', 'optional tags') contributes directly to understanding the tool's function.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given a mutation tool with no annotations, 2 parameters (one nested), 0% schema coverage, and no output schema, the description is inadequate. It doesn't explain what 'storing' entails operationally, what format the information should be in, how tags are used, or what the tool returns. The agent lacks critical context for proper invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for undocumented parameters. It mentions 'information' and 'optional tags' which loosely map to 'content' and 'metadata.tags', but doesn't explain the 'metadata.type' parameter at all or provide any format/constraint details. This partial coverage is insufficient given the schema's complexity with nested objects.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('store') and resource ('new information') with additional functionality ('with optional tags'), making the purpose understandable. However, it doesn't explicitly differentiate from sibling tools like 'retrieve_memory' or 'search_by_tag', which would require mentioning this is specifically for creating/adding new memories rather than retrieving or searching existing ones.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives like 'retrieve_memory' or 'search_by_tag'. It doesn't mention prerequisites, appropriate contexts, or exclusions, leaving the agent to infer usage based solely on the tool name and basic purpose.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
- First observed
retrieve_memory - First observed
search_by_tag - First observed
store_memory
This server cannot be deployed
TDQS
Scored across 3 tools
Each tool has a clearly distinct purpose: retrieve_memory finds memories based on content queries, search_by_tag filters by tags, and store_memory creates new entries. There is no overlap or ambiguity between these three operations.
All tools follow a consistent verb_noun pattern (retrieve_memory, search_by_tag, store_memory) with snake_case throughout. The naming is predictable and uniform across the set.
With only 3 tools, the set feels minimal but functional for a memory service. It covers basic operations (store, retrieve, search), but lacks advanced features like updating or deleting memories, which might be expected in a more comprehensive service.
The tools provide core CRUD-like operations for storing and retrieving memories, but there are notable gaps: no update_memory or delete_memory tools, which limits lifecycle management. Agents can work around this for basic use but may encounter dead ends for modifications.
Maintenance
Related MCP Connectors
Private persistent memory for Claude, ChatGPT & Gemini via MCP - semantic search, zero-code setup.
Persistent memory for AI agents across Claude, ChatGPT and any MCP client.
Persistent memory for Claude Code and Cursor. Stop re-explaining your project every session.
Persistent AI memory shared across Claude, ChatGPT, coding agents, and compatible MCP clients.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceProvides Claude AI with persistent, searchable memory management across sessions using SQL database, semantic analysis with multi-provider LLM support (Anthropic/Ollama), vector search via ChromaDB, and graph-based knowledge relationships through Neo4j integration.1-
- FlicenseAqualityDmaintenanceSupercharges Claude Desktop with persistent semantic memory, sandboxed file I/O, live web search, and local emotional intelligence using a local ChromaDB and Hugging Face model.6-
- AlicenseNot gradedqualityDmaintenanceProvides persistent, searchable memory for Claude Code using local SQLite, semantic embeddings, and full-text search, enabling Claude to recall and retrieve context across sessions and projects without external services.15 npm4MIT
- AlicenseNot gradedqualityFmaintenanceProvides Claude with long-term memory by indexing conversation history, enabling semantic search, decision tracking, and cross-project search.18 npm29MIT