universal-memory-mcp
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@universal-memory-mcpStore that I finished reading Dune yesterday."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
universal-memory-mcp
Persistent memory MCP server for single and multi-agent LLM systems. Gives AI agents long-term memory backed by SQLite with hybrid keyword + semantic search.
Features
Memory types: episodic (events/logs), semantic (facts/knowledge), procedural (how-to/workflows)
Hybrid search: FTS5 keyword search + cosine similarity over embeddings, with configurable weights
Knowledge graph: directed links between memories (caused_by, related_to, contradicts, supports, follows) with BFS traversal
Session checkpoints: save/restore agent state across conversations
Multi-agent support: scope memories by agent_id, session_id, or share globally
Optimistic locking: safe concurrent updates with version conflict detection
Pluggable embeddings: HuggingFace transformers (in-process) or llama-server (external HTTP)
Related MCP server: NeuralVaultCore
Install
uv syncUsage
Run as an MCP server (stdio transport):
uv run python server.pyOr via the wrapper script:
./run.shCLI
This package also installs a memory CLI:
uv run memory doctorIngest Claude/Codex logs
The log ingester reads local Claude/Codex JSONL/text logs, redacts common secrets, extracts durable memories, and stores only the extracted memories with log provenance metadata. Raw logs are not stored.
Dry-run first:
uv run memory ingest-logs codex --dry-run --root ~/.codex/sessions --extractor heuristic
uv run memory ingest-logs claude --dry-run --root ~/.claude/projects --extractor heuristicUse a local llama-server OpenAI-compatible chat endpoint for extraction:
llama-server --model /path/to/chat-model.gguf --port 8080
uv run memory ingest-logs all --llm-url http://localhost:8080 --llm-model localBy default the CLI does not send max_tokens to the chat endpoint, which avoids
truncating extraction responses from thinking models. Set an explicit cap only
when you need one:
uv run memory ingest-logs codex --llm-max-tokens 4096To also compute embeddings through a local embedding server:
llama-server --model embeddinggemma-300m-Q4_0.gguf --port 8787 --embedding --ctx-size 512
MEMORY_ENABLE_EMBEDDINGS=true \
MEMORY_EMBEDDING_BACKEND=llama-server \
MEMORY_EMBEDDING_DIMENSION=768 \
MEMORY_LLAMA_SERVER_URL=http://localhost:8787 \
uv run memory ingest-logs codex --llm-url http://localhost:8080Incremental checkpoints are stored in SQLite, so later runs only consume new events. For polling:
uv run memory watch-logs all --interval 30 --llm-url http://localhost:8080Claude Code config
Add to your MCP settings (~/.claude/settings.json or project .mcp.json):
{
"mcpServers": {
"memory": {
"command": "uv",
"args": ["run", "--directory", "/path/to/universal-memory-mcp", "python", "server.py"]
}
}
}Configuration
All settings via environment variables (prefix MEMORY_):
Variable | Default | Description |
|
| SQLite database path |
|
|
|
|
| HuggingFace model name |
|
| Embedding vector size |
|
| llama-server endpoint |
|
| Set |
|
| Hybrid search keyword weight |
|
| Hybrid search semantic weight |
|
| Minimum cosine similarity for semantic-channel recall results (0 = disabled). Keyword matches always survive. |
|
| Truncate recalled contents to this many chars (0 = disabled). Full text via exact fetch or |
Using llama-server backend
For lower memory usage with a GGUF model:
llama-server --model embeddinggemma-300m-Q4_0.gguf --port 8787 --embedding --ctx-size 512
MEMORY_EMBEDDING_BACKEND=llama-server MEMORY_EMBEDDING_DIMENSION=768 uv run python server.pyMCP Tools
The surface is deliberately small — five tools, tiered by call frequency:
Tool | Description |
| One retrieval tool, three selectors: |
| Store a memory with type, agent/session scope, importance; optional |
| Update with optimistic locking ( |
|
|
|
|
Errors are raised as MCP tool errors (isError: true), never returned as data.
Breaking changes in 0.2.0
get_memory, delete_memory, link_memories, get_linked_memories, create_session, checkpoint_session, restore_session, get_stats, extract_entities, get_entity_neighbors, check_contradictions, dream_status, and run_dream_jobs were folded into the five tools above. recall_memories now returns {mode, count, memories} instead of a bare list, truncates long contents by default, and applies a semantic relevance cutoff (MEMORY_RECALL_MIN_RELEVANCE).
Tests
uv run pytestLicense
This server cannot be installed
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Servers
- Alicense-qualityDmaintenanceA local-first MCP memory server providing persistent, searchable memory for AI agents, powered by SQLite.Last updated381Apache 2.0
- Alicense-qualityCmaintenanceAn MCP server that provides persistent long-term memory for AI agents via local SQLite storage with low token overhead, enabling memory storage, retrieval, and management across sessions.Last updated1MIT
- Flicense-qualityDmaintenanceAn MCP server for managing persistent AI memory using hybrid search (keyword + semantic vector) with SQLite storage and offline-first local embeddings.Last updated
- AlicenseAqualityBmaintenanceA long-term memory MCP server for AI agents that stores memories (facts, decisions, etc.) in a single SQLite database with hybrid search and full edit history, ensuring consistency across sessions.Last updated24MIT
Related MCP Connectors
Cloud-hosted MCP server for durable AI memory
Person-owned, portable AI memory as a remote MCP server, readable and writable by any MCP client.
User-owned memory for AI agents, Copilot, Claude, IDEs, CLIs, and chat apps over remote MCP.
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/hz1ulqu01gmnZH4/universal-memory-mcp'
If you have feedback or need assistance with the MCP directory API, please join our Discord server