daryl-memories
Provides shared GraphRAG memory for Hermes agents, including tools to store episodes, check processing status, run hybrid recall, expand entity context, resolve duplicate entities, and forget memories.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@daryl-memoriesrecall the key findings from the prototype testing"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Mnemosyne
Shared GraphRAG memory MCP server for 3 Hermes agents. Fully local -- no data leaves the network.
Architecture
Machine 1 (HOST: this PC) Machine 2 Machine 3
+--------------------------+ +-------------+ +-------------+
| Neo4j (Docker) | | Hermes | | Hermes |
| Ollama (native) |<---| Agent | | Agent |
| MCP server (Python) | MCP| (client) | | (client) |
| Port 8080 | +-------------+ +-------------+
+--------------------------+Machine 1 runs everything. Machines 2 & 3 are pure MCP clients.
Neo4j handles graph + vector + full-text in one container.
Ollama runs natively (not Docker) for simplicity.
MCP is Hermes's native protocol -- agents get memory tools as first-class capabilities.
Related MCP server: Knowledge Graph Memory Server
Quick Start
# 1. Clone and configure
git clone https://github.com/DarylAndrian/Mnemosyne.git
cd Mnemosyne
cp .env.example .env
# Edit .env with your passwords
# 2. Start Neo4j
docker compose up -d
# 3. Install Python deps
uv venv .venv
uv pip install -r requirements.txt
# 4. Start the server
python -m server.mainThe server starts on http://0.0.0.0:8080/mcp. Health check at /health.
Memory Graph Frontend
An Obsidian-style graph viewer is served at http://<HOST_IP>:8080/ (same port as MCP).
Force-directed graph of entities, colored by type
Click a node: aliases, facts, episodes, 1-3 hop neighborhood
Double-click: expand neighborhood
Recall box: hybrid RAG search (same pipeline as the
recalltool)Filters by entity type, node limit, include episodes
Paste
MCP_API_KEYor an agent token in the top-right to unlock API calls
No build step and no internet needed -- plain HTML/JS with vis-network vendored in
frontend/vendor/. The API endpoints (/api/graph, /api/entity, /api/recall,
/api/episodes, /api/stats) require the same Bearer token as MCP; static files are public.
Environment Variables
Variable | Default | Description |
|
| Neo4j bolt URI |
|
| Neo4j username |
| (required) | Neo4j password |
|
| Ollama API URL |
|
| LLM for entity extraction |
|
| Embedding model (768-dim) |
|
| Bind address |
|
| Listen port |
| (required) | Shared API key for auth |
MCP Tools
remember(content, agent_id, session_id?, tags?, request_id?)
Store a memory. Returns immediately with episode_id and processing: true;
entity/relationship/fact extraction and embedding run in the background.
Detects conflicts with existing facts.
episode_status(episode_id)
Check the background processing status of a stored episode:
pending, processed, failed, or not_found.
Latency & Retries (read this before integrating)
rememberreturns in milliseconds. The DB write happens first; the slow part (LLM extraction + embeddings, 10-20s on CPU-only hosts) runs in the background. No client timeout issues.Always send a
request_id(a UUID per logical store intent). If your client times out and retries with the samerequest_id, the server returns the original episode (deduplicated: true) instead of creating a duplicate. Idempotency keys never expire.Automatic dedup: identical content (SHA-256 of normalized text) or near-identical content (>= 97% token overlap) stored within a 5-minute window returns the existing episode with
deduplicated: true.Use
episode_statusto poll background completion if your agent needs extracted entities/facts to be ready; full-text recall works immediately, vector recall once embedding completes.
recall(query, top_k?, agent_id?)
Hybrid RAG search: vector similarity + keyword + graph neighborhood expansion. Fused with recency and access-count boosts.
context(entity_name, depth_limit?)
Graph neighborhood traversal. Returns all edges connected to an entity within 1-3 hops.
resolve(entity_a, entity_b)
Merge duplicate entities. Re-points all edges, keeps both names as aliases.
forget(memory_id)
Soft-delete an episode. Preserves provenance.
Agent Configuration
Add to your Hermes config:
{
"mcpServers": {
"mnemosyne": {
"url": "http://<HOST_IP>:8080/mcp",
"headers": {
"Authorization": "Bearer <MCP_API_KEY>"
}
}
}
}Infrastructure
Neo4j 5.26 (Community) -- graph + vector + full-text in one container
Ollama 0.32+ -- qwen2.5:3b (extraction) + nomic-embed-text (embeddings)
Python 3.11+ -- FastMCP server with Streamable HTTP transport
Docker Compose -- Neo4j only (Ollama stays native)
Development
# Run integration tests (requires live stack)
.venv/Scripts/python.exe -c "from tests.test_integration import *; ..."License
MIT
Auto-deploy
Pushes to main deploy automatically via GitHub webhook
(POST /api/deploy/webhook, HMAC-SHA256 verified). A detached finisher
pulls, restarts the pm2 process, polls /health, and auto-rolls back to
the previous commit if the new code fails its health check within 120s.
Check GET /api/deploy/status for the last deploy state.
This server cannot be deployed
Maintenance
Related MCP Connectors
Shared long-term memory for AI agents: save and recall context as a searchable knowledge graph.
Persistent memory and knowledge graphs for AI agents. Hybrid search, context checkpoints, and more.
Persistent memory for AI agents. Semantic search, memory graph, W3C DID identity.
- GoMindOAuthcom.gominddb
Persistent knowledge graph for AI agents. Remember, recall, and forget facts.
Related MCP Servers
FlicenseNot gradedqualityNot gradedmaintenanceProvides local-first memory storage and retrieval with automatic embedding, vector search, and knowledge graph capabilities. Enables agents to store memories locally and retrieve relevant context through hybrid search with optional Neo4j graph traversal.-- AlicenseNot gradedqualityDmaintenanceA persistent memory server that implements a local knowledge graph using the Kuzu embedded database to store entities, relationships, and observations. It enables AI models to maintain structured long-term context through searchable nodes and comprehensive tag-based organization.8 npmMIT
- AlicenseNot gradedqualityBmaintenanceA lightweight, powerful local memory server for AI agents supporting text, entities, and relations. Enables persistent codebase understanding and user preference management.145 npm56MIT
- AlicenseNot gradedqualityDmaintenanceProvides persistent knowledge graph memory for AI agents, enabling them to store, recall, and query facts about people, projects, and relationships across sessions.MIT