mcp-memory-agent
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@mcp-memory-agentgive me dinner suggestions"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
mcp-memory-agent
This project is my attempt on giving real memory for local LLMs instead of hoping for context window to last a bit longer as the conversation drags on for smaller quantized models.
Has a four-layer memory architecture for LLM agents : working, episodic, semantic, and a routing layer that decides which tier a query actually needs. Also exposed as MCP tools so any MCP-compatible client can plug into it. Runs entirely locally: Ollama for inference, ChromaDB for vector storage.
Tested standalone via a custom async chat client and as an MCP connector in Claude Desktop.
Why this exists
Not every question or a conversation needs the reference of full chat memory. What if LLMs could remember certain n "events" of conversation to be later referenced for a particular user prompt.
For example, math questions don't need any memory. "What did I ask last time" needs a specific event and "what should i cook tonight" needs a preference instead of an entire search through raw chat logs from three weeks ago. This project treats that distinction as the core engineering problem, not an afterthought — which is also why why it's a router, a consolidation pipeline, and two separate memory stores, not one vector DB with a single search() call.
Layer | Lifetime | Storage | Retrieval |
Working memory | One process | In-RAM list | Sliding window (last N turns) |
Episodic memory | Permanent | ChromaDB (disk) | Semantic similarity |
Semantic memory | Permanent | ChromaDB (disk, separate collection) | Semantic similarity, distilled facts only |
Router | Per-query | — | Heuristic + LLM fallback classification |
Related MCP server: Engram-Mem
MCP Tools
Exposed via memory_mcp_server.py (FastMCP, stdio transport) — any MCP client can call these, not just this repo's demo chat loop:
Tool | What it does |
| Store a fact/event as episodic memory |
| Similarity search over raw episodic memory |
| Distill accumulated episodic memories into standing semantic facts |
| Similarity search over consolidated semantic facts |
| Classify a query as |
Proof: Cross-Session Recall
The core claim of this project — a fact told to the agent dies with working memory, survives in episodic memory, and gets distilled into semantic memory — demonstrated across a full process restart:
Session 1:
You: hi i am harshith
Bot: Hi Harshith, how can I assist you today?
You: i can't eat vegetarian food, only non veg
Bot: Okay, I can help you find recipes related to non-vegetarian food...
--- consolidation ran → Consolidated 3 facts from 2 episodic memories. ---
You: quit ← process fully terminated, working memory destroyedSession 2 — brand new process, zero working memory, nothing re-told:
You: give me dinner suggestions
--- router decision = SEMANTIC ---
Relevant semantic facts:
- User only eats non-vegetarian food.
- User does not eat vegetarian food.
- User's name is Harshith.
Bot: Dinner suggestions for Harshith that are non-vegetarian:
1. Grilled steak with roasted vegetables
2. Chicken stir-fry with mixed vegetables
...Nothing about Harshith or the dietary constraint exists anywhere in session 2's process memory. The router independently chose SEMANTIC, and the answer came from a consolidated fact and not a raw conversation replay, proving the full pipeline end to end.
Memory referenced from chromadb for every new session.

Claude Desktop Integration
A genuine finding Sometimes, Claude fails to refer the mcp tools unless mentioned as seen in the output sample below.

Design Decisions
Hybrid router, not pure-LLM. An early pure-LLM router (few-shot prompt, 4-way classification) was unreliable on a 7B quantized model, it would anchor on whichever example answer appeared last in the prompt, regardless of the actual query. Fixed with a fast heuristic pre-filter (regex for arithmetic/definitions → NONE, keyword matching for temporal phrasing → EPISODIC, preference language → SEMANTIC) and the LLM call reserved for genuinely ambiguous cases only, at temperature=0 for determinism.
Relative paths worked with cli but fails when tested through claude desktop , switching to absolute path.
Tech Stack
LLM: Mistral-7B-Instruct (
mistral:7b-instruct-q5_K_M) via Ollama — fully localVector DB: ChromaDB (persistent local client, two collections)
Embeddings:
sentence-transformers(all-MiniLM-L6-v2)Protocol: MCP Python SDK v1.x (
FastMCP)
Quick Start
# 1. Pull the model
ollama pull mistral:7b-instruct-q5_K_M
# 2. Set up the environment
python -m venv venv
venv\Scripts\activate
pip install -r requirements.txt
# 3. Run the demo chat client (auto-launches the MCP server as a subprocess)
python chat_loop.pyUsing with Claude Desktop
Add to claude_desktop_config.json (Settings → Developer → Edit Config):
{
"mcpServers": {
"episodic_memory": {
"command": "C:\\path\\to\\venv\\Scripts\\python.exe",
"args": ["C:\\path\\to\\memory_mcp_server.py"]
}
}
}Use the full path to your venv's python.exe, not a bare "python" — Claude Desktop doesn't launch from an activated shell. Fully restart Claude Desktop (quit from tray, not just close window) after saving.
Limitations / Future Work
Two contradictory facts will co-exist instead of being replaced (example: if the user says "I'm allergic to peanuts" in one session and "I'm not allergic to peanuts" in another, both memories co-exist).
Evaluation has to be done manually , there's no benchmark to test router or retrieval accuracy.
Consolidation re-processes all episodic memories each run rather than tracking what's already been consolidated.
Single-user only, no session/user scoping in the data model.
No similarity-distance threshold, retrieval always returns top-k regardless of how weak the closest match is.
Project Structure
memory_store.py # Core logic: ChromaDB client, embeddings, store/retrieve/consolidate/route
memory_mcp_server.py # MCP tool wrappers (FastMCP, stdio transport)
chat_loop.py # Demo client: async chat loop wired to the MCP server
test_mcp_client.py # Standalone MCP client for manually testing tool calls
requirements.txt # Pinned dependencies
architecture_diagram/ # System architecture SVG
sample_output/ # Screenshots: CLI recall demo + Claude Desktop integrationThis server cannot be deployed
Maintenance
Related MCP Connectors
Cross-tool persistent memory and context for AI assistants over MCP.
Private-by-default, local-first memory/context/task orchestrator for MCP apps and agents.
Persistent memory for AI agents across Claude, ChatGPT and any MCP client.
- KogniteOAuthdev.kognite
Hosted agent memory: store, search, and recall facts across sessions from any MCP client.
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceProvides a local, persistent long-term memory service for MCP-compatible AI agents, enabling them to store, search, and recall information across sessions.1GPL 3.0
- AlicenseNot gradedqualityCmaintenanceEnables persistent memory for AI agents, combining episodic and semantic memory with LLM reasoning, accessible via MCP.2MIT
- AlicenseNot gradedqualityBmaintenanceProvides persistent semantic memory for AI agents via MCP, enabling them to remember, recall, list, update, and forget memories with vector-based similarity search.ISC
- AlicenseNot gradedqualityBmaintenanceMCP server providing agents with persistent working, episodic, and semantic memory backed by CockroachDB, exposed via tools to store, recall, list, and forget memories.MIT