mcp-memory-agent
README.md
# mcp-memory-agent
**This project is my attempt on giving real memory for local LLMs instead of hoping for context window to last a bit longer as the conversation drags on for smaller quantized models.**
Has a four-layer memory architecture for LLM agents : working, episodic, semantic, and a routing layer that decides which tier a query actually needs. Also exposed as MCP tools so any MCP-compatible client can plug into it. Runs entirely locally: Ollama for inference, ChromaDB for vector storage.
Tested standalone via a custom async chat client and as an MCP connector in Claude Desktop.

---
## Why this exists
Not every question or a conversation needs the reference of full chat memory. What if LLMs could remember certain n "events" of conversation to be later referenced for a particular user prompt.
For example, math questions don't need any memory. "What did I ask last time" needs a specific event and "what should i cook tonight" needs a preference instead of an entire search through raw chat logs from three weeks ago. This project treats that distinction as the core engineering problem, not an afterthought — which is also why why it's a router, a consolidation pipeline, and two separate memory stores, not one vector DB with a single search() call.
| Layer | Lifetime | Storage | Retrieval |
|---|---|---|---|
| **Working memory** | One process | In-RAM list | Sliding window (last N turns) |
| **Episodic memory** | Permanent | ChromaDB (disk) | Semantic similarity |
| **Semantic memory** | Permanent | ChromaDB (disk, separate collection) | Semantic similarity, distilled facts only |
| **Router** | Per-query | — | Heuristic + LLM fallback classification |
## MCP Tools
Exposed via `memory_mcp_server.py` (FastMCP, stdio transport) — any MCP client can call these, not just this repo's demo chat loop:
| Tool | What it does |
|---|---|
| `store_memory_tool` | Store a fact/event as episodic memory |
| `retrieve_relevant_memory_tool` | Similarity search over raw episodic memory |
| `consolidate_memories_tool` | Distill accumulated episodic memories into standing semantic facts |
| `retrieve_semantic_memory_tool` | Similarity search over consolidated semantic facts |
| `route_query_tool` | Classify a query as `NONE` / `EPISODIC` / `SEMANTIC` / `BOTH` |
## Proof: Cross-Session Recall
The core claim of this project — a fact told to the agent dies with working memory, survives in episodic memory, and gets distilled into semantic memory — demonstrated across a full process restart:
**Session 1:**
```
You: hi i am harshith
Bot: Hi Harshith, how can I assist you today?
You: i can't eat vegetarian food, only non veg
Bot: Okay, I can help you find recipes related to non-vegetarian food...
--- consolidation ran → Consolidated 3 facts from 2 episodic memories. ---
You: quit ← process fully terminated, working memory destroyed
```
**Session 2 — brand new process, zero working memory, nothing re-told:**
```
You: give me dinner suggestions
--- router decision = SEMANTIC ---
Relevant semantic facts:
- User only eats non-vegetarian food.
- User does not eat vegetarian food.
- User's name is Harshith.
Bot: Dinner suggestions for Harshith that are non-vegetarian:
1. Grilled steak with roasted vegetables
2. Chicken stir-fry with mixed vegetables
...
```
Nothing about Harshith or the dietary constraint exists anywhere in session 2's process memory. The router independently chose `SEMANTIC`, and the answer came from a *consolidated* fact and not a raw conversation replay, proving the full pipeline end to end.

Memory referenced from chromadb for every new session.

## Claude Desktop Integration
**A genuine finding** Sometimes, Claude fails to refer the mcp tools unless mentioned as seen in the output sample below.


## Design Decisions
**Hybrid router, not pure-LLM.** An early pure-LLM router (few-shot prompt, 4-way classification) was unreliable on a 7B quantized model, it would anchor on whichever example answer appeared last in the prompt, regardless of the actual query. Fixed with a fast heuristic pre-filter (regex for arithmetic/definitions → `NONE`, keyword matching for temporal phrasing → `EPISODIC`, preference language → `SEMANTIC`) and the LLM call reserved for genuinely ambiguous cases only, at `temperature=0` for determinism.
**Relative paths worked with cli but fails when tested through claude desktop , switching to absolute path.**
## Tech Stack
- **LLM**: Mistral-7B-Instruct (`mistral:7b-instruct-q5_K_M`) via [Ollama](https://ollama.com) — fully local
- **Vector DB**: [ChromaDB](https://www.trychroma.com/) (persistent local client, two collections)
- **Embeddings**: `sentence-transformers` (`all-MiniLM-L6-v2`)
- **Protocol**: [MCP Python SDK](https://github.com/modelcontextprotocol/python-sdk) v1.x (`FastMCP`)
## Quick Start
```bash
# 1. Pull the model
ollama pull mistral:7b-instruct-q5_K_M
# 2. Set up the environment
python -m venv venv
venv\Scripts\activate
pip install -r requirements.txt
# 3. Run the demo chat client (auto-launches the MCP server as a subprocess)
python chat_loop.py
```
### Using with Claude Desktop
Add to `claude_desktop_config.json` (Settings → Developer → Edit Config):
```json
{
"mcpServers": {
"episodic_memory": {
"command": "C:\\path\\to\\venv\\Scripts\\python.exe",
"args": ["C:\\path\\to\\memory_mcp_server.py"]
}
}
}
```
Use the full path to your venv's `python.exe`, not a bare `"python"` — Claude Desktop doesn't launch from an activated shell. Fully restart Claude Desktop (quit from tray, not just close window) after saving.
## Limitations / Future Work
- Two contradictory facts will co-exist instead of being replaced (example: if the user says "I'm allergic to peanuts" in one session and "I'm not allergic to peanuts" in another, both memories co-exist).
- Evaluation has to be done manually , there's no benchmark to test router or retrieval accuracy.
- Consolidation re-processes all episodic memories each run rather than tracking what's already been consolidated.
- Single-user only, no session/user scoping in the data model.
- No similarity-distance threshold, retrieval always returns top-k regardless of how weak the closest match is.
## Project Structure
```
memory_store.py # Core logic: ChromaDB client, embeddings, store/retrieve/consolidate/route
memory_mcp_server.py # MCP tool wrappers (FastMCP, stdio transport)
chat_loop.py # Demo client: async chat loop wired to the MCP server
test_mcp_client.py # Standalone MCP client for manually testing tool calls
requirements.txt # Pinned dependencies
architecture_diagram/ # System architecture SVG
sample_output/ # Screenshots: CLI recall demo + Claude Desktop integration
```
This server cannot be deployed
Maintenance
ActivityMaintained
ResponsivenessNo issues