Skip to main content
Glama
harshithnm110

mcp-memory-agent

README.md
# mcp-memory-agent


**This project is my attempt on giving real memory for local LLMs instead of hoping for context window to last a bit longer as the conversation drags on for smaller quantized models.**

Has a four-layer memory architecture for LLM agents : working, episodic, semantic, and a routing layer that decides which tier a query actually needs. Also exposed as MCP tools so any MCP-compatible client can plug into it. Runs entirely locally: Ollama for inference, ChromaDB for vector storage.

Tested standalone via a custom async chat client and as an MCP connector in Claude Desktop.

![Architecture diagram](architecture_diagram/diagram.svg)

---

## Why this exists

Not every question or a conversation needs the reference of full chat memory. What if LLMs could remember certain n "events" of conversation to be later referenced for a particular user prompt.

For example, math questions don't need any memory. "What did I ask last time" needs a specific event and "what should i cook tonight" needs a preference instead of an entire search through raw chat logs from three weeks ago. This project treats that distinction as the core engineering problem, not an afterthought — which is also why why it's a router, a consolidation pipeline, and two separate memory stores, not one vector DB with a single search() call.

| Layer | Lifetime | Storage | Retrieval |
|---|---|---|---|
| **Working memory** | One process | In-RAM list | Sliding window (last N turns) |
| **Episodic memory** | Permanent | ChromaDB (disk) | Semantic similarity |
| **Semantic memory** | Permanent | ChromaDB (disk, separate collection) | Semantic similarity, distilled facts only |
| **Router** | Per-query | — | Heuristic + LLM fallback classification |

## MCP Tools

Exposed via `memory_mcp_server.py` (FastMCP, stdio transport) — any MCP client can call these, not just this repo's demo chat loop:

| Tool | What it does |
|---|---|
| `store_memory_tool` | Store a fact/event as episodic memory |
| `retrieve_relevant_memory_tool` | Similarity search over raw episodic memory |
| `consolidate_memories_tool` | Distill accumulated episodic memories into standing semantic facts |
| `retrieve_semantic_memory_tool` | Similarity search over consolidated semantic facts |
| `route_query_tool` | Classify a query as `NONE` / `EPISODIC` / `SEMANTIC` / `BOTH` |

## Proof: Cross-Session Recall

The core claim of this project — a fact told to the agent dies with working memory, survives in episodic memory, and gets distilled into semantic memory — demonstrated across a full process restart:

**Session 1:**
```
You: hi i am harshith
Bot: Hi Harshith, how can I assist you today?

You: i can't eat vegetarian food, only non veg
Bot: Okay, I can help you find recipes related to non-vegetarian food...

--- consolidation ran → Consolidated 3 facts from 2 episodic memories. ---

You: quit    ← process fully terminated, working memory destroyed
```

**Session 2 — brand new process, zero working memory, nothing re-told:**
```
You: give me dinner suggestions

--- router decision = SEMANTIC ---
Relevant semantic facts:
- User only eats non-vegetarian food.
- User does not eat vegetarian food.
- User's name is Harshith.

Bot: Dinner suggestions for Harshith that are non-vegetarian:
1. Grilled steak with roasted vegetables
2. Chicken stir-fry with mixed vegetables
...
```

Nothing about Harshith or the dietary constraint exists anywhere in session 2's process memory. The router independently chose `SEMANTIC`, and the answer came from a *consolidated* fact and not a raw conversation replay, proving the full pipeline end to end.

![CLI recall demo - peanut allergy](sample_output/cli_recall_demo_1.png)
Memory referenced from chromadb for every new session.
![CLI recall demo - dietary preference](sample_output/cli_recall_demo_2.png)

## Claude Desktop Integration

**A genuine finding** Sometimes, Claude fails to refer the mcp tools unless mentioned as seen in the output sample below.

![Memory tool enabled in Claude Desktop](sample_output/claude_desktop_connector.png)
![Claude Desktop correctly calling memory tools](sample_output/claude_desktop_working.png)

## Design Decisions

**Hybrid router, not pure-LLM.** An early pure-LLM router (few-shot prompt, 4-way classification) was unreliable on a 7B quantized model, it would anchor on whichever example answer appeared last in the prompt, regardless of the actual query. Fixed with a fast heuristic pre-filter (regex for arithmetic/definitions → `NONE`, keyword matching for temporal phrasing → `EPISODIC`, preference language → `SEMANTIC`) and the LLM call reserved for genuinely ambiguous cases only, at `temperature=0` for determinism.

**Relative paths worked with cli but fails when tested through claude desktop , switching to absolute path.** 

## Tech Stack

- **LLM**: Mistral-7B-Instruct (`mistral:7b-instruct-q5_K_M`) via [Ollama](https://ollama.com) — fully local
- **Vector DB**: [ChromaDB](https://www.trychroma.com/) (persistent local client, two collections)
- **Embeddings**: `sentence-transformers` (`all-MiniLM-L6-v2`)
- **Protocol**: [MCP Python SDK](https://github.com/modelcontextprotocol/python-sdk) v1.x (`FastMCP`)

## Quick Start

```bash
# 1. Pull the model
ollama pull mistral:7b-instruct-q5_K_M

# 2. Set up the environment
python -m venv venv
venv\Scripts\activate
pip install -r requirements.txt

# 3. Run the demo chat client (auto-launches the MCP server as a subprocess)
python chat_loop.py
```

### Using with Claude Desktop

Add to `claude_desktop_config.json` (Settings → Developer → Edit Config):

```json
{
  "mcpServers": {
    "episodic_memory": {
      "command": "C:\\path\\to\\venv\\Scripts\\python.exe",
      "args": ["C:\\path\\to\\memory_mcp_server.py"]
    }
  }
}
```

Use the full path to your venv's `python.exe`, not a bare `"python"` — Claude Desktop doesn't launch from an activated shell. Fully restart Claude Desktop (quit from tray, not just close window) after saving.

## Limitations / Future Work

- Two contradictory facts will co-exist instead of being replaced (example: if the user says "I'm allergic to peanuts" in one session and "I'm not allergic to peanuts" in another, both memories co-exist).
- Evaluation has to be done manually , there's no benchmark to test router or retrieval accuracy.
- Consolidation re-processes all episodic memories each run rather than tracking what's already been consolidated.
- Single-user only, no session/user scoping in the data model.
- No similarity-distance threshold, retrieval always returns top-k regardless of how weak the closest match is.

## Project Structure

```
memory_store.py             # Core logic: ChromaDB client, embeddings, store/retrieve/consolidate/route
memory_mcp_server.py        # MCP tool wrappers (FastMCP, stdio transport)
chat_loop.py                # Demo client: async chat loop wired to the MCP server
test_mcp_client.py          # Standalone MCP client for manually testing tool calls
requirements.txt            # Pinned dependencies
architecture_diagram/       # System architecture SVG
sample_output/              # Screenshots: CLI recall demo + Claude Desktop integration
```