Skip to main content
Glama
harshithnm110

mcp-memory-agent

mcp-memory-agent

This project is my attempt on giving real memory for local LLMs instead of hoping for context window to last a bit longer as the conversation drags on for smaller quantized models.

Has a four-layer memory architecture for LLM agents : working, episodic, semantic, and a routing layer that decides which tier a query actually needs. Also exposed as MCP tools so any MCP-compatible client can plug into it. Runs entirely locally: Ollama for inference, ChromaDB for vector storage.

Tested standalone via a custom async chat client and as an MCP connector in Claude Desktop.

Architecture diagram


Why this exists

Not every question or a conversation needs the reference of full chat memory. What if LLMs could remember certain n "events" of conversation to be later referenced for a particular user prompt.

For example, math questions don't need any memory. "What did I ask last time" needs a specific event and "what should i cook tonight" needs a preference instead of an entire search through raw chat logs from three weeks ago. This project treats that distinction as the core engineering problem, not an afterthought — which is also why why it's a router, a consolidation pipeline, and two separate memory stores, not one vector DB with a single search() call.

Layer

Lifetime

Storage

Retrieval

Working memory

One process

In-RAM list

Sliding window (last N turns)

Episodic memory

Permanent

ChromaDB (disk)

Semantic similarity

Semantic memory

Permanent

ChromaDB (disk, separate collection)

Semantic similarity, distilled facts only

Router

Per-query

—

Heuristic + LLM fallback classification

Related MCP server: Engram-Mem

MCP Tools

Exposed via memory_mcp_server.py (FastMCP, stdio transport) — any MCP client can call these, not just this repo's demo chat loop:

Tool

What it does

store_memory_tool

Store a fact/event as episodic memory

retrieve_relevant_memory_tool

Similarity search over raw episodic memory

consolidate_memories_tool

Distill accumulated episodic memories into standing semantic facts

retrieve_semantic_memory_tool

Similarity search over consolidated semantic facts

route_query_tool

Classify a query as NONE / EPISODIC / SEMANTIC / BOTH

Proof: Cross-Session Recall

The core claim of this project — a fact told to the agent dies with working memory, survives in episodic memory, and gets distilled into semantic memory — demonstrated across a full process restart:

Session 1:

You: hi i am harshith
Bot: Hi Harshith, how can I assist you today?

You: i can't eat vegetarian food, only non veg
Bot: Okay, I can help you find recipes related to non-vegetarian food...

--- consolidation ran → Consolidated 3 facts from 2 episodic memories. ---

You: quit    ← process fully terminated, working memory destroyed

Session 2 — brand new process, zero working memory, nothing re-told:

You: give me dinner suggestions

--- router decision = SEMANTIC ---
Relevant semantic facts:
- User only eats non-vegetarian food.
- User does not eat vegetarian food.
- User's name is Harshith.

Bot: Dinner suggestions for Harshith that are non-vegetarian:
1. Grilled steak with roasted vegetables
2. Chicken stir-fry with mixed vegetables
...

Nothing about Harshith or the dietary constraint exists anywhere in session 2's process memory. The router independently chose SEMANTIC, and the answer came from a consolidated fact and not a raw conversation replay, proving the full pipeline end to end.

CLI recall demo - peanut allergy Memory referenced from chromadb for every new session. CLI recall demo - dietary preference

Claude Desktop Integration

A genuine finding Sometimes, Claude fails to refer the mcp tools unless mentioned as seen in the output sample below.

Memory tool enabled in Claude Desktop Claude Desktop correctly calling memory tools

Design Decisions

Hybrid router, not pure-LLM. An early pure-LLM router (few-shot prompt, 4-way classification) was unreliable on a 7B quantized model, it would anchor on whichever example answer appeared last in the prompt, regardless of the actual query. Fixed with a fast heuristic pre-filter (regex for arithmetic/definitions → NONE, keyword matching for temporal phrasing → EPISODIC, preference language → SEMANTIC) and the LLM call reserved for genuinely ambiguous cases only, at temperature=0 for determinism.

Relative paths worked with cli but fails when tested through claude desktop , switching to absolute path.

Tech Stack

  • LLM: Mistral-7B-Instruct (mistral:7b-instruct-q5_K_M) via Ollama — fully local

  • Vector DB: ChromaDB (persistent local client, two collections)

  • Embeddings: sentence-transformers (all-MiniLM-L6-v2)

  • Protocol: MCP Python SDK v1.x (FastMCP)

Quick Start

# 1. Pull the model
ollama pull mistral:7b-instruct-q5_K_M

# 2. Set up the environment
python -m venv venv
venv\Scripts\activate
pip install -r requirements.txt

# 3. Run the demo chat client (auto-launches the MCP server as a subprocess)
python chat_loop.py

Using with Claude Desktop

Add to claude_desktop_config.json (Settings → Developer → Edit Config):

{
  "mcpServers": {
    "episodic_memory": {
      "command": "C:\\path\\to\\venv\\Scripts\\python.exe",
      "args": ["C:\\path\\to\\memory_mcp_server.py"]
    }
  }
}

Use the full path to your venv's python.exe, not a bare "python" — Claude Desktop doesn't launch from an activated shell. Fully restart Claude Desktop (quit from tray, not just close window) after saving.

Limitations / Future Work

  • Two contradictory facts will co-exist instead of being replaced (example: if the user says "I'm allergic to peanuts" in one session and "I'm not allergic to peanuts" in another, both memories co-exist).

  • Evaluation has to be done manually , there's no benchmark to test router or retrieval accuracy.

  • Consolidation re-processes all episodic memories each run rather than tracking what's already been consolidated.

  • Single-user only, no session/user scoping in the data model.

  • No similarity-distance threshold, retrieval always returns top-k regardless of how weak the closest match is.

Project Structure

memory_store.py             # Core logic: ChromaDB client, embeddings, store/retrieve/consolidate/route
memory_mcp_server.py        # MCP tool wrappers (FastMCP, stdio transport)
chat_loop.py                # Demo client: async chat loop wired to the MCP server
test_mcp_client.py          # Standalone MCP client for manually testing tool calls
requirements.txt            # Pinned dependencies
architecture_diagram/       # System architecture SVG
sample_output/              # Screenshots: CLI recall demo + Claude Desktop integration

Related MCP Connectors

Related MCP Servers