lightrag-memory-mcp
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@lightrag-memory-mcpsearch memory for LLM provider decision"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
lightrag-memory-mcp
English | Русский
Long-term memory for your AI coding agents that answers in a second, not a minute.
lightrag-memory-mcp is a small MCP server that connects Claude Code, Codex, or any other MCP client to a LightRAG knowledge base. It gives your agent the facts and skips the slow part.
Wrapper that calls | lightrag-memory-mcp | |
Memory lookup | 17 s median, 49 s p90, up to 164 s | 0.2–2.5 s |
Lookups lost to client timeouts | ~1 in 4 (30 s limit) | none in our tests |
Reading memory when your LLM is down | fails | still works (falls back to a mode that uses no LLM) |
What the agent gets | a summary written by another LLM | original notes with sources, plus the graph |
Measured on a real-world knowledge base: ~37k graph nodes, ~1.5k documents, LightRAG 1.5.4.
The problem
You give your agent a LightRAG memory so it remembers decisions, context, and the reasons behind them. Then this happens:
Every lookup is slow. Most wrappers call LightRAG's
/query, so LightRAG runs an LLM to write a polished answer. That takes 15–60 seconds, sometimes minutes.Lookups time out. Wrappers often hardcode a 30-second limit. In our logs, about one query in four went past it, so the agent got an error instead of memory.
You pay for two LLMs doing one job. LightRAG's LLM summarizes the facts, and then your agent's LLM reads that summary and writes its own answer. The first step adds latency and cost, and it loses details.
Your memory depends on your LLM provider. If the API is rate-limited, down, or blocked in your region, the agent loses its memory as well.
Related MCP server: codebase-wiki
The fix
Your agent is already a capable LLM. It doesn't need a summary. It needs the facts.
⚡ Fast search with no answer generation.
memory_searchcalls/query/dataand returns entities, relations, and original text chunks directly.🛟 Keeps working without an LLM. If the primary search mode fails (LLM timeout, 5xx,
status=failure), the server retries innaivemode, which never calls an LLM. The response tells the agent that a fallback happened.🎯 Context-friendly output. Raw
/query/dataresponses run to hundreds of kilobytes. The server fits each result into a budget (16,000 characters by default). Original notes with their sources come first, because they hold the dates and the reasons. The graph follows, one line per entity or relation.✍️ Safe writes. LightRAG rejects a repeated source name with HTTP 409, and agents reuse names all the time.
memory_saveadds a timestamp to every source name, so no note is silently lost.🌍 English or Russian. Tool descriptions and responses are in English by default. Set
LIGHTRAG_MCP_LANG=rufor Russian.🪶 Small and auditable. One ~400-line Python file, two pinned dependencies, 20 tests.
How it works
flowchart LR
A["AI agent<br/>(Claude Code, Codex, …)"] -- memory_search --> M["lightrag-memory-mcp"]
M -- "POST /query/data<br/>mode = mix" --> L[("LightRAG")]
L -. "LLM timeout / 5xx / failure" .-> M
M -- "retry: mode = naive<br/>(no LLM call)" --> L
M -- "compact result:<br/>notes + sources first, then graph" --> AHere is what your agent gets back from memory_search (shortened):
[memory: mode mix, 0.4 s]
## Note fragments
[decision-2026-09-07-llm-provider--20260907-181502-a1b2.md]
Date: 2026-09-07. Decision: model A stays the primary model.
WHY: model B gives different results from run to run on the same document…
## Entities
- LLM Provider Layer (concept): the component that picks the LLM provider…
## Relations
- DocService → LLM Provider Layer: uses, depends onIf the LLM is down, the header shows it and the search still returns results:
[memory: mode naive, 0.1 s; fell back to naive: timeout 20 s on /query/data]Quick start
You need a running LightRAG server (default http://localhost:9621) and uv.
git clone https://github.com/p-belov/lightrag-memory-mcp.git
cd lightrag-memory-mcp
uv venv .venv && uv pip install --python .venv/bin/python -r requirements.txtClaude Code
claude mcp add-json lightrag '{"type":"stdio","command":"'"$PWD"'/.venv/bin/python","args":["'"$PWD"'/lightrag_memory_mcp.py"],"env":{"LIGHTRAG_URL":"http://localhost:9621"}}' -s userCodex
Add to ~/.codex/config.toml (use absolute paths):
[mcp_servers.lightrag]
command = "/path/to/lightrag-memory-mcp/.venv/bin/python"
args = ["/path/to/lightrag-memory-mcp/lightrag_memory_mcp.py"]
tool_timeout_sec = 200 # memory_answer may wait up to 180 s for the LLM
[mcp_servers.lightrag.env]
LIGHTRAG_URL = "http://localhost:9621"Tell your agent to use it
Add this to CLAUDE.md or AGENTS.md:
Before answering questions about projects → call lightrag:memory_search.
After a decision or an important fact → call lightrag:memory_save
(include the date, what was decided, and why).Tools
Tool | What it does | LightRAG endpoint |
| Finds facts without generating an answer. Falls back to |
|
| Saves a decision or fact and returns a |
|
| Asks LightRAG's LLM for a written answer. Slow; use it only when you need a summary of a large area. |
|
| Server health, pipeline state, and the status of a save by |
|
Configuration
All settings are environment variables.
Variable | Default | Purpose |
|
| LightRAG server address |
| — | API key, if your server has auth enabled (sent as |
|
| Default search mode: |
|
| Seconds to wait for the primary mode before falling back to |
|
| Timeout for |
|
| Character budget for one |
| generic text | Describes your memory to the agent, for example which projects it covers |
|
| Language of tool descriptions and responses: |
FAQ
Does this replace LightRAG? No. It is a thin client. LightRAG still stores the data and builds the knowledge graph.
Do I still need an LLM?
LightRAG needs one to write memory, because it extracts the graph with an LLM. Reading doesn't need one. mix and hybrid use the LLM only to extract keywords from the query, and if that fails, the server switches to naive, which makes no LLM calls.
Can I still get a written answer?
Yes, call memory_answer. It is slow by design, so keep it for cases where you need a summary.
Which LightRAG versions are supported?
Tested with LightRAG 1.5.4 (API 0313). The server doesn't use GET /documents, which LightRAG 1.5.7 removed.
What language are the tool descriptions in?
English by default. Set LIGHTRAG_MCP_LANG=ru for Russian. This only changes what the server says to the agent. Your notes can be in any language LightRAG handles. PRs with more languages are welcome: all texts live in two dictionaries, MESSAGES and TOOL_DOCS.
Development
uv pip install --python .venv/bin/python pytest
.venv/bin/python -m pytest -qIssues and pull requests are welcome.
License
MIT, see LICENSE.
This server cannot be deployed
Maintenance
Related MCP Connectors
Persistent memory for AI agents. Search and store durable facts, preferences and decisions.
Persistent memory and knowledge graphs for AI agents. Hybrid search, context checkpoints, and more.
Universal memory for AI agents and tools. Save, organize and search context anywhere.
Shared long-term memory for AI agents: save and recall context as a searchable knowledge graph.
Related MCP Servers
- AlicenseNot gradedqualityAmaintenanceProvides persistent, local-first memory with knowledge graph and hybrid search for AI coding agents, reducing token usage by storing decisions, patterns, and codebase context.9MIT
- FlicenseNot gradedqualityDmaintenanceEnables AI agents to index, search, and retrieve architectural documentation and store self-learning notes from codebases.1-
- AlicenseNot gradedqualityBmaintenanceEnables AI agents to search, read, and traverse a local knowledge base of Markdown files using full-text search and relationship graph, reducing token usage.MIT
- AlicenseAqualityCmaintenanceEnables coding agents to persist notes, decisions, requirements, and learnings per project in SQLite and retrieve them across sessions via search and status tools.3MIT