lightrag-memory-mcp
by p-belov
README.md
# lightrag-memory-mcp
**English** | [Русский](README.ru.md)
[](LICENSE)



**Long-term memory for your AI coding agents that answers in a second, not a minute.**
`lightrag-memory-mcp` is a small MCP server that connects Claude Code, Codex, or any other MCP client to a [LightRAG](https://github.com/HKUDS/LightRAG) knowledge base. It gives your agent the facts and skips the slow part.
| | Wrapper that calls `/query` | lightrag-memory-mcp |
|---|---|---|
| Memory lookup | 17 s median, 49 s p90, up to 164 s | **0.2–2.5 s** |
| Lookups lost to client timeouts | ~1 in 4 (30 s limit) | **none in our tests** |
| Reading memory when your LLM is down | fails | **still works** (falls back to a mode that uses no LLM) |
| What the agent gets | a summary written by another LLM | **original notes with sources**, plus the graph |
*Measured on a real-world knowledge base: ~37k graph nodes, ~1.5k documents, LightRAG 1.5.4.*
---
## The problem
You give your agent a LightRAG memory so it remembers decisions, context, and the reasons behind them. Then this happens:
- **Every lookup is slow.** Most wrappers call LightRAG's `/query`, so LightRAG runs an LLM to write a polished answer. That takes 15–60 seconds, sometimes minutes.
- **Lookups time out.** Wrappers often hardcode a 30-second limit. In our logs, about one query in four went past it, so the agent got an error instead of memory.
- **You pay for two LLMs doing one job.** LightRAG's LLM summarizes the facts, and then your agent's LLM reads that summary and writes its own answer. The first step adds latency and cost, and it loses details.
- **Your memory depends on your LLM provider.** If the API is rate-limited, down, or blocked in your region, the agent loses its memory as well.
## The fix
Your agent is already a capable LLM. It doesn't need a summary. It needs the facts.
- ⚡ **Fast search with no answer generation.** `memory_search` calls `/query/data` and returns entities, relations, and original text chunks directly.
- 🛟 **Keeps working without an LLM.** If the primary search mode fails (LLM timeout, 5xx, `status=failure`), the server retries in `naive` mode, which never calls an LLM. The response tells the agent that a fallback happened.
- 🎯 **Context-friendly output.** Raw `/query/data` responses run to hundreds of kilobytes. The server fits each result into a budget (16,000 characters by default). Original notes with their sources come first, because they hold the dates and the reasons. The graph follows, one line per entity or relation.
- ✍️ **Safe writes.** LightRAG rejects a repeated source name with HTTP 409, and agents reuse names all the time. `memory_save` adds a timestamp to every source name, so no note is silently lost.
- 🌍 **English or Russian.** Tool descriptions and responses are in English by default. Set `LIGHTRAG_MCP_LANG=ru` for Russian.
- 🪶 **Small and auditable.** One ~400-line Python file, two pinned dependencies, 20 tests.
## How it works
```mermaid
flowchart LR
A["AI agent<br/>(Claude Code, Codex, …)"] -- memory_search --> M["lightrag-memory-mcp"]
M -- "POST /query/data<br/>mode = mix" --> L[("LightRAG")]
L -. "LLM timeout / 5xx / failure" .-> M
M -- "retry: mode = naive<br/>(no LLM call)" --> L
M -- "compact result:<br/>notes + sources first, then graph" --> A
```
Here is what your agent gets back from `memory_search` (shortened):
```text
[memory: mode mix, 0.4 s]
## Note fragments
[decision-2026-09-07-llm-provider--20260907-181502-a1b2.md]
Date: 2026-09-07. Decision: model A stays the primary model.
WHY: model B gives different results from run to run on the same document…
## Entities
- LLM Provider Layer (concept): the component that picks the LLM provider…
## Relations
- DocService → LLM Provider Layer: uses, depends on
```
If the LLM is down, the header shows it and the search still returns results:
```text
[memory: mode naive, 0.1 s; fell back to naive: timeout 20 s on /query/data]
```
## Quick start
You need a running [LightRAG server](https://github.com/HKUDS/LightRAG) (default `http://localhost:9621`) and [uv](https://docs.astral.sh/uv/).
```bash
git clone https://github.com/p-belov/lightrag-memory-mcp.git
cd lightrag-memory-mcp
uv venv .venv && uv pip install --python .venv/bin/python -r requirements.txt
```
### Claude Code
```bash
claude mcp add-json lightrag '{"type":"stdio","command":"'"$PWD"'/.venv/bin/python","args":["'"$PWD"'/lightrag_memory_mcp.py"],"env":{"LIGHTRAG_URL":"http://localhost:9621"}}' -s user
```
### Codex
Add to `~/.codex/config.toml` (use absolute paths):
```toml
[mcp_servers.lightrag]
command = "/path/to/lightrag-memory-mcp/.venv/bin/python"
args = ["/path/to/lightrag-memory-mcp/lightrag_memory_mcp.py"]
tool_timeout_sec = 200 # memory_answer may wait up to 180 s for the LLM
[mcp_servers.lightrag.env]
LIGHTRAG_URL = "http://localhost:9621"
```
### Tell your agent to use it
Add this to `CLAUDE.md` or `AGENTS.md`:
```text
Before answering questions about projects → call lightrag:memory_search.
After a decision or an important fact → call lightrag:memory_save
(include the date, what was decided, and why).
```
## Tools
| Tool | What it does | LightRAG endpoint |
|---|---|---|
| `memory_search` | Finds facts without generating an answer. Falls back to `naive` (no LLM) on failure. | `POST /query/data` |
| `memory_save` | Saves a decision or fact and returns a `track_id` | `POST /documents/text` |
| `memory_answer` | Asks LightRAG's LLM for a written answer. Slow; use it only when you need a summary of a large area. | `POST /query` |
| `memory_status` | Server health, pipeline state, and the status of a save by `track_id` | `GET /health`, `GET /documents/track_status/{id}` |
## Configuration
All settings are environment variables.
| Variable | Default | Purpose |
|---|---|---|
| `LIGHTRAG_URL` | `http://localhost:9621` | LightRAG server address |
| `LIGHTRAG_API_KEY` | — | API key, if your server has auth enabled (sent as `X-API-Key`) |
| `LIGHTRAG_SEARCH_MODE` | `mix` | Default search mode: `mix`, `hybrid`, `local`, `global`, `naive` |
| `LIGHTRAG_SEARCH_TIMEOUT` | `20` | Seconds to wait for the primary mode before falling back to `naive` |
| `LIGHTRAG_ANSWER_TIMEOUT` | `180` | Timeout for `memory_answer`, seconds |
| `LIGHTRAG_MAX_RESULT_CHARS` | `16000` | Character budget for one `memory_search` result |
| `LIGHTRAG_MCP_INSTRUCTIONS` | generic text | Describes your memory to the agent, for example which projects it covers |
| `LIGHTRAG_MCP_LANG` | `en` | Language of tool descriptions and responses: `en` or `ru` |
## FAQ
**Does this replace LightRAG?**
No. It is a thin client. LightRAG still stores the data and builds the knowledge graph.
**Do I still need an LLM?**
LightRAG needs one to *write* memory, because it extracts the graph with an LLM. *Reading* doesn't need one. `mix` and `hybrid` use the LLM only to extract keywords from the query, and if that fails, the server switches to `naive`, which makes no LLM calls.
**Can I still get a written answer?**
Yes, call `memory_answer`. It is slow by design, so keep it for cases where you need a summary.
**Which LightRAG versions are supported?**
Tested with LightRAG 1.5.4 (API 0313). The server doesn't use `GET /documents`, which LightRAG 1.5.7 removed.
**What language are the tool descriptions in?**
English by default. Set `LIGHTRAG_MCP_LANG=ru` for Russian. This only changes what the server says to the agent. Your notes can be in any language LightRAG handles. PRs with more languages are welcome: all texts live in two dictionaries, `MESSAGES` and `TOOL_DOCS`.
## Development
```bash
uv pip install --python .venv/bin/python pytest
.venv/bin/python -m pytest -q
```
Issues and pull requests are welcome.
## License
MIT, see [LICENSE](LICENSE).
This server cannot be deployed
Maintenance
ActivityMaintained
ResponsivenessNo issues