Skip to main content
Glama
p-belov

lightrag-memory-mcp

by p-belov

lightrag-memory-mcp

English | Русский

License: MIT Python 3.10+ MCP: stdio Tests: 20 passing

Long-term memory for your AI coding agents that answers in a second, not a minute.

lightrag-memory-mcp is a small MCP server that connects Claude Code, Codex, or any other MCP client to a LightRAG knowledge base. It gives your agent the facts and skips the slow part.

Wrapper that calls /query

lightrag-memory-mcp

Memory lookup

17 s median, 49 s p90, up to 164 s

0.2–2.5 s

Lookups lost to client timeouts

~1 in 4 (30 s limit)

none in our tests

Reading memory when your LLM is down

fails

still works (falls back to a mode that uses no LLM)

What the agent gets

a summary written by another LLM

original notes with sources, plus the graph

Measured on a real-world knowledge base: ~37k graph nodes, ~1.5k documents, LightRAG 1.5.4.


The problem

You give your agent a LightRAG memory so it remembers decisions, context, and the reasons behind them. Then this happens:

  • Every lookup is slow. Most wrappers call LightRAG's /query, so LightRAG runs an LLM to write a polished answer. That takes 15–60 seconds, sometimes minutes.

  • Lookups time out. Wrappers often hardcode a 30-second limit. In our logs, about one query in four went past it, so the agent got an error instead of memory.

  • You pay for two LLMs doing one job. LightRAG's LLM summarizes the facts, and then your agent's LLM reads that summary and writes its own answer. The first step adds latency and cost, and it loses details.

  • Your memory depends on your LLM provider. If the API is rate-limited, down, or blocked in your region, the agent loses its memory as well.

Related MCP server: codebase-wiki

The fix

Your agent is already a capable LLM. It doesn't need a summary. It needs the facts.

  • ⚡ Fast search with no answer generation. memory_search calls /query/data and returns entities, relations, and original text chunks directly.

  • 🛟 Keeps working without an LLM. If the primary search mode fails (LLM timeout, 5xx, status=failure), the server retries in naive mode, which never calls an LLM. The response tells the agent that a fallback happened.

  • 🎯 Context-friendly output. Raw /query/data responses run to hundreds of kilobytes. The server fits each result into a budget (16,000 characters by default). Original notes with their sources come first, because they hold the dates and the reasons. The graph follows, one line per entity or relation.

  • ✍️ Safe writes. LightRAG rejects a repeated source name with HTTP 409, and agents reuse names all the time. memory_save adds a timestamp to every source name, so no note is silently lost.

  • 🌍 English or Russian. Tool descriptions and responses are in English by default. Set LIGHTRAG_MCP_LANG=ru for Russian.

  • 🪶 Small and auditable. One ~400-line Python file, two pinned dependencies, 20 tests.

How it works

flowchart LR
    A["AI agent<br/>(Claude Code, Codex, …)"] -- memory_search --> M["lightrag-memory-mcp"]
    M -- "POST /query/data<br/>mode = mix" --> L[("LightRAG")]
    L -. "LLM timeout / 5xx / failure" .-> M
    M -- "retry: mode = naive<br/>(no LLM call)" --> L
    M -- "compact result:<br/>notes + sources first, then graph" --> A

Here is what your agent gets back from memory_search (shortened):

[memory: mode mix, 0.4 s]
## Note fragments
[decision-2026-09-07-llm-provider--20260907-181502-a1b2.md]
Date: 2026-09-07. Decision: model A stays the primary model.
WHY: model B gives different results from run to run on the same document…
## Entities
- LLM Provider Layer (concept): the component that picks the LLM provider…
## Relations
- DocService → LLM Provider Layer: uses, depends on

If the LLM is down, the header shows it and the search still returns results:

[memory: mode naive, 0.1 s; fell back to naive: timeout 20 s on /query/data]

Quick start

You need a running LightRAG server (default http://localhost:9621) and uv.

git clone https://github.com/p-belov/lightrag-memory-mcp.git
cd lightrag-memory-mcp
uv venv .venv && uv pip install --python .venv/bin/python -r requirements.txt

Claude Code

claude mcp add-json lightrag '{"type":"stdio","command":"'"$PWD"'/.venv/bin/python","args":["'"$PWD"'/lightrag_memory_mcp.py"],"env":{"LIGHTRAG_URL":"http://localhost:9621"}}' -s user

Codex

Add to ~/.codex/config.toml (use absolute paths):

[mcp_servers.lightrag]
command = "/path/to/lightrag-memory-mcp/.venv/bin/python"
args = ["/path/to/lightrag-memory-mcp/lightrag_memory_mcp.py"]
tool_timeout_sec = 200  # memory_answer may wait up to 180 s for the LLM

[mcp_servers.lightrag.env]
LIGHTRAG_URL = "http://localhost:9621"

Tell your agent to use it

Add this to CLAUDE.md or AGENTS.md:

Before answering questions about projects → call lightrag:memory_search.
After a decision or an important fact → call lightrag:memory_save
(include the date, what was decided, and why).

Tools

Tool

What it does

LightRAG endpoint

memory_search

Finds facts without generating an answer. Falls back to naive (no LLM) on failure.

POST /query/data

memory_save

Saves a decision or fact and returns a track_id

POST /documents/text

memory_answer

Asks LightRAG's LLM for a written answer. Slow; use it only when you need a summary of a large area.

POST /query

memory_status

Server health, pipeline state, and the status of a save by track_id

GET /health, GET /documents/track_status/{id}

Configuration

All settings are environment variables.

Variable

Default

Purpose

LIGHTRAG_URL

http://localhost:9621

LightRAG server address

LIGHTRAG_API_KEY

—

API key, if your server has auth enabled (sent as X-API-Key)

LIGHTRAG_SEARCH_MODE

mix

Default search mode: mix, hybrid, local, global, naive

LIGHTRAG_SEARCH_TIMEOUT

20

Seconds to wait for the primary mode before falling back to naive

LIGHTRAG_ANSWER_TIMEOUT

180

Timeout for memory_answer, seconds

LIGHTRAG_MAX_RESULT_CHARS

16000

Character budget for one memory_search result

LIGHTRAG_MCP_INSTRUCTIONS

generic text

Describes your memory to the agent, for example which projects it covers

LIGHTRAG_MCP_LANG

en

Language of tool descriptions and responses: en or ru

FAQ

Does this replace LightRAG? No. It is a thin client. LightRAG still stores the data and builds the knowledge graph.

Do I still need an LLM? LightRAG needs one to write memory, because it extracts the graph with an LLM. Reading doesn't need one. mix and hybrid use the LLM only to extract keywords from the query, and if that fails, the server switches to naive, which makes no LLM calls.

Can I still get a written answer? Yes, call memory_answer. It is slow by design, so keep it for cases where you need a summary.

Which LightRAG versions are supported? Tested with LightRAG 1.5.4 (API 0313). The server doesn't use GET /documents, which LightRAG 1.5.7 removed.

What language are the tool descriptions in? English by default. Set LIGHTRAG_MCP_LANG=ru for Russian. This only changes what the server says to the agent. Your notes can be in any language LightRAG handles. PRs with more languages are welcome: all texts live in two dictionaries, MESSAGES and TOOL_DOCS.

Development

uv pip install --python .venv/bin/python pytest
.venv/bin/python -m pytest -q

Issues and pull requests are welcome.

License

MIT, see LICENSE.

Related MCP Connectors

Related MCP Servers