LangGraph RAG MCP
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@LangGraph RAG MCPhow do I use StateGraph in LangGraph?"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
LangGraph RAG MCP
A local documentation retrieval server for MCP-compatible coding assistants. It indexes LangGraph's official Markdown documentation, searches it with local embeddings, and returns excerpts with source URLs.
The server does not call an LLM or require an Anthropic/OpenAI API key. Your MCP host generates the final answer. LangGraph is the documentation source, not an orchestration dependency.
If you only need hosted, up-to-date documentation without maintaining a local index, LangChain also provides official documentation MCP servers.
Requirements
Python 3.11–3.13 and uv, or Docker with Docker Compose.
Internet access for installation, documentation updates, and the initial Hugging Face model download.
Enough disk space and RAM for CPU inference with
BAAI/bge-large-en-v1.5. Model weights are downloaded separately from the Python packages and can be substantial.
After the model is cached and the index is built, retrieval is local. Set HF_HUB_OFFLINE=1 for an explicitly offline runtime.
Related MCP server: MCP Documentation Server
Quick start
From the repository root:
uv sync --locked --extra embeddings
uv run --locked --extra embeddings langgraph-rag index
uv run --locked --extra embeddings langgraph-rag query "How do interrupts work?" --k 5
uv run --locked --extra embeddings langgraph-rag statusRun the MCP server:
uv run --locked --extra embeddings langgraph-rag serveThe server uses stdio: it waits for an MCP host, not browser requests. Logs go to stderr; stdout is reserved for protocol messages. Starting it and discovering tools do not download the model or require an existing index. Searching requires both.
Keep --extra embeddings on uv run commands that need the local model. Running uv sync without that extra intentionally installs only the lightweight development/test environment.
Data and configuration
By default, data is stored at $XDG_DATA_HOME/langgraph-rag-mcp, falling back to ~/.local/share/langgraph-rag-mcp. This is independent of the working directory.
For project-local data, use the same directory for indexing and serving:
uv run --locked --extra embeddings langgraph-rag --data-dir ./data index
uv run --locked --extra embeddings langgraph-rag --data-dir ./data serveGlobal CLI options must precede the subcommand. Alternatively, export LANGGRAPH_MCP_DATA_DIR as an absolute path.
Environment variable | Default | Purpose |
| XDG data directory | Index snapshots and embedding cache |
|
| Documentation index |
|
| Local embedding model |
|
| Pinned Hugging Face model revision |
|
| Maximum embedding-tokenizer tokens, including special tokens |
|
| Chunk overlap budget |
|
| Maximum number of discovered Markdown pages |
|
| HTTP timeout in seconds |
| Hugging Face default | Downloaded model/tokenizer cache |
Configuration comes from exported environment variables and CLI options; .env files are not loaded implicitly. --index-url overrides the source URL. For another model, set both its name and an appropriate revision, then rebuild the index. The pinned default revision belongs specifically to BGE.
MCP integration
For hosts using the mcpServers configuration format, merge an entry like this into the host's settings:
{
"mcpServers": {
"langgraph-docs": {
"command": "uv",
"args": [
"run",
"--project", "/absolute/path/to/langgraph-rag-mcp",
"--locked",
"--extra", "embeddings",
"langgraph-rag", "serve"
]
}
}
}Use the absolute path to uv if it is not on your editor's PATH. For VS Code, use the servers key instead of mcpServers in its MCP configuration. If you override the data directory or model during indexing, pass the same settings through the host's env configuration.
Tools and resources
langgraph_query_tool(query, k=3): returnsquery,indexed_at, and aresultslist. Each result includes an ID, title, section, source URL, content, and cosine similarity.kis bounded to 1–20 and queries to 4,000 characters. Similarity is not a confidence score.read_page(source, offset=0, limit=12000): reads only a source already present in the local index. Pagination is in characters, with a maximum of 20,000 per request. It cannot fetch arbitrary URLs.docs://langgraph/status: index timestamp, counts, source and embedding configuration.docs://langgraph/full: compatibility resource for small corpora, limited to 200,000 characters. Prefer search and paginated reads for the complete LangGraph corpus.
Tools provide structured output as well as text content for host compatibility. Both tools are advertised as read-only. The embedding model and vector index are reused across requests; newly published snapshots are picked up automatically.
Docker
Build the image and create the index before connecting your host:
docker compose build
docker compose run --rm --no-deps -T langgraph-mcp index
docker compose run --rm --no-deps -T langgraph-mcp query "How do checkpoints work?"Then use the wrapper as the MCP host command:
{
"mcpServers": {
"langgraph-docs": {
"command": "/bin/bash",
"args": ["/absolute/path/to/langgraph-rag-mcp/run-mcp-docker.sh"]
}
}
}The wrapper resolves the project directory itself, does not allocate a TTY, and starts the server directly in a disposable container. An optional --build argument rebuilds the image first, sending build output to stderr. There is no background tail process, sleep, or docker exec lifecycle.
The image uses a locked CPU-only PyTorch installation on Linux, runs as UID 10001, and excludes .env, Git metadata, notebooks, tests and local data from the build context. Named volumes preserve the index (index-data) and downloaded models (model-cache). These Docker volumes are separate from a native installation's data directory.
Indexing behavior
Discover Markdown pages from the configured
llms.txt, including nested indexes within its directory scope.Fetch pages with timeouts, HTTP status checks and a 2 MiB per-response limit. Redirects remain on the configured HTTPS origin; relocated Markdown pages retain Markdown retrieval. HTML/error pages are rejected.
Remove documentation-index/footer boilerplate and deduplicate identical page content.
Split by Markdown headings, then enforce the embedding tokenizer's token budget. Source, title and section metadata travel with each chunk.
Reuse embeddings from a local SQLite cache keyed by content and embedding configuration. Search vectors remain in Parquet, served by scikit-learn cosine nearest-neighbor search.
Publish a new immutable snapshot and atomically switch
current.jsononly after the build succeeds. A failed build leaves the previous index active. Concurrent writers are prevented with a file lock.
Running index again refetches documentation to detect changes, but embeds only new/changed chunks. It does not implement conditional HTTP requests yet. If neither content nor configuration changed, the snapshot is reused. Pages removed from the source index disappear from the new search snapshot.
Old snapshots and cached embeddings are retained; automatic pruning is intentionally not implemented. A manifest records the model revision, chunking settings, source and index timestamp. Serving rejects configuration mismatches rather than silently mixing incompatible embeddings.
Retrieval benchmark
The evaluate command measures retrieval with 20 manually authored questions: ten topics with equivalent English and Portuguese queries. The versioned dataset is bundled with the package at src/langgraph_rag_mcp/data/benchmark.json; no evaluation service or LLM API is involved.
With an existing Docker index, run:
docker compose run --rm --no-deps -T \
-e HF_HUB_OFFLINE=1 -e OMP_NUM_THREADS=4 -e MKL_NUM_THREADS=4 \
-e TOKENIZERS_PARALLELISM=false \
langgraph-mcp evaluate --output /data/baseline.jsonFor a native installation:
uv run --locked --extra embeddings langgraph-rag evaluateWithout --output, stdout contains the complete JSON report. With it, the full report is saved and stdout contains a summary. The output directory must exist, and existing report files are never overwritten: choose a new filename for each comparison. Docker reports remain in the index-data volume.
Use --dataset /path/to/questions.json for another labeled dataset and --k 10 to change the maximum retrieval cutoff (1–20). The dataset has name, version, description and a nonempty cases list. Each case requires a unique id, language (en or pt), query, and a nonempty list of expected_sources. All expected sources must exist in the current index; missing labels fail evaluation rather than silently counting as misses.
Reported metrics, both overall and per language:
Hit@k: fraction of questions with at least one expected source among the first k chunks.
MRR@k: mean reciprocal rank of the first expected source; zero for a miss.
Duplicate-source fraction@k: mean
1 - unique sources / returned chunks. Multiple chunks from one page may be useful, so this measures source diversity, not duplicate text or an error rate.Warm-query latency p50/p95: wall-clock search latency after one excluded warm-up query. The warm-up duration, including lazy model/index loading, is reported separately. p95 uses nearest-rank estimation.
The evaluator preserves the actual chunk ranking: it does not deduplicate results before scoring. Reports include all retrieved excerpts, cases missed at the maximum cutoff, the full index manifest, dataset and implementation SHA-256 fingerprints, package versions, and relevant thread/offline settings. Evaluation aborts if the index changes during the run.
This is a small diagnostic dataset, not a general accuracy claim or an exhaustive relevance judgment. Finding the expected page does not prove that the returned passage answers the question. The paired translations are correlated, and the suite does not assess generated answers, unsupported questions or abstention. Review failures before changing retrieval, and do not edit labels to improve a measured score. Change the dataset version when changing its questions or expected sources.
Development and verification
The default development installation does not include PyTorch, model weights or an LLM SDK:
uv sync --locked
uv run --locked pytest
uv run --locked ruff check .
uv run --locked ruff format --check .
uv run --locked mypy
uv build
bash -n run-mcp-docker.sh
docker compose config --quietTests use deterministic embeddings and mock HTTP responses, but exercise real Parquet persistence, scikit-learn retrieval, MCP structured responses, legacy client compatibility, and subprocess stdio transport. They need no credentials, network access or GPU.
CI runs these checks on Python 3.11, 3.12 and 3.13, and separately builds/smoke-tests the container. GitHub Actions are pinned to commit SHAs. Dependency resolution has a recorded exclude-newer cutoff in pyproject.toml; upgrades should be deliberate, reviewed, locked and tested rather than selecting unbounded latest versions.
Code is organized under src/langgraph_rag_mcp: configuration and schemas, source discovery, chunking, embedding cache, snapshot publication, vector storage, retrieval, MCP server and CLI.
Migrating from the notebook version
rag-tool.ipynbis preserved as a historical experiment, not the supported ingestion path. Its old dependencies, imports and Claude example were not migrated; do not run it to build the new index.langgraph-mcp.pyremains a compatibility entry point after installing the package. Preferlanggraph-rag servefor new host configurations.Existing root-level
sklearn_vectorstore.parquetandllms_full.txtfiles are not overwritten or imported. Build a new index withlanggraph-rag index; the previous chunking and metadata do not satisfy the new index format.requirements.txtis a compatibility adapter forpip install -r requirements.txt; useuv sync --locked --extra embeddingsfor reproducible installations.No Anthropic API key is required. The obsolete
tail/docker execCompose setup has been replaced.
The default embedding model is English-oriented. Use the retrieval benchmark above to assess Portuguese-to-English searches on the current corpus; deterministic unit tests alone cannot establish semantic quality. Hybrid search, reranking and comparisons with lighter or multilingual embedding models remain separate future experiments.
Resources
Available Tools
2 toolslanggraph_query_toolCRead-onlyIdempotent
Search local LangGraph docs; return excerpts, source URLs and cosine similarity.
| Name | Required | Description | Default |
|---|---|---|---|
| k | No | ||
| query | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| query | Yes | |
| results | Yes | |
| indexed_at | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and openWorldHint=false, so safety and determinism are covered by structured data. The description adds the useful fact that results include similarity scores and source URLs, but says nothing about corpus scope, freshness, or result limits beyond the k parameter.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with zero filler; the action and the return payload lead. It is terse to the point of under-specification, but on the conciseness dimension itself it performs well.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and annotations cover the safety profile. However, for a retrieval tool with a sibling and fully undocumented parameters, the description omits usage routing and parameter meaning, leaving real gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for two parameters, so the description must compensate and does not: it never explains what 'k' controls (result count) or what 'query' should contain. An agent can guess k from the maximum of 20, but the description provides no added meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb (Search) and resource (local LangGraph docs) and even previews the return shape (excerpts, source URLs, cosine similarity). It is clear on its own, but it does not distinguish itself from the sibling read_page tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to use this search tool versus the sibling read_page, and no prerequisites or exclusions. The agent must infer the retrieval-vs-fetch distinction from the names alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_pageARead-onlyIdempotent
Read an indexed source URL from search results, with character-based pagination.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| offset | No | ||
| source | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| title | Yes | |
| offset | Yes | |
| source | Yes | |
| content | Yes | |
| next_offset | Yes | |
| total_characters | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint=false, so the safety profile is covered. The description adds a genuine behavioral trait beyond that: pagination is character-based, telling the agent how offset/limit are measured. It does not state error behavior for unindexed URLs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence with no filler, front-loading what is read and the pagination model. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and annotations cover safety. However, for a 3-parameter tool with zero schema coverage, the description leaves the required "source" format and the limit/offset defaults undocumented, so it is only minimally sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry the load. "Character-based pagination" usefully clarifies the unit of limit/offset, but defaults (limit 12000, offset 0), the 20000 maximum, and the expected format/meaning of the required "source" value are all left unstated.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Read) and resource (an indexed source URL from search results), which is unambiguous on its own. It does not name or contrast with the sibling langgraph_query_tool, so it stops short of full sibling differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"from search results" implies this tool is used downstream of a search step on already-indexed URLs, giving implied usage context. There is no explicit when-to-use, when-not-to-use, or named alternative, so the guidance remains inferential.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
v0.2.0- First observed
langgraph_query_tool - First observed
read_page
TDQS
Scored across 2 tools
langgraph_query_tool and read_page have clearly distinct purposes: one searches indexed docs and returns excerpts/URLs, the other reads a specific URL with pagination. No overlap; the intended workflow is search then read.
Both use snake_case, but langgraph_query_tool includes a domain prefix and redundant _tool suffix while read_page is a simple verb_noun. The inconsistency in naming pattern makes it mixed but still readable.
Two tools is thin for a RAG server; a typical retrieval interface might include a list-sources or metadata tool. The pair covers core search-and-read but feels minimal.
The search and read tools cover the essential retrieval lifecycle (query -> fetch full content). Minor gaps exist, such as listing all indexed documents or retrieving metadata, but agents can work around them.
Maintenance
Related MCP Connectors
Versioned documentation registry and semantic search for AI tools and coding assistants.
Ingest, manage, and retrieve documents for RAG-powered AI applications
Medical RAG: semantic search for clinical guidelines, drug interactions, diagnoses & EHR data.
Medical RAG: semantic search for clinical guidelines, drug interactions, diagnoses & EHR data.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceA simple Model Context Protocol server that enables searching and retrieving relevant documentation snippets from Langchain, Llama Index, and OpenAI official documentation.-
- FlicenseNot gradedqualityDmaintenanceA customized MCP server that enables integration between LLM applications and documentation sources, providing AI-assisted access to LangGraph and Model Context Protocol documentation.-
- AlicenseNot gradedqualityDmaintenanceA modular RAG (Retrieval-Augmented Generation) service framework with pluggable architecture and full observability, enabling AI assistants to perform document Q\&A, semantic search, and knowledge base construction through the Model Context Protocol.MIT
- FlicenseNot gradedqualityCmaintenanceProvides MCP-compatible hosts with retrieval-augmented access to LangGraph documentation, enabling Claude and other assistants to answer questions with relevant, source-attributed context.-