Gemini CLI RAG MCP
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Gemini CLI RAG MCPhow do I configure API keys for different models?"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Gemini CLI RAG MCP
A local MCP documentation search server. It extracts versioned official Gemini CLI documentation, generates local embeddings, and returns relevant passages with source URLs. The connected MCP client generates the final answer; this server does not call a generative model.
Pipeline
extract.pydownloads an explicit official tag/commit, or reads a local docs directory. It selects user-facing.mdand.mdxfiles and writes JSONL sections with provenance.create_vectorstore.pyconsumes that JSONL and builds a new model-specific index directory.rag.pyshares pinned model configuration, token-aware chunking, input validation, and index compatibility checks between indexing and serving.vector_store.pystores vectors and metadata in Parquet and performs exact cosine-neighbor search with scikit-learn. LangChain is not required.gemini_cli_mcp.pyexposes documentation search and a full-documentation resource over MCP stdio.
The default embedding candidate is intfloat/multilingual-e5-small (384 dimensions). BAAI/bge-large-en-v1.5 remains available as the bge-large-en comparison profile (1024 dimensions). Both use pinned model revisions and a maximum of 512 input tokens.
Related MCP server: RagDocs MCP Server
Requirements
Python 3.11+; local verification uses the Conda environment
project_envwith Python 3.11.GitHub CLI (
gh) with working authentication for official documentation downloads. Local extraction does not require it.Internet access for the initial model download. Embedding inference runs locally on CPU.
Docker and Compose are optional. The Python server itself does not require Node.js or a Gemini CLI installation.
Install dependencies
The direct dependencies are pinned in requirements.txt. requirements.lock contains the resolved Linux CPU dependencies and distribution hashes.
With an activated Python environment:
python -m pip install torch==2.14.0 --index-url https://download.pytorch.org/whl/cpu
python -m pip install --require-hashes -r requirements.lockAlternatively, with uv and the existing Conda environment:
uv pip install --python "$CONDA_PREFIX/bin/python" --torch-backend cpu --require-hashes -r requirements.lockThe CPU-specific lock was generated with:
uv pip compile requirements.txt --python-version 3.11 --torch-backend cpu --exclude-newer 2026-09-15 --no-annotate --no-header --generate-hashes --output-file requirements.lockExtract documentation
Authenticate gh if needed, then choose an explicit stable tag:
python extract.py --ref v0.60.0 --output gemini_cli_documents-v0.60.0-r2.jsonlThe tag is resolved to a full commit before downloading. Moving branches and preview tags are rejected; an explicit full commit SHA is also accepted. Downloads are temporary and do not depend on the legacy gemini-cli submodule.
The r2 dataset includes MDX installation and authentication guides that were absent from the first Markdown-only extraction. The v0.60.0 corpus contains 84 source documents and 1,310 sections.
For local input:
python extract.py --source-dir gemini-cli/docs --output gemini_cli_documents-local.jsonlLocal files are not assigned an unverified official revision or URL. Extraction and its tests use only the Python standard library.
The initial corpus includes installation, authentication, configuration, commands, tools, MCP, extensions, hooks, skills, sandboxing, and troubleshooting. Contributor workflows, internal development documentation, translations, and changelogs are excluded. Markdown headings define sections while code examples, tables, and links remain in the text. Single-line MDX component imports are omitted; component labels and content are retained as text. MDX is not rendered or executed.
Each JSONL row contains id, page_content, and metadata. Metadata includes the source file, heading hierarchy, line range, document/content hashes, collection timestamp, and official version/commit where verified. Existing output files are never overwritten. Missing or empty input and download errors exit nonzero.
Build a new index
python create_vectorstore.py --input gemini_cli_documents-v0.60.0-r2.jsonl --model e5-smallThe default destination is indexes/e5-small-v0.60.0. To compare the previous embedding model on the same corpus:
python create_vectorstore.py --input gemini_cli_documents-v0.60.0-r2.jsonl --model bge-large-en --output indexes/bge-large-en-v0.60.0An index directory contains:
vectors.parquet: normalized embeddings, chunk text, and source metadata.documents.jsonl: the original extracted sections used to build this index.manifest.json: model/revision, dimensions, query/document prefixes, chunking settings, source revisions, counts, and artifact checksums.
Existing directories are never overwritten. A failed build may leave an incomplete directory without a manifest; use a new destination for a retry. The historical gemini_cli_docs.txt and gemini_cli_vectorstore.parquet are not modified or used by the new server.
Chunking uses the selected model's tokenizer. The 512-token budget includes the section title, model-specific prefix, and special tokens. Long sections are divided with up to 40 tokens of overlap, preferring paragraph boundaries. Query and passage inputs that still exceed the model limit are rejected rather than silently truncated.
E5 uses query: and passage: prefixes. BGE uses its recommended retrieval instruction for queries. Model weights are loaded with remote code disabled and safetensors enabled. Changing the model or its preprocessing requires a new index, even if the embedding dimensions happen to match.
Run the MCP server
The client starts the server as a stdio subprocess. Do not allocate a TTY for the server process. Model and index loading is lazy and reused across queries in that process; diagnostics go to stderr, not protocol stdout.
Settings:
Variable | Default | Purpose |
|
| Must match the index manifest. |
|
| Directory containing the complete index bundle. |
|
| CPU thread count for PyTorch. |
| unset | Set to |
For a client such as Gemini CLI, configure the appropriate mcpServers entry with absolute paths:
{
"mcpServers": {
"gemini_docs": {
"command": "/absolute/path/to/python-environment/bin/python",
"args": ["/absolute/path/to/gemini-cli-rag-mcp/gemini_cli_mcp.py"],
"env": {
"RAG_MODEL": "e5-small",
"RAG_INDEX_DIR": "/absolute/path/to/gemini-cli-rag-mcp/indexes/e5-small-v0.60.0"
}
}
}
}The existing MCP interfaces are retained:
gemini_cli_query_tool(query: str): returns three retrieved chunks, including source file, section, version, commit, URL, and cosine similarity. Similarity is not a confidence probability.docs://gemini-cli/full: returns the extracted sections from the same index bundle, without loading embedding weights.
The manifest and artifact checksums are validated before an index is opened. Missing, corrupted, and incompatible indexes fail explicitly; the server never silently falls back to the historical index. Restart the MCP process after changing its index or model settings.
Retrieval always returns nearest passages for a nonempty valid query. There is no calibrated relevance threshold yet; the client should not assume every returned passage answers the question. The full resource can be large, so search is preferable for focused questions.
Docker
Build the index on the host first, then:
docker compose up -d --buildCompose mounts the E5 index read-only and provides a named Hugging Face model cache. No Anthropic or other LLM API key is needed. It refuses to create a missing host index directory automatically. The image excludes datasets, indexes, Git history, and dotenv files from its build context.
The existing on-demand docker exec workflow is preserved: Compose keeps a helper container alive, and the MCP client launches the server with:
{
"mcpServers": {
"gemini_docs": {
"command": "docker",
"args": ["exec", "-i", "gemini-cli-mcp-container", "python", "gemini_cli_mcp.py"]
}
}
}Verification and model comparison
Offline unit tests, including mocked encoders and index round trips:
python -B -m unittest discover -vUsing the existing environment, prefix Python commands with conda run -n project_env. After building the default index and caching its model, run the real MCP stdio test:
RUN_MCP_INTEGRATION=1 python -B -m unittest -v test_mcp.StdioIntegrationTestsThe evaluation fixture contains 15 topics with paired Portuguese and English questions (30 queries). Evaluate each model separately, without another indexing job competing for the CPU:
HF_HUB_OFFLINE=1 python evaluate_retrieval.py --index indexes/e5-small-v0.60.0 --model e5-small --output indexes/e5-small-v0.60.0/evaluation.json
HF_HUB_OFFLINE=1 python evaluate_retrieval.py --index indexes/bge-large-en-v0.60.0 --model bge-large-en --output indexes/bge-large-en-v0.60.0/evaluation.jsonEvaluation checks coverage of every original section and the token budget of every chunk before querying. Reports include source hit rates at 1/3/5, source MRR at 5, warm query latency, model/index load time, and peak process RSS on Linux. It validates that every expected source exists in the selected corpus. Output reports are never overwritten.
This is a small, manually authored source-level smoke benchmark, not a held-out section-relevance benchmark or an evaluation of generated answers. BGE is evaluated with corrected preprocessing, not the original oversized-chunk pipeline. Model selection should account for these limitations, not just the headline score.
This server cannot be deployed
Maintenance
Related MCP Connectors
Ingest, manage, and retrieve documents for RAG-powered AI applications
Versioned documentation registry and semantic search for AI tools and coding assistants.
Run AI customer support from your terminal: conversations, knowledge base, and chat widget.
Docs Q&A: search 169 data and AI guides, fetch any page as markdown. Read-only, keyless.
Related MCP Servers
- FlicenseNot gradedqualityFmaintenanceProvides curated documentation access via the Gemini API, enabling users to query and interact with technical docs effectively by overcoming context and search limitations.14-
- AlicenseNot gradedqualityNot gradedmaintenanceProvides RAG capabilities for semantic document search using Qdrant vector database and Ollama/OpenAI embeddings, allowing users to add, search, list, and delete documentation with metadata support.16 npm16-
- AlicenseNot gradedqualityNot gradedmaintenanceEnables serving and querying documentation with AI capabilities, allowing users to search, ask questions, and get AI-powered answers from their documentation files. Built for seamless deployment on the Dedalus platform with OpenAI integration for enhanced document analysis.MIT
- FlicenseNot gradedqualityDmaintenanceEnables semantic search and retrieval of information from technical documentation PDFs using RAG-powered natural language queries with Ollama embeddings and LLMs.6-