rag-retriever-mcp
by code-lawyer
README.md
# rag-retriever
A lightweight, **local-first document retrieval engine** that mounts to an agent
as an **MCP tool**. Drop files in; the agent searches them and answers with **its
own LLM**. There is no LLM in here — this is only the "front half" of RAG
(extract → chunk → embed → store + similarity search).
```
your agent (owns the LLM)
│ calls MCP tool: search("question")
▼
rag-retriever ──► extract ─► chunk ─► embed ─► LanceDB
▲ │
└────────── returns relevant passages ◄────────┘
│
agent reads passages → answers with its own LLM
```
Built from the same proven pieces as Open Notebook (file extraction + bge-m3
embeddings + vector search), minus the heavyweight backend, UI, and answer/podcast
generation you don't need.
## Why this shape
- **One LLM, not two.** The retriever never answers; your agent does. You keep
full control of reasoning, prompts, and cost.
- **Local-first.** Default backend (`fastembed`) runs entirely offline, no server.
- **Pluggable embeddings.** Switch between fully local and a China-friendly cloud
API (SiliconFlow) with one env var — no code change.
## Install
```bash
cd rag-retriever
uv sync
cp .env.example .env # then pick your embedding backend
```
## Configure the embedding backend (`.env`)
| `RAG_EMBED_BACKEND` | What it uses | Notes |
|---|---|---|
| `local` (default) | fastembed (ONNX, in-process) | 100% offline, no server, heavier first install |
| `ollama` | local Ollama daemon | `ollama serve` + `ollama pull bge-m3` |
| `openai` | OpenAI-compatible API (e.g. SiliconFlow) | needs `RAG_OPENAI_API_KEY`; text leaves the machine |
> ⚠️ Index-time and query-time must use the **same backend + model**. Changing the
> model means re-indexing everything.
## Use (CLI, for testing)
```bash
uv run rag-retriever index "C:\path\to\docs" # a file or a whole folder
uv run rag-retriever search "什么是表见代理" -k 5
uv run rag-retriever list
uv run rag-retriever stats
```
## Mount as an MCP server (the real entry point)
Run `uv run rag-retriever-mcp` (stdio). Register it with your MCP client. For
Claude Code, add to your MCP config:
```json
{
"mcpServers": {
"rag-retriever": {
"command": "uv",
"args": ["run", "--directory", "D:\\Vibe Coding Items\\rag-retriever", "rag-retriever-mcp"]
}
}
}
```
Tools exposed: `index_path`, `search`, `list_sources`, `stats`.
## Supported files
pdf, docx, pptx, xlsx, html, md, txt, csv, json, epub (via markitdown).
**Scanned / image-only PDFs** need an OCR engine (tesseract) installed separately;
without it they extract empty and are reported as skipped.
## Layout
```
rag_retriever/
config.py # env-driven config; picks the embedding backend
extract.py # file -> text (markitdown)
chunk.py # token-based chunking with overlap
embed.py # local | ollama | openai-compatible backends
store.py # LanceDB vector store (embedded, no server)
pipeline.py # ingest + search orchestration (no LLM)
server.py # MCP server (agent-facing)
cli.py # manual CLI
```
TDQS
A4/5.0
Scored across 4 tools
Disambiguation5/5
Each tool has a distinct purpose: indexing files, searching passages, listing indexed sources, and showing system stats—no overlap or ambiguity.
Naming Consistency5/5
All tool names follow a consistent verb_noun snake_case pattern (e.g., index_path, list_sources) with clear, predictable naming.
Tool Count5/5
With only 4 tools, the set is well-scoped for a document retrieval MCP, covering core operations without unnecessary bloat.
Completeness4/5
The set covers key operations (index, search, list, stats) but lacks a tool to delete indexed documents, which is a minor but notable gap.
Maintenance
ActivityInactive
ResponsivenessNo issues