rag-starter
Offers an optional OpenAI embedding backend for indexing and retrieving documents, allowing the server to use OpenAI models instead of local ONNX embeddings.
rag-starter — chat with your documents (RAG), with citations
A production-ready starter that turns a folder of documents into a cited Q&A service. Drop in your PDFs / Markdown / text, ask questions, get answers grounded in the source — every claim traceable to the exact passage it came from.
Exposed two ways from one codebase:
MCP server — plug it into Claude Desktop / Claude Code / any MCP host and chat with your docs.
HTTP API (FastAPI) — call it from any app.
Built once, reskinned per client. Swap the
data/folder, tweakconfig.py, ship.
Why it's different
Keyless by default. Embeddings run locally (ONNX MiniLM) — no API key, no per-query cost, runs offline. Demo it anywhere in seconds.
Citations, not hallucinations. Retrieval returns ranked passages tagged
[source#chunk]/[file.pdf p3]. A missing answer returns "Not found in the documents" — never a fabricated one.Answer synthesis is optional. With an
ANTHROPIC_API_KEYit writes a cited answer for you; without one it returns passages for the host LLM to answer. Either way the RAG works.Idempotent ingestion. Re-ingesting a file updates it in place (no duplicates).
Related MCP server: PinRAG
Quickstart
pip install -r requirements.txt # or: pip install -e .
PYTHONUTF8=1 python tests/test_smoke.py # proves retrieval works on the sample docsAs an HTTP API
rag-starter-api # uvicorn on 127.0.0.1:8000
# then:
curl -X POST localhost:8000/ingest -H "content-type: application/json" -d '{"path":"./data"}'
curl -X POST localhost:8000/search -H "content-type: application/json" -d '{"query":"refund policy"}'
curl -X POST localhost:8000/answer -H "content-type: application/json" -d '{"query":"how long is the refund window?"}'As an MCP server (Claude Desktop)
Add to claude_desktop_config.json → mcpServers:
{
"rag-starter": {
"command": "python",
"args": ["-m", "rag_starter"],
"cwd": "C:/path/to/rag-starter"
}
}Restart Claude Desktop, then: "Ingest the folder ./data", then "What's the refund policy?"
Tools / endpoints
MCP tool | HTTP | Purpose |
|
| Index a file or folder (txt/md/pdf). |
|
| Top-k cited passages (keyless). |
|
| Cited answer (with key) or passages (without). |
— |
| Liveness + indexed-chunk count + embedding backend. |
Configuration
Everything is in config.py (or env / .env — see .env.example): chunk size & overlap,
top-k, embedding backend (default local vs openai), and the answer model. A client
reskin is usually just: replace data/, set RAG_COLLECTION, re-ingest.
Architecture
documents ──ingest.py──> chunks(+citations) ──store.py──> Chroma (local, persistent)
│
question ─────────────────retrieve.py──> top-k cited passages┤
├─> rag_search (host LLM answers)
answer.py ──┴─> rag_answer (server synthesizes, optional)Reuses the mcp-factory contract (Result, formatting, cache) so it composes with the
rest of the catalog.
License
MIT.
Available Tools
3 toolsrag_answerARead-onlyIdempotent
Answer a question grounded in the documents, with citations.
With ANTHROPIC_API_KEY set, returns a synthesized cited answer; otherwise returns the ranked passages for you to answer from.
Examples: - "Summarize the cancellation terms" -> query='cancellation terms' - "Does the plan include support?" -> query='support included in plan'
| Name | Required | Description | Default |
|---|---|---|---|
| params | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so safety is covered. The description adds genuinely new behavioral context: output varies by environment (synthesized cited answer vs. raw ranked passages), which the structured fields do not convey. It says nothing about the retrieval scope or latency, but the mode-switching disclosure is the important part.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core purpose in the first clause, then adds the mode caveat and two short examples. Every element is functionally useful, though the second sentence could be tightened slightly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return format need not be explained, yet the description still usefully distinguishes the two response shapes. Combined with the parameter schema and annotations, an agent has enough to invoke this correctly; only the sibling relationship to rag_search is left unresolved.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The examples demonstrate how to phrase the `query` argument from a user question, which is real added meaning beyond the schema's field description. However `k` and `response_format` are never mentioned in the description, so those rely entirely on the nested schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb+resource+output: answers a question grounded in the provided documents and includes citations. It also clarifies the two possible outputs (synthesized answer vs. ranked passages). It does not name or differentiate from siblings like rag_search, which it partially overlaps with when synthesis is unavailable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explains the condition that selects each mode (ANTHROPIC_API_KEY present vs. absent) and provides worked examples showing how a natural-language question maps to the `query` value. It gives no explicit guidance on when to prefer this over rag_search, leaving the sibling boundary to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
rag_ingestAIdempotent
Index a document or folder so it can be searched and cited.
Examples: - "Index the folder ./data" -> path='./data' - "Add the handbook pdf" -> path='docs/handbook.pdf'
| Name | Required | Description | Default |
|---|---|---|---|
| params | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the full safety profile (readOnlyHint=false, idempotentHint=true, destructiveHint=false, openWorldHint=false), so the bar is lower. The description adds the consequence of the write ('so it can be searched and cited'), which is useful, but says nothing about re-indexing behavior, supported formats (left to the schema), or cost/latency of ingestion.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One declarative sentence front-loads the action and its payoff, followed by two short, well-formatted examples. No filler, no restating of the tool name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described, and annotations cover the mutation's safety profile. The remaining gap is operational context (what happens on re-ingest, whether indexing is synchronous, format constraints), which is minor but would help an agent set expectations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Reported schema description coverage is 0%, so the description must compensate, and it does: the two examples map natural-language requests onto concrete path values ('Index the folder ./data' -> path='./data'), clarifying that path accepts both files and folders. The response_format parameter is not addressed, but its schema description is self-explanatory.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Index') and resource ('a document or folder') plus the downstream purpose ('so it can be searched and cited'). It is distinguishable from rag_search/rag_answer by implication (this is the tool that adds content to the index), but it never names the siblings or explicitly contrasts itself with them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied rather than stated: the examples show how to invoke it, but nothing says when to use rag_ingest rather than rag_search or rag_answer, or that ingestion is a prerequisite for those tools. No prerequisites, limits, or exclusions are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
rag_searchBRead-onlyIdempotent
Return the most relevant passages for a question, each with a citation.
Use this to ground your own answer: read the passages and cite their [tags].
Examples: - "What is the refund policy?" -> query='refund policy' - "How do I reset my password?" -> query='reset password'
| Name | Required | Description | Default |
|---|---|---|---|
| params | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/ idempotent/ non-destructive/ closed-world, so the safety profile is covered. The description usefully adds that results carry citations and should be cited by [tags], but it discloses nothing about ranking quality, empty-result behavior, or latency, so it only modestly exceeds the structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the purpose in sentence one, followed by usage guidance and two compact examples; every element is short and earns its place. Sentence two slightly overlaps the citation point already made in sentence one, a minor redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-shape detail is not required, and the single nested parameter is well covered by the schema. The notable gap is sibling routing: with rag_answer and rag_ingest present, the description should state plainly when to prefer this tool, and it only implies it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description contributes query-narrowing examples ('refund policy', 'reset password') that clarify how to phrase the query parameter. It says nothing about k (default/max) or response_format, so coverage of the remaining parameters rests entirely on the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence gives a specific verb and resource — returns relevant passages with citations — so an agent knows exactly what comes back. It does not explicitly name or distinguish itself from the sibling rag_answer, so the reader must infer the split from 'ground your own answer.'
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Use this to ground your own answer: read the passages and cite their [tags]' implies the when-to-use case (agent composes the answer itself rather than calling rag_answer), and the query examples show input formulation. However, no alternative is named and there is no explicit 'use rag_answer instead when you want a synthesized answer' routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
v0.1.0- First observed
rag_answer - First observed
rag_ingest - First observed
rag_search
TDQS
Scored across 3 tools
The three tools map to distinct lifecycle stages (index, retrieve, synthesize), but rag_search and rag_answer overlap: both take a query and can return ranked passages, since rag_answer falls back to passages without an API key. The descriptions do clarify the difference, so confusion is limited.
All tools use a consistent rag_ prefix followed by a clear verb (ingest, search, answer) in uniform snake_case. The pattern is predictable and immediately readable.
Three tools is a tight, well-scoped set for a RAG starter, each earning its place. It is at the lower end but appropriately minimal for the stated purpose.
The core ingest-search-answer flow is covered, but there is no way to remove, delete, or list indexed documents, and no re-index/update operation. These gaps around document lifecycle management could force agents to work around stale content.
Maintenance
Related MCP Connectors
- docs2mcpOAuthcom.docs2mcp
Query your own PDFs and documents from any MCP client. Every answer cites the page it came from.
Remote ChromaDB vector database MCP server with streamable HTTP transport
Cloud or self-hosted knowledge for AI agents: hybrid search, reranking, GraphRAG, scoped MCP tools.
- WauldoOAuthcom.wauldo
Stateless agentic tools over MCP: concept extraction, long-context, knowledge graph, planning.
Related MCP Servers
- AlicenseAqualityDmaintenanceProvides local Retrieval-Augmented Generation (RAG) capabilities using Ollama for embeddings and ChromaDB for vector storage. It enables users to ingest and perform semantic searches across PDF, Markdown, and TXT documents within MCP-compatible clients.422 npmMIT
- AlicenseAqualityCmaintenanceRAG MCP server for PDFs, YouTube, GitHub repos, and Discord exports. Index documents and query with citations via LangChain and Chroma.52MIT
- AlicenseAqualityDmaintenanceLocal-first RAG indexing and semantic search MCP server. Enables document retrieval and context-aware queries using local embedding models.35 npmMIT
- AlicenseBqualityAmaintenanceLocal end-to-end RAG system for agentic code editors, exposing retrieval-augmented generation via MCP to any compatible client.331MIT