Skip to main content
Glama

rag-starter — chat with your documents (RAG), with citations

rag-starter MCP server

License: MIT Python MCP

A production-ready starter that turns a folder of documents into a cited Q&A service. Drop in your PDFs / Markdown / text, ask questions, get answers grounded in the source — every claim traceable to the exact passage it came from.

Exposed two ways from one codebase:

  • MCP server — plug it into Claude Desktop / Claude Code / any MCP host and chat with your docs.

  • HTTP API (FastAPI) — call it from any app.

Built once, reskinned per client. Swap the data/ folder, tweak config.py, ship.

Why it's different

  • Keyless by default. Embeddings run locally (ONNX MiniLM) — no API key, no per-query cost, runs offline. Demo it anywhere in seconds.

  • Citations, not hallucinations. Retrieval returns ranked passages tagged [source#chunk] / [file.pdf p3]. A missing answer returns "Not found in the documents" — never a fabricated one.

  • Answer synthesis is optional. With an ANTHROPIC_API_KEY it writes a cited answer for you; without one it returns passages for the host LLM to answer. Either way the RAG works.

  • Idempotent ingestion. Re-ingesting a file updates it in place (no duplicates).

Related MCP server: PinRAG

Quickstart

pip install -r requirements.txt          # or: pip install -e .
PYTHONUTF8=1 python tests/test_smoke.py   # proves retrieval works on the sample docs

As an HTTP API

rag-starter-api            # uvicorn on 127.0.0.1:8000
# then:
curl -X POST localhost:8000/ingest  -H "content-type: application/json" -d '{"path":"./data"}'
curl -X POST localhost:8000/search  -H "content-type: application/json" -d '{"query":"refund policy"}'
curl -X POST localhost:8000/answer  -H "content-type: application/json" -d '{"query":"how long is the refund window?"}'

As an MCP server (Claude Desktop)

Add to claude_desktop_config.json → mcpServers:

{
  "rag-starter": {
    "command": "python",
    "args": ["-m", "rag_starter"],
    "cwd": "C:/path/to/rag-starter"
  }
}

Restart Claude Desktop, then: "Ingest the folder ./data", then "What's the refund policy?"

Tools / endpoints

MCP tool

HTTP

Purpose

rag_ingest(path)

POST /ingest

Index a file or folder (txt/md/pdf).

rag_search(query, k)

POST /search

Top-k cited passages (keyless).

rag_answer(query, k)

POST /answer

Cited answer (with key) or passages (without).

—

GET /health

Liveness + indexed-chunk count + embedding backend.

Configuration

Everything is in config.py (or env / .env — see .env.example): chunk size & overlap, top-k, embedding backend (default local vs openai), and the answer model. A client reskin is usually just: replace data/, set RAG_COLLECTION, re-ingest.

Architecture

documents ──ingest.py──> chunks(+citations) ──store.py──> Chroma (local, persistent)
                                                              │
question ─────────────────retrieve.py──> top-k cited passages┤
                                                              ├─> rag_search  (host LLM answers)
                                                  answer.py ──┴─> rag_answer  (server synthesizes, optional)

Reuses the mcp-factory contract (Result, formatting, cache) so it composes with the rest of the catalog.

License

MIT.

Available Tools

3 tools
rag_answerA
Read-onlyIdempotent

Answer a question grounded in the documents, with citations.

With ANTHROPIC_API_KEY set, returns a synthesized cited answer; otherwise returns the ranked passages for you to answer from.

Examples: - "Summarize the cancellation terms" -> query='cancellation terms' - "Does the plan include support?" -> query='support included in plan'

ParametersJSON Schema
NameRequiredDescriptionDefault
paramsYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so safety is covered. The description adds genuinely new behavioral context: output varies by environment (synthesized cited answer vs. raw ranked passages), which the structured fields do not convey. It says nothing about the retrieval scope or latency, but the mode-switching disclosure is the important part.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core purpose in the first clause, then adds the mode caveat and two short examples. Every element is functionally useful, though the second sentence could be tightened slightly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return format need not be explained, yet the description still usefully distinguishes the two response shapes. Combined with the parameter schema and annotations, an agent has enough to invoke this correctly; only the sibling relationship to rag_search is left unresolved.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The examples demonstrate how to phrase the `query` argument from a user question, which is real added meaning beyond the schema's field description. However `k` and `response_format` are never mentioned in the description, so those rely entirely on the nested schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise verb+resource+output: answers a question grounded in the provided documents and includes citations. It also clarifies the two possible outputs (synthesized answer vs. ranked passages). It does not name or differentiate from siblings like rag_search, which it partially overlaps with when synthesis is unavailable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explains the condition that selects each mode (ANTHROPIC_API_KEY present vs. absent) and provides worked examples showing how a natural-language question maps to the `query` value. It gives no explicit guidance on when to prefer this over rag_search, leaving the sibling boundary to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

rag_ingestA
Idempotent

Index a document or folder so it can be searched and cited.

Examples: - "Index the folder ./data" -> path='./data' - "Add the handbook pdf" -> path='docs/handbook.pdf'

ParametersJSON Schema
NameRequiredDescriptionDefault
paramsYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare the full safety profile (readOnlyHint=false, idempotentHint=true, destructiveHint=false, openWorldHint=false), so the bar is lower. The description adds the consequence of the write ('so it can be searched and cited'), which is useful, but says nothing about re-indexing behavior, supported formats (left to the schema), or cost/latency of ingestion.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One declarative sentence front-loads the action and its payoff, followed by two short, well-formatted examples. No filler, no restating of the tool name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described, and annotations cover the mutation's safety profile. The remaining gap is operational context (what happens on re-ingest, whether indexing is synchronous, format constraints), which is minor but would help an agent set expectations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Reported schema description coverage is 0%, so the description must compensate, and it does: the two examples map natural-language requests onto concrete path values ('Index the folder ./data' -> path='./data'), clarifying that path accepts both files and folders. The response_format parameter is not addressed, but its schema description is self-explanatory.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Index') and resource ('a document or folder') plus the downstream purpose ('so it can be searched and cited'). It is distinguishable from rag_search/rag_answer by implication (this is the tool that adds content to the index), but it never names the siblings or explicitly contrasts itself with them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied rather than stated: the examples show how to invoke it, but nothing says when to use rag_ingest rather than rag_search or rag_answer, or that ingestion is a prerequisite for those tools. No prerequisites, limits, or exclusions are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 3 tool updatesv0.1.0
    • First observedrag_answer
    • First observedrag_ingest
    • First observedrag_search

TDQS

A3.7/5.0

Scored across 3 tools

Disambiguation4/5

The three tools map to distinct lifecycle stages (index, retrieve, synthesize), but rag_search and rag_answer overlap: both take a query and can return ranked passages, since rag_answer falls back to passages without an API key. The descriptions do clarify the difference, so confusion is limited.

Naming Consistency5/5

All tools use a consistent rag_ prefix followed by a clear verb (ingest, search, answer) in uniform snake_case. The pattern is predictable and immediately readable.

Tool Count4/5

Three tools is a tight, well-scoped set for a RAG starter, each earning its place. It is at the lower end but appropriately minimal for the stated purpose.

Completeness3/5

The core ingest-search-answer flow is covered, but there is no way to remove, delete, or list indexed documents, and no re-index/update operation. These gaps around document lifecycle management could force agents to work around stale content.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    D
    maintenance
    Provides local Retrieval-Augmented Generation (RAG) capabilities using Ollama for embeddings and ChromaDB for vector storage. It enables users to ingest and perform semantic searches across PDF, Markdown, and TXT documents within MCP-compatible clients.
    4
    22 npm
    MIT
  • A
    license
    A
    quality
    D
    maintenance
    Local-first RAG indexing and semantic search MCP server. Enables document retrieval and context-aware queries using local embedding models.
    3
    5 npm
    MIT