Skip to main content
Glama
RITIKA-SHARMAA

RAG-MCP

RAG-MCP

A Model Context Protocol server that gives an LLM client searchable access to a local document corpus. It indexes a directory of Markdown and text files with BM25 at startup and exposes four tools over stdio: search, read a passage, read a document, list the corpus.

The retrieval is the whole product. Generation is the client's job, so this server never calls a model and needs no API key.

Quick start

Requires Python 3.11 or newer.

python -m venv .venv
.venv/bin/pip install -e ".[dev]"
.venv/bin/python -m pytest          # 67 tests
RAG_CORPUS_DIR=corpus .venv/bin/python -m rag_mcp

Run on its own it will sit waiting for JSON-RPC on stdin, which is correct: an MCP server is launched by its client, not used directly.

Connecting it to a client

Add it to the client's MCP server config. For Claude Desktop that is claude_desktop_config.json; for Claude Code it is .mcp.json in the project.

{
  "mcpServers": {
    "rag-mcp": {
      "command": "/absolute/path/to/.venv/bin/python",
      "args": ["-m", "rag_mcp"],
      "env": { "RAG_CORPUS_DIR": "/absolute/path/to/your/docs" }
    }
  }
}

Point RAG_CORPUS_DIR at any directory of notes. The corpus/ directory in this repo holds three documents about MCP and retrieval so the server does something useful the first time it starts.

Related MCP server: agent-bm25-knowledge-mcp

Tools

Tool

Arguments

Returns

search_documents

query, optional top_k (1 to 25)

Ranked passages with chunk id, source file, BM25 score and a snippet centred on the match

get_chunk

chunk_id

The full passage plus its character offsets in the source file

get_document

source

The whole document, truncated with a flag if very large

list_documents

none

Every indexed document with size and chunk count, plus anything that was skipped

Responses are JSON in a text block. Two details are deliberate. Search returns the chunk_id so the model can fetch the exact passage before quoting it, and an empty result set comes back with an explicit note saying nothing matched, rather than as an empty list a model might paper over.

How retrieval works

Tokenising. Lowercase, split on non-word characters, drop tokens shorter than two characters and a small stop list, then strip one trailing plural s. Documents and queries go through the same function, which matters more than it sounds: if the two sides normalise differently, terms silently stop matching and results get quietly worse rather than visibly broken.

Chunking. Documents are split into overlapping windows of 180 words with 40 words of overlap, both configurable. Overlap exists so a fact sitting on a chunk boundary still appears whole in one chunk. Each chunk keeps the character offsets of its span, so any result can be traced back to an exact region of the source.

Ranking. BM25 Okapi, k1 = 1.5, b = 0.75:

idf(t) = ln(1 + (N - df(t) + 0.5) / (df(t) + 0.5))

score(D,Q) = sum over t in Q of
             idf(t) * f(t,D) * (k1 + 1) / (f(t,D) + k1 * (1 - b + b * |D| / avgdl))

k1 controls how fast term frequency saturates, so the tenth occurrence of a word adds far less than the second. b controls length normalisation, so a long document does not win simply by containing more words.

Design decisions

BM25 rather than embeddings. No model download, no vector store, no API key, so the server starts in the time it takes to walk a directory and runs anywhere the client runs. Every score is explainable from the formula, so a surprising ranking is debuggable. The cost is real: BM25 matches words, not meaning, and a query phrased entirely in synonyms will miss. On technical documentation, where the person asking usually shares the vocabulary of the text, that trade is worth taking. It would not be for a corpus of customer emails.

The index is in memory and built once. For a documentation corpus this is a few megabytes and it removes a whole class of failure: no stale index, no vector store to be unreachable. The cost is that changing a file means restarting the server.

Failing loudly at startup. A missing corpus directory exits non-zero instead of serving an empty index. An empty index answers every query with "no results", which reads as a retrieval bug and hides a configuration mistake.

Logging to stderr, always. Stdout carries the JSON-RPC stream. One stray print corrupts the protocol and the client disconnects with a parse error that points nowhere near the cause.

Dispatch separated from transport. tools.py takes a corpus, a tool name and a dict, and returns content. It has no session object in it, so every behaviour is unit-testable without starting a server, and server.py stays pure wiring.

Tests

.venv/bin/python -m pytest -q

67 tests. They cover the tokeniser's normalisation rules, chunk overlap and offset round-tripping, BM25 properties that are easy to get wrong (term frequency saturation, length normalisation, non-negative idf, deterministic tie breaks), corpus building including skipped and non-UTF-8 files, every tool's success and failure path, the server handlers themselves, verifying that a bad tool call comes back as a tool error the model can read rather than a crashed connection, and one end to end test that runs a real client session against the server over in-memory streams. That last one is the only test that catches a mistake in the wiring itself, such as a handler that is never registered or a result shape the client cannot parse.

Built against the 2.x Python SDK, where the low level Server takes its handlers as constructor callbacks rather than decorators and Tool uses input_schema rather than inputSchema. The dependency is pinned to <3 so a future major release cannot silently break it.

CI runs the suite on Python 3.11, 3.12 and 3.13 on every push.

What is not built

Stated plainly, because a README that hides its gaps is worse than no README.

  • No semantic search. Lexical matching only. See the trade above.

  • No stemming beyond plurals. "retrieval" and "retrieve" are separate terms.

  • No incremental reindexing. The index reflects the corpus at startup. A file watcher with an incremental rebuild is the obvious next step.

  • No hybrid ranking or reranking. A second pass over the top results would likely improve precision more than any tuning of k1 and b.

  • No PDF or HTML ingestion. Markdown, text and reStructuredText only.

  • Whole-corpus scan per query. Scoring iterates every chunk rather than using an inverted index with posting lists. Fine at this size, wrong at a large one.

Available Tools

4 tools
get_chunkA

Return the full text of one passage by its chunk id, as returned by search_documents, together with its character offsets in the source file.

ParametersJSON Schema
NameRequiredDescriptionDefault
chunk_idYesChunk id in the form 'path/to/file.md#3'.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses the return content (full text and character offsets) and implies a read-only operation via 'Return'. It does not mention error handling or side effects, but for a simple retrieval tool, this is adequate. The explicit mention of offsets adds useful behavioral detail beyond the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, well-structured sentence that front-loads the action ('Return the full text') and then specifies the id source and output details. No wasted words; every phrase contributes.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter retrieval tool with no output schema and no annotations, this description covers the essential behavior (returning text and offsets) and the id's source. It does not explain error conditions or size limits, but those are minor for this simple tool. Overall, an agent can successfully invoke it with confidence.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so baseline is 3. The description adds value by specifying that the chunk_id comes from search_documents, clarifying its provenance and how to obtain it. This is beyond the schema's format description. It doesn't deeply elaborate on the parameter's format, but the provenance cue is useful.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific verb ('Return'), the resource ('full text of one passage'), and the identifier source ('as returned by search_documents'). It distinguishes get_chunk from siblings by focusing on a single passage by chunk id, which is distinct from get_document (whole document) or list_documents/search_documents (listing/searching).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'as returned by search_documents' provides clear usage context, implying this tool is used after a search to fetch a specific chunk. It does not explicitly name alternatives or when-not-to-use, but the context is strong enough for an agent to infer correct usage. It lacks explicit exclusions, which prevents a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_documentA

Return the full text of one indexed document by its source path. Large documents are truncated, and the response says so.

ParametersJSON Schema
NameRequiredDescriptionDefault
sourceYesSource path relative to the corpus root.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full behavioral burden. It discloses that large documents are truncated and that the response indicates this, which is a key behavioral trait. It does not cover all possible behaviors (e.g., error handling), but it adds meaningful transparency beyond the schema, so a 4 is appropriate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences with zero waste. The primary purpose is front-loaded, and the truncation caveat is placed second. It is concise, well-structured, and every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter tool with no output schema, the description is largely complete. It states what the tool does, the key identifier, and the major limitation (truncation). It does not explain what 'says so' means in the response or mention error cases, but given the simplicity of the tool, this is adequate and likely sufficient for an agent to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% – the parameter 'source' is described as 'Source path relative to the corpus root.' The description adds no additional semantic meaning beyond that, just restating the role of the parameter. Per the baseline rule for high schema coverage, a 3 is correct.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'return', the resource 'full text of one indexed document', and the identifier 'source path'. It distinguishes this from siblings like get_chunk (chunk retrieval) and list_documents (listing) by specifying it returns the entire document text. The purpose is unambiguous and specific.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage when the full text of a single document is needed, but it does not explicitly mention alternatives or when not to use it. For example, it does not say 'use get_chunk for a specific chunk' or 'use search_documents for query-based retrieval'. The context is clear but lacks explicit exclusion guidance, so it earns a 3.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_documentsA

List every indexed document with its size and chunk count. Use it to find out what this corpus actually covers before searching.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It states that the tool lists every indexed document with size and chunk count, which conveys the core behavior. However, it does not mention pagination, ordering, or whether the result is returned as a single batch – relevant for large corpora. Since it's a read operation with no side effects, this is acceptable but not exhaustive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is exactly two sentences: the first states the function, the second gives the usage context. It is front-loaded with the core action and every sentence earns its place. There is no fluff or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (0 params, no output schema, no annotations), the description provides sufficient information: what it does (lists documents) and what fields are included (size, chunk count). It also gives usage context. It lacks explicit return format, but the mention of fields implies the structure. For a straightforward list tool, this is adequate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters and the schema is an empty object, so schema description coverage is 100%. With no parameters to describe, the description adds nothing parameter-specific, and the baseline of 4 applies because there is nothing to compensate for.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('list'), a precise resource ('every indexed document'), and the output fields (size and chunk count). It clearly distinguishes this from search_documents (searching), get_chunk (retrieving a chunk), and get_document (retrieving a document) – it is the corpus-level overview tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit guidance on when to use it: 'Use it to find out what this corpus actually covers before searching.' This implies it is a preliminary discovery step and implicitly differentiates from search_documents, which would be used for actual searching. It doesn't explicitly name alternatives, but the context is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_documentsA

Search the indexed corpus for passages relevant to a query and return ranked results with their source file and a snippet. Use this first, then call get_chunk or get_document to read the full text of anything you intend to quote or rely on.

ParametersJSON Schema
NameRequiredDescriptionDefault
queryYesNatural language or keyword query.
top_kNoHow many passages to return. Defaults to 5.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It states what the tool returns (ranked results with source file and snippet) and implies a read-only search operation. It does not mention side effects, permissions, or rate limits, but for a search tool this is reasonably transparent. It lacks explicit confirmation of read-only nature, but the term 'search' strongly implies it.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, zero wasted words. The first sentence states the core function and output; the second provides usage guidance. The purpose is front-loaded, and every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with two parameters and no output schema, the description covers the essential elements: what it does, what it returns (source file and snippet), and how to proceed for full text. It lacks explicit detail on the exact structure of the ranked results or error conditions, but it is sufficient for an agent to invoke it correctly. The reference to sibling tools adds necessary routing context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already describes both parameters (query as natural language/keyword, top_k with min/max/default). The description adds context about the output (source file, snippet) and mentions 'ranked results' but does not add semantic detail about the parameters beyond what the schema provides. With 100% schema coverage, the baseline is 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action: 'Search the indexed corpus for passages relevant to a query' and specifies the output: 'return ranked results with their source file and a snippet.' It also differentiates from siblings by naming get_chunk and get_document as follow-up tools, so an agent can distinguish this retrieval tool from content-access tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit guidance is provided: 'Use this first, then call get_chunk or get_document to read the full text of anything you intend to quote or rely on.' This tells the agent when to use this tool and what to use next, and implicitly when not to use it (when you already have a document, use get_chunk/get_document). No ambiguity remains.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 4 tool updatesv0.1.0
    • First observedget_chunk
    • First observedget_document
    • First observedlist_documents
    • First observedsearch_documents

TDQS

A4.4/5.0

Scored across 4 tools

Disambiguation5/5

Each tool has a clearly distinct purpose: search returns ranked passages, get_chunk retrieves a specific chunk, get_document retrieves a full document, and list_documents provides an overview. There is no overlap or ambiguity between them.

Naming Consistency5/5

All tool names follow the same verb_noun pattern with lowercase and underscores: search_documents, get_chunk, get_document, list_documents. The convention is perfectly consistent.

Tool Count5/5

Four tools form a tight, well-scoped set for a RAG server. Each tool earns its place by covering the essential retrieval workflow without redundancy or bloat.

Completeness5/5

The tool surface covers the full retrieval lifecycle: discover what's indexed (list_documents), search the corpus (search_documents), read a specific passage (get_chunk), and read the full source (get_document). No obvious gaps exist for the stated purpose.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    B
    maintenance
    Local-first Markdown vault retrieval for agents. Read-only MCP stdio server exposing search, get, status, and doctor over Obsidian-compatible Markdown with hybrid BM25/vector/wikilink/title retrieval and first-class CJK support.
    11
    MIT
  • A
    license
    A
    quality
    B
    maintenance
    Enables fast, low-token code search for AI coding agents via a local BM25 engine built on SQLite FTS5, with support for camelCase, snake_case, and Japanese text. Provides a stateless MCP stdio server and a Hermes adapter for multi-agent environments.
    1
    MIT
  • A
    license
    A
    quality
    A
    maintenance
    Local deterministic BM25 memory for AI agents — offline-first, no API key, SHA-256 content-addressed shards, stdio MCP transport. Same query always returns the same ranked result.
    11
    42 npm
    MIT