Skip to main content
Glama
udarshmarthala

hybrid-rag-mcp

hybrid-rag-mcp

A production-minded Model Context Protocol (MCP) server that gives any LLM agent a high-quality search_docs tool backed by hybrid retrieval (BM25 + dense vectors, fused with Reciprocal Rank Fusion) and cross-encoder reranking — with a real evaluation harness, prompt-injection guardrails, and OpenTelemetry tracing built in.

One tool, exposed over MCP, that is measurably good and demonstrably safe.


Documentation

Full documentation lives in docs/. Quick links:

Overview · Installation · Usage

Get started

Architecture · API Reference · Configuration

Reference

Retrieval · Evaluation · Security · Observability

Deep dives

Development · FAQ

Extend & troubleshoot


Related MCP server: docs2db-mcp-server

Why this project

Most retrieval demos stop at "embed the query, take cosine top-k." Real systems don't. This repo shows the parts that actually matter in production:

Capability

What it demonstrates

MCP server (FastMCP, stdio transport)

Tool-calling integration any MCP client (Claude Desktop, IDEs, custom agents) can use

Hybrid retrieval (BM25 + dense, RRF fusion)

You understand lexical vs. semantic search and how to combine them

Cross-encoder reranking

You can improve precision@k, not just recall — and measure it

Eval harness (recall@k, MRR)

You prove quality with numbers, before/after each stage

Prompt-injection guardrail

You treat tool inputs as untrusted (OWASP LLM01)

OpenTelemetry tracing

You can debug and observe an agent tool in production

Architecture

flowchart LR
    A[MCP Client] -- search_docs query --> B[FastMCP Server]
    B --> G{Injection guardrail}
    G -- flagged --> R[Reject + reason]
    G -- clean --> P[Retrieval pipeline]
    subgraph P [Retrieval pipeline]
        C[BM25 lexical] --> F[RRF fusion]
        D[Dense vector] --> F
        F --> E[Cross-encoder rerank]
    end
    P --> H[Top-k passages]
    H --> A
    B -. spans .-> T[(OpenTelemetry)]

Quickstart

Requires Python 3.10+ (tested on 3.13).

# 1. Create an isolated environment
python3.13 -m venv .venv && source .venv/bin/activate

# 2. Install the core (BM25 works immediately — no model downloads)
pip install -e .

# 3. Run the tests and the retrieval eval
make test
make eval

# 4. Start the MCP server (stdio)
make run

The server ships with a small sample corpus (src/hybrid_rag_mcp/corpus/sample_docs.jsonl) so everything runs end-to-end on first clone. Swap in your own corpus to make it yours.

Optional: enable dense vectors + reranking

The advanced retrieval stages activate automatically when their (heavier) dependencies are installed; otherwise the pipeline gracefully degrades to BM25-only.

pip install -e ".[full]"   # sentence-transformers, flashrank, opentelemetry

Connect it to an MCP client

Add this to your client's MCP config (example for Claude Desktop claude_desktop_config.json):

{
  "mcpServers": {
    "hybrid-rag": {
      "command": "/absolute/path/to/.venv/bin/python",
      "args": ["-m", "hybrid_rag_mcp.server"]
    }
  }
}

Evaluation

make eval scores retrieval quality on evals/qa_dataset.jsonl and prints a table so you can see the contribution of each stage. Reranking should lift precision — prove it:

Config

Recall@5

MRR

BM25 only

run make eval

.

+ dense (RRF fusion)

.

.

+ cross-encoder rerank

.

.

Fill this table with your real numbers and screenshot it in your write-up. Numbers win interviews.

Security / red-teaming

make redteam runs a battery of prompt-injection payloads (redteam/injection_payloads.jsonl) through the input guardrail and reports how many were caught. Extend the payload set and the detector rules — closing the gap between them is the interesting part.

Observability

Every tool call is wrapped in an OpenTelemetry span (query, stage latencies, result count). By default spans print to the console; point OTEL_EXPORTER_OTLP_ENDPOINT at a collector (Jaeger, Grafana Tempo, Langfuse) to visualize traces.

Repository layout

hybrid-rag-mcp/
├── src/hybrid_rag_mcp/
│   ├── server.py            # FastMCP entrypoint, exposes search_docs
│   ├── pipeline.py          # composes guardrail -> retrieve -> fuse -> rerank
│   ├── retrieval/           # bm25 / vector / hybrid (RRF) / rerank
│   ├── security/            # prompt-injection guardrail
│   ├── observability/       # OpenTelemetry tracing helpers
│   └── corpus/              # sample corpus (jsonl)
├── evals/                   # recall@k + MRR harness, QA dataset, promptfoo config
├── redteam/                 # injection payloads + runner
└── tests/                   # pytest

Roadmap — make it yours

These are deliberately left for you to implement and defend in interviews:

  • Replace the sample corpus with a real one (your notes, a docs site, arXiv abstracts).

  • Add a second MCP tool (e.g., fetch_document(id) or summarize(query)).

  • Swap the embedding model and benchmark quality vs. latency.

  • Add caching for embeddings and rerank scores; measure cost/latency savings.

  • Wire traces into Jaeger or Langfuse and add a screenshot to the README.

  • Expand the red-team set and report your catch rate over time.

  • Migrate to the MCP SDK 2.0 MCPServer API once it stabilizes (currently pinned to the stable 1.x FastMCP line for maximum tutorial/Claude Desktop compatibility).

License

MIT — see LICENSE.

Available Tools

1 tool
search_docsA

Search the document corpus with hybrid retrieval and reranking.

Args: query: A natural-language search query. top_k: How many passages to return (default 5).

Returns: Passages as {id, title, text, score} objects, most relevant first.

ParametersJSON Schema
NameRequiredDescriptionDefault
queryYes
top_kNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must carry the full behavioral burden. It does disclose the return format ({id, title, text, score}) and ordering ('most relevant first'), and mentions 'hybrid retrieval and reranking.' However, it does not explicitly state whether the operation is read-only, nor does it mention behavior for empty results, invalid queries, or rate limits. These omissions are non-trivial for a tool with no annotation safety hints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and well-structured: a one-sentence purpose, then Args, then Returns. Every line earns its place, and the format matches the docstring convention, making it easy to parse. No redundant flourishes or unnecessary detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with only two parameters and a clearly described output, the description is nearly complete. It covers purpose, parameters, return shape, and relevance ordering. The main gaps are edge-case behavior (no results, malformed queries) and lack of an explicit read-only statement, but these are minor relative to the simplicity of the tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description fully compensates. It explicitly defines 'query' as a natural-language search query and 'top_k' as how many passages to return with its default of 5. This gives the agent actionable semantics that the bare input schema lacks, making both parameters self-explanatory.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific action: 'Search the document corpus with hybrid retrieval and reranking.' It names a clear verb, a clear resource, and even the retrieval technique, making the tool's function unambiguous. There are no sibling tools to distinguish from, but the purpose is stated with enough specificity that none is needed.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly identifies the tool's context: it is for searching the document corpus using a natural-language query. It implies when an agent would use it (whenever document search is needed) and does not require exclusions since no sibling tools exist. It stops short of explicit 'Use this when...' phrasing, but the context is clear and the parameters reinforce the intended usage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool updatev0.1.0
    • First observedsearch_docs

TDQS

A4.1/5.0

Scored across 1 tool

Disambiguation5/5

With only one tool, there is no possibility of confusion or overlap between tools. The single tool's purpose is clear and distinct by definition.

Naming Consistency5/5

The single tool name 'search_docs' follows a consistent verb_noun convention. Since there is only one tool, no conflicting patterns exist.

Tool Count3/5

One tool is on the thin edge of what is acceptable. For a narrow search-only server this could be sufficient, but for a 'hybrid-rag' service a single tool feels minimal.

Completeness2/5

The server provides only search, with no document ingestion, update, deletion, or management capabilities. In a typical RAG workflow, agents would need upload or indexing tools to make the search useful, so the surface has significant gaps.

Maintenance

ActivitySlowing
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    An MCP server implementation that provides tools for retrieving and processing documentation through vector search, enabling AI assistants to augment their responses with relevant documentation context
    8 npm
    265
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    An MCP server that indexes documents and serves relevant context to LLMs via Retrieval Augmented Generation (RAG).
    15 npm
    37
    MIT
  • A
    license
    Not graded
    quality
    A
    maintenance
    An MCP server that provides tools for retrieving and processing documentation through vector search, enabling AI assistants to augment their responses with relevant documentation context.
    19 npm
    MIT