hybrid-rag-mcp
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@hybrid-rag-mcpSearch the docs for details on cross-encoder reranking."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
hybrid-rag-mcp
A production-minded Model Context Protocol (MCP) server that gives any LLM agent
a high-quality search_docs tool backed by hybrid retrieval (BM25 + dense vectors,
fused with Reciprocal Rank Fusion) and cross-encoder reranking — with a real
evaluation harness, prompt-injection guardrails, and OpenTelemetry tracing
built in.
One tool, exposed over MCP, that is measurably good and demonstrably safe.
Documentation
Full documentation lives in docs/. Quick links:
Get started | |
Reference | |
Deep dives | |
Extend & troubleshoot |
Related MCP server: docs2db-mcp-server
Why this project
Most retrieval demos stop at "embed the query, take cosine top-k." Real systems don't. This repo shows the parts that actually matter in production:
Capability | What it demonstrates |
MCP server ( | Tool-calling integration any MCP client (Claude Desktop, IDEs, custom agents) can use |
Hybrid retrieval (BM25 + dense, RRF fusion) | You understand lexical vs. semantic search and how to combine them |
Cross-encoder reranking | You can improve precision@k, not just recall — and measure it |
Eval harness (recall@k, MRR) | You prove quality with numbers, before/after each stage |
Prompt-injection guardrail | You treat tool inputs as untrusted (OWASP LLM01) |
OpenTelemetry tracing | You can debug and observe an agent tool in production |
Architecture
flowchart LR
A[MCP Client] -- search_docs query --> B[FastMCP Server]
B --> G{Injection guardrail}
G -- flagged --> R[Reject + reason]
G -- clean --> P[Retrieval pipeline]
subgraph P [Retrieval pipeline]
C[BM25 lexical] --> F[RRF fusion]
D[Dense vector] --> F
F --> E[Cross-encoder rerank]
end
P --> H[Top-k passages]
H --> A
B -. spans .-> T[(OpenTelemetry)]Quickstart
Requires Python 3.10+ (tested on 3.13).
# 1. Create an isolated environment
python3.13 -m venv .venv && source .venv/bin/activate
# 2. Install the core (BM25 works immediately — no model downloads)
pip install -e .
# 3. Run the tests and the retrieval eval
make test
make eval
# 4. Start the MCP server (stdio)
make runThe server ships with a small sample corpus (src/hybrid_rag_mcp/corpus/sample_docs.jsonl)
so everything runs end-to-end on first clone. Swap in your own corpus to make it yours.
Optional: enable dense vectors + reranking
The advanced retrieval stages activate automatically when their (heavier) dependencies are installed; otherwise the pipeline gracefully degrades to BM25-only.
pip install -e ".[full]" # sentence-transformers, flashrank, opentelemetryConnect it to an MCP client
Add this to your client's MCP config (example for Claude Desktop
claude_desktop_config.json):
{
"mcpServers": {
"hybrid-rag": {
"command": "/absolute/path/to/.venv/bin/python",
"args": ["-m", "hybrid_rag_mcp.server"]
}
}
}Evaluation
make eval scores retrieval quality on evals/qa_dataset.jsonl and prints a table so you
can see the contribution of each stage. Reranking should lift precision — prove it:
Config | Recall@5 | MRR |
BM25 only | run | . |
+ dense (RRF fusion) | . | . |
+ cross-encoder rerank | . | . |
Fill this table with your real numbers and screenshot it in your write-up. Numbers win interviews.
Security / red-teaming
make redteam runs a battery of prompt-injection payloads (redteam/injection_payloads.jsonl)
through the input guardrail and reports how many were caught. Extend the payload set and the
detector rules — closing the gap between them is the interesting part.
Observability
Every tool call is wrapped in an OpenTelemetry span (query, stage latencies, result count).
By default spans print to the console; point OTEL_EXPORTER_OTLP_ENDPOINT at a collector
(Jaeger, Grafana Tempo, Langfuse) to visualize traces.
Repository layout
hybrid-rag-mcp/
├── src/hybrid_rag_mcp/
│ ├── server.py # FastMCP entrypoint, exposes search_docs
│ ├── pipeline.py # composes guardrail -> retrieve -> fuse -> rerank
│ ├── retrieval/ # bm25 / vector / hybrid (RRF) / rerank
│ ├── security/ # prompt-injection guardrail
│ ├── observability/ # OpenTelemetry tracing helpers
│ └── corpus/ # sample corpus (jsonl)
├── evals/ # recall@k + MRR harness, QA dataset, promptfoo config
├── redteam/ # injection payloads + runner
└── tests/ # pytestRoadmap — make it yours
These are deliberately left for you to implement and defend in interviews:
Replace the sample corpus with a real one (your notes, a docs site, arXiv abstracts).
Add a second MCP tool (e.g.,
fetch_document(id)orsummarize(query)).Swap the embedding model and benchmark quality vs. latency.
Add caching for embeddings and rerank scores; measure cost/latency savings.
Wire traces into Jaeger or Langfuse and add a screenshot to the README.
Expand the red-team set and report your catch rate over time.
Migrate to the MCP SDK 2.0
MCPServerAPI once it stabilizes (currently pinned to the stable 1.xFastMCPline for maximum tutorial/Claude Desktop compatibility).
License
MIT — see LICENSE.
Available Tools
1 toolsearch_docsA
Search the document corpus with hybrid retrieval and reranking.
Args: query: A natural-language search query. top_k: How many passages to return (default 5).
Returns:
Passages as {id, title, text, score} objects, most relevant first.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | ||
| top_k | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must carry the full behavioral burden. It does disclose the return format ({id, title, text, score}) and ordering ('most relevant first'), and mentions 'hybrid retrieval and reranking.' However, it does not explicitly state whether the operation is read-only, nor does it mention behavior for empty results, invalid queries, or rate limits. These omissions are non-trivial for a tool with no annotation safety hints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and well-structured: a one-sentence purpose, then Args, then Returns. Every line earns its place, and the format matches the docstring convention, making it easy to parse. No redundant flourishes or unnecessary detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with only two parameters and a clearly described output, the description is nearly complete. It covers purpose, parameters, return shape, and relevance ordering. The main gaps are edge-case behavior (no results, malformed queries) and lack of an explicit read-only statement, but these are minor relative to the simplicity of the tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description fully compensates. It explicitly defines 'query' as a natural-language search query and 'top_k' as how many passages to return with its default of 5. This gives the agent actionable semantics that the bare input schema lacks, making both parameters self-explanatory.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific action: 'Search the document corpus with hybrid retrieval and reranking.' It names a clear verb, a clear resource, and even the retrieval technique, making the tool's function unambiguous. There are no sibling tools to distinguish from, but the purpose is stated with enough specificity that none is needed.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly identifies the tool's context: it is for searching the document corpus using a natural-language query. It implies when an agent would use it (whenever document search is needed) and does not require exclusions since no sibling tools exist. It stops short of explicit 'Use this when...' phrasing, but the context is clear and the parameters reinforce the intended usage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
v0.1.0- First observed
search_docs
TDQS
Scored across 1 tool
With only one tool, there is no possibility of confusion or overlap between tools. The single tool's purpose is clear and distinct by definition.
The single tool name 'search_docs' follows a consistent verb_noun convention. Since there is only one tool, no conflicting patterns exist.
One tool is on the thin edge of what is acceptable. For a narrow search-only server this could be sufficient, but for a 'hybrid-rag' service a single tool feels minimal.
The server provides only search, with no document ingestion, update, deletion, or management capabilities. In a typical RAG workflow, agents would need upload or indexing tools to make the search useful, so the surface has significant gaps.
Maintenance
Related MCP Connectors
MCP server for langchain documentation, generated by doc2mcp.
Read-only MCP server for the OrchestKit docs: full-text search + Markdown fetch. No auth.
MCP server for building and testing AI agents with multi-model experimentation and insights.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceAn MCP server implementation that provides tools for retrieving and processing documentation through vector search, enabling AI assistants to augment their responses with relevant documentation context8 npm265MIT
- AlicenseNot gradedqualityDmaintenanceMCP server for semantic and hybrid search over RHEL documentation using docs2db RAG, with cross-encoder reranking and support for multiple MCP clients.4Apache 2.0
- AlicenseNot gradedqualityDmaintenanceAn MCP server that indexes documents and serves relevant context to LLMs via Retrieval Augmented Generation (RAG).15 npm37MIT
- AlicenseNot gradedqualityAmaintenanceAn MCP server that provides tools for retrieving and processing documentation through vector search, enabling AI assistants to augment their responses with relevant documentation context.19 npmMIT