Skip to main content
Glama

hybrid-rag-mcp

CI Python 3.11+ License: MIT Release

MCP server with hybrid RAG (vector Qdrant + lexical BM25 via RRF), a multi-step agent with offline fallback via Ollama, and stdio or streamable HTTP transport.

Focus: index technical documents (Markdown, TXT, PDF) and answer questions with cited sources, 100% locally, without sending documents to third parties.

Evidence: recall@1 0.917 · answer judge 1.92/2 · 80 tests · mypy strict · CI with quality gates.

Leia em português.

Contents

Related MCP server: rag-mcp

Highlights

  • Hybrid search: embeddings (local Qdrant, no Docker) + in-house BM25 (smoothed idf), fused by RRF.

  • Multi-step agent: if the first search's context is insufficient, the model signals [MORE_CONTEXT], the agent issues a follow-up search and retries with incremental source memory.

  • Optional re-ranking: cross-encoder (Ollama /api/rerank, e.g. bge-reranker-v2-m3) with graceful degradation.

  • Resilient fallback: cloud provider (OpenAI-compatible) first, local Ollama as backup when the API is down.

  • Persistence: chunks live in Qdrant; the BM25 index is restored on demand (first search) with no re-ingestion.

  • Incremental ingestion: re-running ingest only embeds what changed (content-hash idempotent) and prunes orphans — cheap in CI and redeploys.

  • Token optimization: semantic cache (JSONL + cosine, with TTL) returns previously generated answers without re-calling the LLM; a Caveman-style compressor (PT/EN) trims the source context before the prompt.

  • Traceability: per-step agent trace + audit.jsonl (question, provider, iterations, latency, sources, cache).

  • Measurable: recall@k / nDCG@k evaluation pipeline with a quality gate in CI.

  • Two transports: stdio (local RPC) and streamable HTTP (http://host:port/mcp).

Architecture

flowchart LR
    C["MCP Client<br/>stdio or HTTP"] -->|tools: ingest / search / ask / metrics| M["MCP Server<br/>hybrid-rag-mcp"]
    M --> AGE["Multi-step agent<br/>loop with [MORE_CONTEXT]"]
    M --> I["ingest"]
    I --> C1["Chunker<br/>sections + sentences"]
    C1 --> E["Embeddings<br/>Ollama bge-m3"]
    E --> Q1[("Qdrant<br/>vector search")]
    C1 --> K["In-house BM25<br/>lexical search"]
    AGE --> RET["Hybrid search"]
    RET --> Q1 & K
    Q1 & K --> RRF["RRF fusion"]
    RRF --> RR["Optional reranker<br/>Ollama /api/rerank"]
    RR --> LLM["FallbackLLM<br/>cloud -> Ollama"]
    LLM --> AUD["audit.jsonl<br/>trace + iterations + sources"]

Metrics (quality gate in CI)

Evaluation pipeline over 36 queries — 12 hand-curated in eval/dataset.jsonl + 24 auto-generated from real documents (eval/dataset.real.jsonl, see Corpus). CI trigger: fails if recall@1 < 0.8.

k

recall@k

nDCG@k

1

0.917

0.917

3

1.000

0.865

5

1.000

0.945

Honest numbers over real text: the right source ranks top-1 in 91.7% of cases and always in the top-3. tools/grid_search.py sweeps RRF weights/top_k to reach this result (lexical weight 1.5) — history in eval/grid_results.json. Run locally with python -m hybrid_rag_mcp.eval.

Answer quality (llm-as-judge)

Retrieval proves the right chunk surfaces; this proves the final answer is correct: 12 questions over the corpus (eval/answers.jsonl), each with an expected answer, scored 0/1/2 by Ollama itself — run with python -m hybrid_rag_mcp.judge (needs Ollama up).

metric

value

average

1.92 (0..2)

score 2

11/12

score 1

1/12 (ans-06: missed "telemetry doesn't go to PostgreSQL")

score 0 / no verdict

0

Compression validated: with CONTEXT_COMPRESSION=1 the average holds at 1.92; with 2, 2.00 (12×2) — pruning function words does not degrade answers (cache off in all 3 runs; n=12, a 1-item swing is noise — the conclusion is "doesn't hurt", not "helps").

Scale (retrieval, embedded mode)

Methodology (tools/scale_bench.py): 11 real docs replicated with unique salt up to ~25k chunks — recall stays on the curated eval; here only latency, throughput and disk (36 real queries, no LLM). Isolated index, embedded Qdrant on a 4GB RAM box.

metric

32,760 indexed chunks

ingest

566s (58 chunks/s), 337MB index on disk

1-thread search

p50 191ms · p95 621ms · mean 373ms

8-thread search

p50 640ms · p95 4308ms

p99 single ≈5s

one-time warmup cost (first search restores the BM25 index; varies across runs)

Re-measured with vectorized BM25 (before: p50 269/p95 663 single, p50 1313/p95 2334 concurrent). p50 fell ~30% single-threaded and ~50% concurrent — numpy releases the GIL on the vectorized path; tails (p95/p99) swing ±2x between runs on this box, noise not signal.

Honest readings (including our own correction): the initial guess was that pure-Python BM25 was the bottleneck — measured, it wasn't: at 32k docs score_all cost ~17ms of ~400ms per search (the bulk is Ollama embedding + local HNSW). We vectorized anyway (~10x: 17.5→1.7ms on 30k synthetic docs, parity < 1e-9): the win compounds at 10x scale, where the pure loop would dominate. Contention under concurrency (~5x with 8 threads — GIL + local Qdrant + embedding on Ollama) stands. Known ceiling: at 128k chunks the box OOMed — above ~100k chunks or on little RAM, use server mode (QDRANT_URL, see Deploy).

Context optimization (cache and compression)

Semantic cache — before generating, ask consults data/cache.jsonl in two layers:

  1. exact normalization (identical questions, ignoring case/whitespace) and

  2. similarity — the question is embedded (same bge-m3 as search) and compared by cosine against the entries; above CACHE_SIM_THRESHOLD (0.92) the saved answer is returned (TTL CACHE_TTL_SEC, CACHE_MAX_ENTRIES cap). A hit skips generation entirely — the biggest token saving. Hits are marked · cache in the answer and recorded in the audit ("cache": true).

Caveman-style context compression — CONTEXT_COMPRESSION (0/1/2) removes predictable function words (connectives/fillers at level 1; + articles and auxiliaries at level 2) only from the copy that goes into the LLM prompt: search, re-ranking and displayed sources keep the original text, and numbers, proper nouns and negations are never removed. Deterministic, multilingual (PT/EN), zero dependencies — principle inspired by Caveman ("strip grammar, keep facts"), reimplemented in-house.

python tools/optimizers_report.py        # dashboard: cache + compression + evaluated integrations
python tools/optimizers_report.py --json # same output as JSON

Corpus

  • examples/corpus/ — 11 documents: 5 fictional (operations/security/database/infra/events) + 6 real public-domain books (Project Gutenberg): Chekhov, Machado de Assis, Aluísio Azevedo, Eça de Queirós, Jane Austen and Conan Doyle, mixing PT and EN.

  • tools/fetch_corpus.py — downloads the Gutenberg catalog and auto-generates the evaluation dataset: each query is a real excerpt from a document and the expected document is that excerpt's exact source (self-supervised ground truth, no manual curation).

Running

Prerequisites: Python 3.11+, running Ollama (ollama serve).

# 1. Local models (once)
ollama pull bge-m3        # embeddings
ollama pull qwen3:8b      # generation (or another)
ollama pull bge-reranker-v2-m3   # optional: Ollama >= 0.36 only

# 2. Install
python -m venv .venv && source .venv/bin/activate
pip install -e ".[dev]"

# 3. ~1 min demo (ingest + search + ask)
bash examples/demo.sh
[ingest] 11 docs, 1964 chunks (+0 novos, 1964 no índice)
[search] 'Qual a porta padrão do servidor?'
  - [pqpbr-operations.txt] score=0.065: ## Configuração de Rede O servidor escuta na porta 8443...
[ask] 'De quantas em quantas horas são os backups?'
  [ollama/qwen3:8b] Os backups incrementais são executados a cada 4 horas [pqpbr-operations.txt].
  fontes: gutenberg-dom-casmurro.txt, pqpbr-operations.txt

As an MCP client

stdio: example client — python examples/client.py "Qual a porta padrão?"

Progress streaming (tokens + steps): python examples/client_stream.py "Qual a política de manutenção do banco?" — the server emits step events and response tokens in real time via notifications/progress (the client sends a progressToken in the request).

HTTP:

# terminal 1
python -m hybrid_rag_mcp --transport http --host 127.0.0.1 --port 8000

# terminal 2
python examples/client_http.py "Qual a porta padrão?"

HTTP auth (recommended when exposing on a network): generate a token (openssl rand -hex 32), export MCP_AUTH_TOKEN on the server and the client — without the Authorization: Bearer header the server answers 401. With no token configured, it behaves as before (localhost only).

MCP_AUTH_TOKEN=... python -m hybrid_rag_mcp --transport http --port 8000
MCP_AUTH_TOKEN=... python examples/client_http.py "Qual a porta padrão?"

Register in any MCP client (Claude Desktop, editors, agents):

{
  "mcpServers": {
    "hybrid-rag": {
      "command": ".venv/bin/python",
      "args": ["-m", "hybrid_rag_mcp"],
      "env": { "PYTHONPATH": "src" }
    }
  }
}

Cloud fallback (optional)

Copy .env.example to .env and fill in CLOUD_BASE_URL + CLOUD_API_KEY + CLOUD_MODEL (any OpenAI-compatible endpoint). Cloud takes priority; Ollama answers automatically if the API fails or goes offline. For cloud embeddings, also set CLOUD_EMBED_MODEL — note: switching models changes the dimension and requires recreating the index (re-run ingest). For a multi-API cascade (e.g. Nvidia → Gemini → Ollama), use LLM_CHAIN with the list in priority order — the first to answer wins, you pay only it per question.

Deploy

systemd (bare metal): deploy/hybrid-rag.service is a user unit that brings up the HTTP transport on 127.0.0.1:8000 with automatic restart — adjust the paths (%h = your home) and install with systemctl --user enable --now pointing at the file. .env is optional (cloud/token only if you export them).

Docker (Qdrant + RAG): docker compose up -d --build brings up Qdrant and the server already pointed at it (QDRANT_URL=http://qdrant:6333, Ollama via host.docker.internal). No Docker installed here so the compose was not run locally — validated by inspection; CI keeps covering embedded mode.

Data hygiene: audit.jsonl rotates by size (AUDIT_MAX_BYTES, default 5MB, keeps AUDIT_KEEP=3 backups); the semantic cache is already capped (CACHE_MAX_ENTRIES).

Observability

In-process metrics, zero dependencies (src/hybrid_rag_mcp/metrics.py): counters (search_total, ask_total, ask_cache_hits, *_errors, ask_provider_<name>) + latencies (search_ms, ask_ms, ingest_ms).

  • HTTP: GET /metrics in Prometheus format (inherits the bearer auth). Scrape example:

    scrape_configs:
      - job_name: hybrid-rag
        static_configs: [{targets: ["127.0.0.1:8000"]}]
        # authorization: {credentials: <MCP_AUTH_TOKEN>}  # if auth is on
  • Both transports: metrics MCP tool with a human summary.

  • Logs: one JSON line per search/ask on stderr ({"op": "ask", "elapsed_ms": 123.4, "cache_hit": false, ...}) — stdout belongs to the MCP protocol in stdio mode and is never polluted.

MCP tools

Tool

Description

ingest

Indexes md/txt/pdf from a directory into both indexes (Qdrant + BM25). Incremental: only re-embeds changed chunks.

search

Hybrid search (RRF, optional re-ranking) returning snippets + sources.

ask

Multi-step agent: retrieves, generates, detects insufficient context, re-searches and answers citing sources (with audit log). Consults the semantic cache before generating.

metrics

Server metrics summary (counters + p50/p95 latencies since boot).

Extending

Nothing here is mandatory — every piece has a local default and can be swapped without forking:

  • Compressor: any (str) -> str callable via CONTEXT_COMPRESSOR=my_package:clean (beats the built-in 0/1/2 levels). E.g. a sidecar like Headroom wrapped in a function.

  • Embeddings/LLM: EmbeddingProvider / LLMProvider ABCs (src/hybrid_rag_mcp/providers/); cloud via CLOUD_* with no code changes. Switching embeddings changes the dimension: wipe the index and re-run ingest, or set QDRANT_RECREATE_ON_DIM_CHANGE=1.

  • Reranker: score(query, texts) -> list[float] duck-type (see providers/rerank.py); empty = off, with graceful degradation.

Structure

src/hybrid_rag_mcp/
├── server.py          # MCP server (stdio + streamable HTTP, optional bearer auth)
├── config.py          # .env configuration (pydantic-settings)
├── eval.py            # recall@k / nDCG@k evaluation
├── judge.py           # Answer evaluation via llm-as-judge (score 0/1/2)
├── bench.py           # Scale-bench blocks (percentiles, synthetic corpus)
├── metrics.py         # In-process metrics (counters + p50/p95, Prometheus)
├── rag/
│   ├── agent.py       # Multi-step loop (source memory, [MORE_CONTEXT])
│   ├── chunker.py     # Chunking by markdown sections + sentences
│   ├── engine.py      # Orchestration: ingest → search → ask + audit
│   ├── hybrid.py      # RRF fusion + re-ranking
│   └── ingestion.py   # md/txt/pdf reading
├── providers/
│   ├── embed.py       # Embeddings via Ollama
│   ├── llm.py         # FallbackLLM (cloud → Ollama)
│   └── rerank.py      # Optional cross-encoder (graceful degradation)
├── optimize/
│   ├── cache.py       # Semantic cache (JSONL + cosine + TTL)
│   └── compress.py    # Caveman-style compressor (PT/EN)
└── stores/
    ├── vector.py      # Embedded Qdrant (no Docker, persistent)
    └── lexic.py       # BM25 with smoothed idf

Architecture decisions with context and evidence: docs/adr/ (RRF, storage, cache, judge, lazy-init, external integrations, cascade scope, BM25 tradeoff, open seams). Contributing guide: CONTRIBUTING.md. Release history: CHANGELOG.md.

Quality

  • 98 unit tests (pytest) with no network/Ollama — chunking, RRF, BM25 (incl. vectorized × brute-force parity), persistence, eval metrics, agent loop, semantic cache, compressor, lexical-index thread-safety, judge parsing/aggregation, HTTP auth, audit rotation, observability, pluggable compressor, cloud embeddings, dimension check and LLM cascade.

  • CI in 2 jobs: test (ruff + mypy strict + pytest with ≥75% coverage + pip-audit + stdio/HTTP smoke) and eval (real Ollama + recall@1 >= 0.8 gate). Weekly Dependabot (pip + actions).

  • Two storage modes: embedded (default, no Docker, 1 process at a time) or server (docker compose up -d + QDRANT_URL=http://localhost:6333) for simultaneous sessions — see docker-compose.yml.

Available Tools

4 tools
askC

Responde com RAG e faz streaming do progresso/tokens via progress notifications.

ParametersJSON Schema
NameRequiredDescriptionDefault
top_kNo
questionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does disclose one valuable behavior: streaming progress/tokens via progress notifications. However, it does not state whether the operation is read-only, how long it may take, or any limitations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no wasted words: it states the core RAG answering behavior first, then the streaming detail. It is appropriately concise, though a bit more structure could improve clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations, no schema descriptions, and sibling tools that could overlap, this description is too thin. It omits parameter semantics, usage context, and any distinction from search; the output schema helps with return values but not with selection or invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the description never mentions question or top_k. The agent must infer their meaning entirely from the schema's bare property names.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies the action (responds) and the method (RAG), and the streaming mention adds a distinguishing behavioral trait. It stops short of a 5 because it does not explicitly contrast with sibling tools like search.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance is provided, and the sibling tools are never referenced. The verb 'ask' offers only an implied use case, which is not enough for an agent to reliably choose this over search.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ingestA

Indexa documentos (md/txt/pdf) de um diretório nos índices vetorial e BM25.

ParametersJSON Schema
NameRequiredDescriptionDefault
corpus_dirNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses the main write behavior (indexing into vector and BM25 stores) but does not mention side effects, such as whether existing documents are replaced, whether indexing is incremental, if the directory is read recursively, or any permission requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with no wasted words. It front-loads the action and packs in file types and target indexes efficiently.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple tool with one parameter and an output schema, the description covers the core purpose and parameter meaning. Gaps remain around usage timing, reindexing behavior, and directory handling, but these are not severe enough to render it unusable.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. 'de um diretório' clarifies that corpus_dir is the source directory, adding meaning to the parameter name alone. However, it does not explain the default behavior, path format, or how recursive/required the directory is.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('Indexa' - indexes), a clear resource (documents md/txt/pdf from a directory), and the target (vector and BM25 indexes). This clearly differentiates it from siblings 'search' and 'ask', which are query operations.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The intended use is implied: call this tool to index documents before searching or asking. However, there is no explicit guidance on when to use it versus alternatives, prerequisites, or scenarios where it should not be used.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

metricsA

Resumo das métricas do servidor (contadores + latências p50/p95 desde o boot).

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does add useful behavioral context: it returns a summary of counters and p50/p95 latencies since boot, implying a read-only metrics retrieval. It does not explicitly state that there are no side effects or describe output shape, but the low-risk nature of the tool makes this acceptable.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single compact sentence front-loads the purpose and packs in the metric categories and time window without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a parameterless tool with an output schema, this description provides all essential invocation knowledge: what is returned and its time scope. Nothing needed to call the tool correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so there is no parameter meaning for the description to add beyond the schema. The no-parameter baseline of 4 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies the tool as a server-metrics summary and specifies the included data (counters and p50/p95 latencies) and the since-boot window, which is enough to separate it from ask/ingest/search. It lacks an explicit action verb such as 'returns' or 'gets', so it is slightly below the top anchor.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given for when to use this tool or when a sibling tool would be preferable. The 'since boot' detail describes the metric window, not a selection rule, leaving the agent to infer usage from the tool name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool update
    • Addedmetrics
  2. 3 tool updatesv0.1.0
    • First observedask
    • First observedingest
    • First observedsearch

TDQS

B3.4/5.0

Scored across 4 tools

Disambiguation4/5

As ferramentas têm propósitos majoritariamente distintos: ingest adiciona documentos, search faz consulta híbrida crua, ask gera resposta com RAG e metrics expõe estatísticas. A única ambiguidade possível é entre ask e search, já que ambos envolvem recuperação, mas as descrições deixam clara a diferença entre resposta gerada e trechos brutos.

Naming Consistency4/5

Os nomes são todos curtos e em minúsculas, seguindo um estilo simples e consistente. Há uma pequena mistura entre verbos (ask, ingest, search) e substantivo (metrics), mas o padrão geral é previsível e uniforme.

Tool Count4/5

Com 4 ferramentas, o servidor é enxuto e cada uma cobre uma função essencial do ciclo RAG: ingestão, busca, geração e observabilidade. O número é apropriado para o escopo declarado, sem excesso ou carência evidente.

Completeness4/5

O fluxo principal de RAG está coberto: ingestão, busca híbrida, perguntas com streaming e métricas. Faltam operações secundárias como remover/atualizar documentos ou listar o corpus, mas essas são lacunas menores que não impedem o uso básico do servidor.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • F
    license
    A
    quality
    D
    maintenance
    Enables indexing local documents (PDF, Markdown, text, code) into a knowledge base and querying them via semantic search using local embeddings, all running privately on your machine.
    4
    -
  • A
    license
    A
    quality
    C
    maintenance
    Enables semantic search and question answering over a knowledge base using hybrid retrieval and grounded answers, all running offline with no API keys.
    4
    MIT
  • A
    license
    A
    quality
    A
    maintenance
    Indexes your project's markdown documentation and exposes it to AI agents via local hybrid search (lexical + semantic) with progressive disclosure tools.
    3
    463 npm
    5
    MIT