hybrid-rag-mcp
This MCP server provides local hybrid RAG over technical documents with indexing, search, question-answering, and metrics.
Index md/txt/pdf documents from a directory into Qdrant vector + BM25 indexes (incremental, idempotent, orphan pruning).
Perform hybrid search (vector + BM25 via RRF) with optional cross-encoder re-ranking.
Ask questions via a multi-step agent that re-searches when context is insufficient and answers with cited sources.
Use local Ollama by default, with optional cloud fallback/cascade (OpenAI-compatible).
Reduce token use with a semantic cache and Caveman-style context compression.
Observe with traces, audit.jsonl, counters, p50/p95 latencies, and Prometheus /metrics.
Expose MCP tools: ingest, search, ask, metrics.
Connect over stdio or streamable HTTP, with optional bearer auth and progress streaming.
Evaluate retrieval/answer quality (recall@k, nDCG@k, llm-as-judge) and run quality gates in CI.
Extend embeddings, LLM, reranker, and compressor via pluggable providers.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@hybrid-rag-mcpWhat is the default port for the server?"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
hybrid-rag-mcp
MCP server with hybrid RAG (vector Qdrant + lexical BM25 via RRF), a multi-step agent with offline fallback via Ollama, and stdio or streamable HTTP transport.
Focus: index technical documents (Markdown, TXT, PDF) and answer questions with cited sources, 100% locally, without sending documents to third parties.
Evidence: recall@1 0.917 · answer judge 1.92/2 · 80 tests · mypy strict · CI with quality gates.
Leia em português.
Contents
Related MCP server: rag-mcp
Highlights
Hybrid search: embeddings (local Qdrant, no Docker) + in-house BM25 (smoothed idf), fused by RRF.
Multi-step agent: if the first search's context is insufficient, the model signals
[MORE_CONTEXT], the agent issues a follow-up search and retries with incremental source memory.Optional re-ranking: cross-encoder (Ollama
/api/rerank, e.g.bge-reranker-v2-m3) with graceful degradation.Resilient fallback: cloud provider (OpenAI-compatible) first, local Ollama as backup when the API is down.
Persistence: chunks live in Qdrant; the BM25 index is restored on demand (first search) with no re-ingestion.
Incremental ingestion: re-running
ingestonly embeds what changed (content-hash idempotent) and prunes orphans — cheap in CI and redeploys.Token optimization: semantic cache (JSONL + cosine, with TTL) returns previously generated answers without re-calling the LLM; a Caveman-style compressor (PT/EN) trims the source context before the prompt.
Traceability: per-step agent
trace+audit.jsonl(question, provider, iterations, latency, sources, cache).Measurable:
recall@k/nDCG@kevaluation pipeline with a quality gate in CI.Two transports: stdio (local RPC) and streamable HTTP (
http://host:port/mcp).
Architecture
flowchart LR
C["MCP Client<br/>stdio or HTTP"] -->|tools: ingest / search / ask / metrics| M["MCP Server<br/>hybrid-rag-mcp"]
M --> AGE["Multi-step agent<br/>loop with [MORE_CONTEXT]"]
M --> I["ingest"]
I --> C1["Chunker<br/>sections + sentences"]
C1 --> E["Embeddings<br/>Ollama bge-m3"]
E --> Q1[("Qdrant<br/>vector search")]
C1 --> K["In-house BM25<br/>lexical search"]
AGE --> RET["Hybrid search"]
RET --> Q1 & K
Q1 & K --> RRF["RRF fusion"]
RRF --> RR["Optional reranker<br/>Ollama /api/rerank"]
RR --> LLM["FallbackLLM<br/>cloud -> Ollama"]
LLM --> AUD["audit.jsonl<br/>trace + iterations + sources"]Metrics (quality gate in CI)
Evaluation pipeline over 36 queries — 12 hand-curated in eval/dataset.jsonl + 24 auto-generated from real documents (eval/dataset.real.jsonl, see Corpus).
CI trigger: fails if recall@1 < 0.8.
k | recall@k | nDCG@k |
1 | 0.917 | 0.917 |
3 | 1.000 | 0.865 |
5 | 1.000 | 0.945 |
Honest numbers over real text: the right source ranks top-1 in 91.7% of cases and always in the top-3. tools/grid_search.py sweeps RRF weights/top_k to reach this result (lexical weight 1.5) — history in eval/grid_results.json. Run locally with python -m hybrid_rag_mcp.eval.
Answer quality (llm-as-judge)
Retrieval proves the right chunk surfaces; this proves the final answer is correct:
12 questions over the corpus (eval/answers.jsonl), each with an expected
answer, scored 0/1/2 by Ollama itself — run with
python -m hybrid_rag_mcp.judge (needs Ollama up).
metric | value |
average | 1.92 (0..2) |
score 2 | 11/12 |
score 1 | 1/12 (ans-06: missed "telemetry doesn't go to PostgreSQL") |
score 0 / no verdict | 0 |
Compression validated: with CONTEXT_COMPRESSION=1 the average holds at 1.92;
with 2, 2.00 (12×2) — pruning function words does not degrade answers
(cache off in all 3 runs; n=12, a 1-item swing is noise —
the conclusion is "doesn't hurt", not "helps").
Scale (retrieval, embedded mode)
Methodology (tools/scale_bench.py): 11 real docs replicated with unique salt
up to ~25k chunks — recall stays on the curated eval; here only latency,
throughput and disk (36 real queries, no LLM). Isolated index, embedded
Qdrant on a 4GB RAM box.
metric | 32,760 indexed chunks |
ingest | 566s (58 chunks/s), 337MB index on disk |
1-thread search | p50 191ms · p95 621ms · mean 373ms |
8-thread search | p50 640ms · p95 4308ms |
p99 single ≈5s | one-time warmup cost (first search restores the BM25 index; varies across runs) |
Re-measured with vectorized BM25 (before: p50 269/p95 663 single, p50 1313/p95 2334 concurrent). p50 fell ~30% single-threaded and ~50% concurrent — numpy releases the GIL on the vectorized path; tails (p95/p99) swing ±2x between runs on this box, noise not signal.
Honest readings (including our own correction): the initial guess was that
pure-Python BM25 was the bottleneck — measured, it wasn't: at 32k docs
score_all cost ~17ms of ~400ms per search (the bulk is Ollama embedding +
local HNSW). We vectorized anyway (~10x: 17.5→1.7ms on 30k synthetic docs,
parity < 1e-9): the win compounds at 10x scale, where the pure loop would
dominate. Contention under concurrency (~5x with 8 threads — GIL + local
Qdrant + embedding on Ollama) stands. Known ceiling: at 128k chunks the box
OOMed — above ~100k chunks or on little RAM, use server mode
(QDRANT_URL, see Deploy).
Context optimization (cache and compression)
Semantic cache — before generating, ask consults data/cache.jsonl in two layers:
exact normalization (identical questions, ignoring case/whitespace) and
similarity — the question is embedded (same bge-m3 as search) and compared by cosine against the entries; above
CACHE_SIM_THRESHOLD(0.92) the saved answer is returned (TTLCACHE_TTL_SEC,CACHE_MAX_ENTRIEScap). A hit skips generation entirely — the biggest token saving. Hits are marked· cachein the answer and recorded in the audit ("cache": true).
Caveman-style context compression — CONTEXT_COMPRESSION (0/1/2) removes
predictable function words (connectives/fillers at level 1; + articles and
auxiliaries at level 2) only from the copy that goes into the LLM prompt:
search, re-ranking and displayed sources keep the original text, and numbers,
proper nouns and negations are never removed. Deterministic, multilingual
(PT/EN), zero dependencies — principle inspired by
Caveman ("strip grammar, keep
facts"), reimplemented in-house.
python tools/optimizers_report.py # dashboard: cache + compression + evaluated integrations
python tools/optimizers_report.py --json # same output as JSONCorpus
examples/corpus/— 11 documents: 5 fictional (operations/security/database/infra/events) + 6 real public-domain books (Project Gutenberg): Chekhov, Machado de Assis, Aluísio Azevedo, Eça de Queirós, Jane Austen and Conan Doyle, mixing PT and EN.tools/fetch_corpus.py— downloads the Gutenberg catalog and auto-generates the evaluation dataset: each query is a real excerpt from a document and the expected document is that excerpt's exact source (self-supervised ground truth, no manual curation).
Running
Prerequisites: Python 3.11+, running Ollama (ollama serve).
# 1. Local models (once)
ollama pull bge-m3 # embeddings
ollama pull qwen3:8b # generation (or another)
ollama pull bge-reranker-v2-m3 # optional: Ollama >= 0.36 only
# 2. Install
python -m venv .venv && source .venv/bin/activate
pip install -e ".[dev]"
# 3. ~1 min demo (ingest + search + ask)
bash examples/demo.sh[ingest] 11 docs, 1964 chunks (+0 novos, 1964 no índice)
[search] 'Qual a porta padrão do servidor?'
- [pqpbr-operations.txt] score=0.065: ## Configuração de Rede O servidor escuta na porta 8443...
[ask] 'De quantas em quantas horas são os backups?'
[ollama/qwen3:8b] Os backups incrementais são executados a cada 4 horas [pqpbr-operations.txt].
fontes: gutenberg-dom-casmurro.txt, pqpbr-operations.txtAs an MCP client
stdio: example client — python examples/client.py "Qual a porta padrão?"
Progress streaming (tokens + steps):
python examples/client_stream.py "Qual a política de manutenção do banco?" — the server
emits step events and response tokens in real time via notifications/progress
(the client sends a progressToken in the request).
HTTP:
# terminal 1
python -m hybrid_rag_mcp --transport http --host 127.0.0.1 --port 8000
# terminal 2
python examples/client_http.py "Qual a porta padrão?"HTTP auth (recommended when exposing on a network): generate a token
(openssl rand -hex 32), export MCP_AUTH_TOKEN on the server and the
client — without the Authorization: Bearer header the server answers 401.
With no token configured, it behaves as before (localhost only).
MCP_AUTH_TOKEN=... python -m hybrid_rag_mcp --transport http --port 8000
MCP_AUTH_TOKEN=... python examples/client_http.py "Qual a porta padrão?"Register in any MCP client (Claude Desktop, editors, agents):
{
"mcpServers": {
"hybrid-rag": {
"command": ".venv/bin/python",
"args": ["-m", "hybrid_rag_mcp"],
"env": { "PYTHONPATH": "src" }
}
}
}Cloud fallback (optional)
Copy .env.example to .env and fill in CLOUD_BASE_URL + CLOUD_API_KEY + CLOUD_MODEL
(any OpenAI-compatible endpoint). Cloud takes priority; Ollama answers automatically
if the API fails or goes offline. For cloud embeddings, also set
CLOUD_EMBED_MODEL — note: switching models changes the dimension and requires
recreating the index (re-run ingest). For a multi-API cascade (e.g. Nvidia →
Gemini → Ollama), use LLM_CHAIN with the list in priority order — the first
to answer wins, you pay only it per question.
Deploy
systemd (bare metal): deploy/hybrid-rag.service is a user unit that brings
up the HTTP transport on 127.0.0.1:8000 with automatic restart — adjust the paths
(%h = your home) and install with
systemctl --user enable --now pointing at the file. .env is optional
(cloud/token only if you export them).
Docker (Qdrant + RAG): docker compose up -d --build brings up Qdrant and the
server already pointed at it (QDRANT_URL=http://qdrant:6333, Ollama via
host.docker.internal). No Docker installed here so the compose was not run
locally — validated by inspection; CI keeps covering embedded mode.
Data hygiene: audit.jsonl rotates by size
(AUDIT_MAX_BYTES, default 5MB, keeps AUDIT_KEEP=3 backups); the semantic
cache is already capped (CACHE_MAX_ENTRIES).
Observability
In-process metrics, zero dependencies (src/hybrid_rag_mcp/metrics.py):
counters (search_total, ask_total, ask_cache_hits, *_errors,
ask_provider_<name>) + latencies (search_ms, ask_ms, ingest_ms).
HTTP:
GET /metricsin Prometheus format (inherits the bearer auth). Scrape example:scrape_configs: - job_name: hybrid-rag static_configs: [{targets: ["127.0.0.1:8000"]}] # authorization: {credentials: <MCP_AUTH_TOKEN>} # if auth is onBoth transports:
metricsMCP tool with a human summary.Logs: one JSON line per
search/askon stderr ({"op": "ask", "elapsed_ms": 123.4, "cache_hit": false, ...}) — stdout belongs to the MCP protocol in stdio mode and is never polluted.
MCP tools
Tool | Description |
| Indexes |
| Hybrid search (RRF, optional re-ranking) returning snippets + sources. |
| Multi-step agent: retrieves, generates, detects insufficient context, re-searches and answers citing sources (with audit log). Consults the semantic cache before generating. |
| Server metrics summary (counters + p50/p95 latencies since boot). |
Extending
Nothing here is mandatory — every piece has a local default and can be swapped without forking:
Compressor: any
(str) -> strcallable viaCONTEXT_COMPRESSOR=my_package:clean(beats the built-in 0/1/2 levels). E.g. a sidecar like Headroom wrapped in a function.Embeddings/LLM:
EmbeddingProvider/LLMProviderABCs (src/hybrid_rag_mcp/providers/); cloud viaCLOUD_*with no code changes. Switching embeddings changes the dimension: wipe the index and re-runingest, or setQDRANT_RECREATE_ON_DIM_CHANGE=1.Reranker:
score(query, texts) -> list[float]duck-type (seeproviders/rerank.py); empty = off, with graceful degradation.
Structure
src/hybrid_rag_mcp/
├── server.py # MCP server (stdio + streamable HTTP, optional bearer auth)
├── config.py # .env configuration (pydantic-settings)
├── eval.py # recall@k / nDCG@k evaluation
├── judge.py # Answer evaluation via llm-as-judge (score 0/1/2)
├── bench.py # Scale-bench blocks (percentiles, synthetic corpus)
├── metrics.py # In-process metrics (counters + p50/p95, Prometheus)
├── rag/
│ ├── agent.py # Multi-step loop (source memory, [MORE_CONTEXT])
│ ├── chunker.py # Chunking by markdown sections + sentences
│ ├── engine.py # Orchestration: ingest → search → ask + audit
│ ├── hybrid.py # RRF fusion + re-ranking
│ └── ingestion.py # md/txt/pdf reading
├── providers/
│ ├── embed.py # Embeddings via Ollama
│ ├── llm.py # FallbackLLM (cloud → Ollama)
│ └── rerank.py # Optional cross-encoder (graceful degradation)
├── optimize/
│ ├── cache.py # Semantic cache (JSONL + cosine + TTL)
│ └── compress.py # Caveman-style compressor (PT/EN)
└── stores/
├── vector.py # Embedded Qdrant (no Docker, persistent)
└── lexic.py # BM25 with smoothed idfArchitecture decisions with context and evidence: docs/adr/
(RRF, storage, cache, judge, lazy-init, external integrations, cascade scope,
BM25 tradeoff, open seams). Contributing guide: CONTRIBUTING.md.
Release history: CHANGELOG.md.
Quality
98 unit tests (
pytest) with no network/Ollama — chunking, RRF, BM25 (incl. vectorized × brute-force parity), persistence, eval metrics, agent loop, semantic cache, compressor, lexical-index thread-safety, judge parsing/aggregation, HTTP auth, audit rotation, observability, pluggable compressor, cloud embeddings, dimension check and LLM cascade.CI in 2 jobs:
test(ruff + mypy strict + pytest with ≥75% coverage +pip-audit+ stdio/HTTP smoke) andeval(real Ollama +recall@1 >= 0.8gate). Weekly Dependabot (pip + actions).Two storage modes: embedded (default, no Docker, 1 process at a time) or server (
docker compose up -d+QDRANT_URL=http://localhost:6333) for simultaneous sessions — seedocker-compose.yml.
Available Tools
4 toolsaskC
Responde com RAG e faz streaming do progresso/tokens via progress notifications.
| Name | Required | Description | Default |
|---|---|---|---|
| top_k | No | ||
| question | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does disclose one valuable behavior: streaming progress/tokens via progress notifications. However, it does not state whether the operation is read-only, how long it may take, or any limitations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no wasted words: it states the core RAG answering behavior first, then the streaming detail. It is appropriately concise, though a bit more structure could improve clarity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations, no schema descriptions, and sibling tools that could overlap, this description is too thin. It omits parameter semantics, usage context, and any distinction from search; the output schema helps with return values but not with selection or invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description never mentions question or top_k. The agent must infer their meaning entirely from the schema's bare property names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the action (responds) and the method (RAG), and the streaming mention adds a distinguishing behavioral trait. It stops short of a 5 because it does not explicitly contrast with sibling tools like search.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance is provided, and the sibling tools are never referenced. The verb 'ask' offers only an implied use case, which is not enough for an agent to reliably choose this over search.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ingestA
Indexa documentos (md/txt/pdf) de um diretório nos índices vetorial e BM25.
| Name | Required | Description | Default |
|---|---|---|---|
| corpus_dir | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses the main write behavior (indexing into vector and BM25 stores) but does not mention side effects, such as whether existing documents are replaced, whether indexing is incremental, if the directory is read recursively, or any permission requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence with no wasted words. It front-loads the action and packs in file types and target indexes efficiently.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple tool with one parameter and an output schema, the description covers the core purpose and parameter meaning. Gaps remain around usage timing, reindexing behavior, and directory handling, but these are not severe enough to render it unusable.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. 'de um diretório' clarifies that corpus_dir is the source directory, adding meaning to the parameter name alone. However, it does not explain the default behavior, path format, or how recursive/required the directory is.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Indexa' - indexes), a clear resource (documents md/txt/pdf from a directory), and the target (vector and BM25 indexes). This clearly differentiates it from siblings 'search' and 'ask', which are query operations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The intended use is implied: call this tool to index documents before searching or asking. However, there is no explicit guidance on when to use it versus alternatives, prerequisites, or scenarios where it should not be used.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
metricsA
Resumo das métricas do servidor (contadores + latências p50/p95 desde o boot).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does add useful behavioral context: it returns a summary of counters and p50/p95 latencies since boot, implying a read-only metrics retrieval. It does not explicitly state that there are no side effects or describe output shape, but the low-risk nature of the tool makes this acceptable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact sentence front-loads the purpose and packs in the metric categories and time window without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a parameterless tool with an output schema, this description provides all essential invocation knowledge: what is returned and its time scope. Nothing needed to call the tool correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so there is no parameter meaning for the description to add beyond the schema. The no-parameter baseline of 4 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the tool as a server-metrics summary and specifies the included data (counters and p50/p95 latencies) and the since-boot window, which is enough to separate it from ask/ingest/search. It lacks an explicit action verb such as 'returns' or 'gets', so it is slightly below the top anchor.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given for when to use this tool or when a sibling tool would be preferable. The 'since boot' detail describes the metric window, not a selection rule, leaving the agent to infer usage from the tool name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
searchA
Busca híbrida (vetorial + BM25 via RRF) no corpus indexado. Retorna trechos e fontes.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | ||
| top_k | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It goes beyond a bare 'search' by revealing the hybrid retrieval approach (vector + BM25 via RRF) and specifying that the output contains excerpts and sources. This gives an agent useful expectations about both behavior and returns.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense sentence with no filler. It front-loads the core purpose and retrieval method, then immediately gives the return type. Every word contributes meaningful information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter search tool with an output schema present, the description is largely complete: it names the corpus, explains the hybrid mechanism, and states what is returned. It could add usage direction relative to siblings, but that gap is already covered under usage guidelines.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description needed to compensate, but it does not explain query or top_k. The parameter names and types are self-explanatory enough for a minimal viable call, and the default for top_k is in the schema, but the agent gets no additional semantic guidance from the description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific operation ('hybrid search') and a clear resource ('the indexed corpus'), and further distinguishes the tool by naming the retrieval mechanism (vector + BM25 via RRF) and the return type (excerpts and sources). This sets it apart from the sibling tools ingest and ask.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The intended use is implied: search over an indexed corpus. However, there is no explicit guidance about when to prefer this tool over ask or ingest, nor any mention of what each sibling is better suited for. The context is reasonably clear but the routing decision is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
- Added
metrics
3 tool updates
v0.1.0- First observed
ask - First observed
ingest - First observed
search
TDQS
Scored across 4 tools
As ferramentas têm propósitos majoritariamente distintos: ingest adiciona documentos, search faz consulta híbrida crua, ask gera resposta com RAG e metrics expõe estatísticas. A única ambiguidade possível é entre ask e search, já que ambos envolvem recuperação, mas as descrições deixam clara a diferença entre resposta gerada e trechos brutos.
Os nomes são todos curtos e em minúsculas, seguindo um estilo simples e consistente. Há uma pequena mistura entre verbos (ask, ingest, search) e substantivo (metrics), mas o padrão geral é previsível e uniforme.
Com 4 ferramentas, o servidor é enxuto e cada uma cobre uma função essencial do ciclo RAG: ingestão, busca, geração e observabilidade. O número é apropriado para o escopo declarado, sem excesso ou carência evidente.
O fluxo principal de RAG está coberto: ingestão, busca híbrida, perguntas com streaming e métricas. Faltam operações secundárias como remover/atualizar documentos ou listar o corpus, mas essas são lacunas menores que não impedem o uso básico do servidor.
Maintenance
Related MCP Connectors
- KumbukaOAuthai.kumbuka
Governed, auditable knowledge your team curates for its AI assistants, self-hostable
Ingest, manage, and retrieve documents for RAG-powered AI applications
Search your knowledge bases from any AI assistant using hybrid RAG.
Versioned documentation registry and semantic search for AI tools and coding assistants.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceEnables intelligent ingestion and querying of PDF, Markdown, and text files using hybrid search that combines keyword matching and semantic embeddings with citations.2-
- FlicenseAqualityDmaintenanceEnables indexing local documents (PDF, Markdown, text, code) into a knowledge base and querying them via semantic search using local embeddings, all running privately on your machine.4-
- AlicenseAqualityCmaintenanceEnables semantic search and question answering over a knowledge base using hybrid retrieval and grounded answers, all running offline with no API keys.4MIT
- AlicenseAqualityAmaintenanceIndexes your project's markdown documentation and exposes it to AI agents via local hybrid search (lexical + semantic) with progressive disclosure tools.3463 npm5MIT