bdc-doc-mcp
BDC Doc RAG
bdc-assist 的文档 RAG MCP
bdc_doc_mcp/config.py env-driven embeddings/LLM/Chroma (replaces utils/__init__.set_emb_llm)
bdc_doc_mcp/ingest.py .pkl/.md/.mdx/.txt/.pdf → embeddings → Chroma (replaces utils/chroma/utils.py)
bdc_doc_mcp/api.py FastAPI: /health /search
bdc_doc_mcp/mcp_server.py search_docs MCP tool for AI agents — self-contained, same search as the API
bdc_doc_mcp/preproc/ source-specific preprocessing pipeline
tests/ self-checks + API / agent notebooks
data/ preproc output (*.pkl), ingest input设置
uv sync
cp .env.example .env # then fill in keys/URLs源仓库
仅预处理(--sources all)时需要;API/MCP 服务器和摄取现有 .pkl 文件无需它们。克隆到本仓库旁边(或将环境变量指向它们):
git clone https://github.com/stagecc/interim-bdc-website ../interim-bdc-website # BDC_WEBSITE_DIR
git clone https://github.com/stagecc/bdc-gitbook ../bdc-gitbook # BDC_GITBOOK_DIR模型
补全使用 Azure 上的 OpenAI API(默认为 gpt-4o-mini)
嵌入使用 Sterling 上的 Ollama(通过 RENCI VPN 连接)
kubectl -n ner port-forward svc/ollama 11434:11434或使用本地 Ollama 搭配 groonga/bge-m3-Q4_K_M-GGUF 模型。
Related MCP server: okfy
摄取
从每个源完整重建(需要克隆两个源仓库 — 见设置;写入 data/*.pkl,然后加载它们):
uv run python -m bdc_doc_mcp.preproc.pipeline --sources all --ingest --reset单个文件或目录:
uv run python -m bdc_doc_mcp.ingest ./data/docs.pkl --doc-type docs # BDC_Chatbot preproc .pkl
uv run python -m bdc_doc_mcp.ingest ../interim-bdc-website/src/pages --doc-type page --reset嵌入模型在集合内不可互换 — bge-m3 是 1024 维,text-embedding-3-small 是 1536 维。切换模型意味着 --reset 并完全重新摄取。
API
uv run uvicorn bdc_doc_mcp.api:app --port 8000 # docs at /docs端点 | 请求体 | 返回 |
| — |
|
|
| 排序后的分块 + 元数据 + 分数 |
mode 是 embedding(默认;语义相似度,分数 = 距离,越低越好)或 keyword(模糊字面词匹配 — 忽略大小写/标点并容忍小拼写错误,因此 picsure 能找到 "PIC-SURE";分数 = 出现次数,越高越好 — 用于精确名称/缩写)。
doc_type 是要搜索的类型 CSV(例如 page,faq)。省略时,仅搜索 docs、page、faq 和 video — 显式指定 fellow、update 或 event 才能搜索它们。
date_from/date_to(YYYY-MM-DD,含端点)按日期过滤;只有 event 和 update 文档带有日期,因此日期过滤器会隐式缩小到这些类型。
该服务在设计上仅用于搜索;摄取通过 CLI 离线进行(见摄取),回答是调用方的职责 — 代理自带 LLM。
MCP
uv run python -m bdc_doc_mcp.mcp_server # stdio
uv run python -m bdc_doc_mcp.mcp_server --http # streamable HTTP, port MCP_PORT (default 8001)暴露一个工具 search_docs — 与 API 相同的搜索,但直接查询 Chroma,因此无需运行 API 服务。需要已摄取的 .chroma_db + 嵌入。
Stdio 客户端(Claude Desktop/Code、Cursor)自行启动服务器 — 注册它:
{"mcpServers": {"bdc-doc-mcp": {
"command": "uv",
"args": ["--directory", "/path/to/bdc-doc-mcp", "run", "python", "-m", "bdc_doc_mcp.mcp_server"]
}}}网络客户端:改为运行 --http 并将它们指向 http://host:8001/mcp。
冒烟测试:uv run python tests/test_mcp.py
预处理
bdc_doc_mcp/preproc/ 是 BDC_Chatbot 流水线的移植:
模块 | 来源 | 移植自 (BDC_Chatbot) | 说明 |
| interim-bdc-website MDX |
| fellows、events、latest-updates、pages |
| bdc-gitbook markdown |
| 按标题层级分块;需要克隆仓库 |
| bdcatalyst.freshdesk.com |
| 实时抓取 |
| Google Sheet + Drive SRT |
| 带时间戳 URL 的视频转录 |
| — | — | LLM 分块上下文化 + 摘要器 |
| — |
| 编排器 |
--no-contextualize 跳过每个分块的 LLM 调用(更快,检索效果较弱)。
源路径来自 BDC_WEBSITE_DIR / BDC_GITBOOK_DIR。
测试
uv run python tests/test_ingest.py # batching + chunk-id logic, no network
uv run python tests/test_keyword.py # keyword ranking, pure function, no DB or API
uv run python tests/test_mcp.py # starts the server over stdio and exercises its tools; needs .chroma_db + embeddings笔记本(每个在空闲端口上启动 API 并在结束时关闭;两者都需要已摄取的 .chroma_db):
tests/api_test.ipynb— 纯 API 演练:/health、/search、doc_type过滤器。只需要本地嵌入。tests/agent_test.ipynb— 一个工具调用代理(deepagents):配置的 LLM 将search_docs作为 LangChain 工具,并决定何时调用它。还需要补全提供方可访问。
Available Tools
1 toolsearch_docsA
Search the BDC (NHLBI BioData Catalyst) documentation database.
Returns the top-k matching chunks with content, metadata (source, doc_type, datetime when available), and a score.
query is the search text. In embedding mode phrase it as a question or topic (e.g. "how do I bring my own data"); in keyword mode give the literal terms to match.
k is the number of chunks to return (default 5). Raise it (10-20) for broad or multi-part questions; each chunk is a small section of a document.
mode toggles the search engine:
"embedding" (default): semantic similarity — best for questions, topics, and paraphrased wording. score is a distance (lower = more similar).
"keyword": fuzzy literal word matching — ignores case and punctuation ("picsure" finds "PIC-SURE") and tolerates small typos — best for exact names, acronyms, tool names, or error messages the embedding may blur. Chunks matching more of the query terms rank first; score is the total number of occurrences (higher = better).
doc_type is a CSV string of types to search (e.g. "page,faq" or "video"). Available types:
docs: BDC GitBook platform documentation — user guides, how-tos, and technical reference (bdcatalyst.gitbook.io)
page: key pages of the BDC website — about/overview, joining BDC, analyzing & sharing data, usage costs and terms
faq: Freshdesk help-desk FAQ articles (support questions & answers)
video: transcripts of BDC YouTube tutorials/webinars, with timestamped links into the video
fellow: BDC Fellows profiles — fellowship recipients and their research projects
update: dated news posts ("latest updates") from the BDC website
event: dated BDC events — webinars, workshops, deadlines When doc_type is omitted, only docs, page, faq, and video are searched — name fellow, update, or event explicitly to search them.
date_from / date_to ("YYYY-MM-DD", inclusive) filter by date. Only event and update docs carry a date, so a date filter implicitly narrows to those types. Results are ranked by relevance, NOT date — for "recent"/"latest" questions, always set date_from to bound the range, then compare the dates returned.
| Name | Required | Description | Default |
|---|---|---|---|
| k | No | ||
| mode | No | embedding | |
| query | Yes | ||
| date_to | No | ||
| doc_type | No | ||
| date_from | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description fully discloses behavioral traits: default doc types when omitted, score interpretation (distance vs occurrences), the effect of date filters, and ranking by relevance not date. It also notes that only event and update docs carry dates, further clarifying behavior. No annotation contradiction exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with bullet points for modes and types, and clear paragraphs for date and ranking behavior. It is lengthy but every sentence carries essential information, and it is front-loaded with the purpose and return content. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (6 parameters, no output schema, no annotations), the description is complete. It explains return format, scoring meaning, type-specific behavior, and parameter interactions. It fully equips an agent to invoke the tool correctly for a variety of use cases.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it does thoroughly. It explains query phrasing for each mode, k's range and purpose, mode options with detailed semantics, doc_type as a CSV list with each type's meaning, and date_from/date_to format and inclusive behavior. This adds far more meaning than the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it searches the BDC (NHLBI BioData Catalyst) documentation database and returns top-k matching chunks with content, metadata, and a score. It names the specific resource and what is returned, making the tool's purpose unambiguous even without sibling tools for differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit guidance on when to use embedding vs keyword mode, how to adjust k for broad questions, when to explicitly name doc_type values, and how to use date filters for recency queries. It also warns that date filters implicitly narrow to types with dates and explains ranking behavior, giving clear when-to-use and when-not-to-use instruction.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
v0.1.0- First observed
search_docs
TDQS
Scored across 1 tool
With only a single tool, there is no possibility of ambiguity. The tool has a clearly defined purpose for searching documentation.
The tool name 'search_docs' follows a consistent verb_noun pattern and is descriptive. Since it is the only tool, naming is inherently consistent.
The server exposes only one tool, which is extremely thin. Even though the tool is multi-functional, a single tool does not constitute a well-scoped set; most servers with this purpose would benefit from at least a couple of complementary tools (e.g., retrieving a document by ID or listing available types).
The search tool covers multiple documentation sources and provides filtering and multiple modes, which addresses the core purpose. However, it lacks any other operation such as fetching a specific document, listing available doc types, or managing content, leaving notable gaps for a documentation server.
Maintenance
Related MCP Connectors
Team docs served to AI agents over MCP - search, Markdown reads, version pinning, read audit.
Agentic search over your Dewey document collections from any MCP-compatible client.
Agent-driven search: build, import, tune, search, and score result quality — all over MCP.
Make your knowledge agent-ready. One MCP endpoint, 5 connectors, 3 search modes.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceProvides semantic search over markdown documentation using RAG, allowing natural language queries and integration with MCP clients.1MIT
- AlicenseNot gradedqualityBmaintenanceEnables AI agents to search, read, and traverse documentation bundles in Open Knowledge Format via MCP tools.506 npm73MIT
- AlicenseNot gradedqualityDmaintenanceProvides RAG (Retrieval Augmented Generation) access to technical documentation through MCP, enabling LLMs to search and retrieve relevant documentation on-demand.4MIT
- AlicenseNot gradedqualityAmaintenanceCrawl documentation sites, index them with hybrid search, and expose them as MCP tools so LLM agents can search and retrieve current docs.MIT