Skip to main content
Glama
KimTsegzc

xinPlugin_Chroma_fastMCP

by KimTsegzc

xinPlugin_Chroma_fastMCP

The vector retrieval layer for the MinIO knowledge base: it extracts text from documents uploaded to MinIO (PDF/txt/md), chunks it, and stores the embeddings in Chroma. Through FastMCP it exposes "retrieval" as MCP tools, letting DSH (DeepSeek Harness)'s dsh-mcp-client connect to them so the agent can retrieve the original text directly during Q&A.

Composition

File

Purpose

chroma_store.py

Chroma persistence + chunking (with page/line numbers) + hybrid retrieval (semantic + character-bigram BM25, RRF-fused)

ingest.py

CLI ingestion: python ingest.py <文件> [source名], outputs a JSON summary (for MinIO-plugin wiring)

server.py

FastMCP stdio service exposing search / ingest_file / list_sources

requirements.txt

chromadb / fastmcp / pypdf

Related MCP server: Modular RAG MCP Server

Installation

pip install -r requirements.txt
# 首次检索会下载默认 embedding(all-MiniLM-L6-v2,约 80MB,缓存在 ~/.cache/chroma)

Usage

# 入库
python ingest.py "广州十五五规划.pdf" "广州十五五规划.pdf"

# 检索(或经 MCP 工具 search)
python -c "from chroma_store import search; import json; print(json.dumps(search('广州 人工智能+ 大模型 算力 数据要素', 6), ensure_ascii=False))"

MCP tools

  • search(query, top_k=6): hybrid semantic + keyword retrieval; returns the original snippet and its provenance (file + page + line number).

  • ingest_file(path, source_name): ingests a local file into the store.

  • list_sources(): lists the ingested sources.

On the DSH side, dsh-mcp-client connects to server.py over stdio, and the tool names look like mcp__chroma__search.

How retrieval works

  • Chunking: text is extracted page by page; header/footer/page-number noise is filtered out; every 6 lines become a chunk (1-line overlap), with source/page/line_start/line_end recorded in the metadata.

  • Hybrid retrieval: Chroma semantic vectors (cosine) + character-bigram BM25 sparse retrieval, fused by RRF — when Chinese semantic embeddings are weak, BM25 keeps exact keywords such as "large models/compute/data elements" covered, so provenance stays reliable.

  • Ingesting invalidates the sparse cache immediately; re-ingesting the same source overwrites the entry.

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables AI agents to search, deep-read, and build knowledge bases from Markdown, PDF, DOCX, and PPTX documents via MCP tools for retrieval, document navigation, and ingestion.
    36 npm
    635
    MIT
  • F
    license
    Not graded
    quality
    D
    maintenance
    MCP server that indexes a knowledge base into Chroma and provides search tools for retrieving document fragments via vector embeddings.
    -