Documind MCP Server
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Documind MCP ServerWhat are the main conclusions of the latest report?"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
DocuMind — Intelligent Document Q&A System
DocuMind is a production-style Retrieval-Augmented Generation (RAG) application for asking natural-language questions over PDF documents. It parses PDFs with PyMuPDF, creates overlapping semantic chunks, embeds them with sentence-transformers/all-MiniLM-L6-v2, stores vectors in ChromaDB, retrieves the most relevant passages, and generates grounded answers with AWS Bedrock.
The project also exposes the retrieval layer through an MCP server so MCP-compatible clients can search and answer from the same document store without going through the REST API.
Architecture
PDF upload
|
v
PyMuPDF extraction
|
v
Overlapping text chunks
|
v
MiniLM embeddings
|
v
ChromaDB vector store
|
+---------------------> MCP tools
|
User question
|
v
Semantic top-k retrieval
|
v
AWS Bedrock LLM
|
v
Grounded answer + page citationsRelated MCP server: DocAgent-MCP
Tech stack
Python, FastAPI
React + Vite
PyMuPDF
Hugging Face Sentence Transformers (
all-MiniLM-L6-v2)ChromaDB
AWS Bedrock
MCP (Model Context Protocol)
Docker / Docker Compose
Repository layout
backend/ FastAPI ingestion and Q&A API
frontend/ React interface
mcp_server/ MCP tools backed by the same vector store
tests/ Basic PDF parsing test
data/chroma/ Local persistent ChromaDB data (ignored by Git)Local setup
1. Prerequisites
Python 3.11+
Node.js 20+
AWS credentials configured locally
Access to the Bedrock model specified in
.env
2. Environment
cp .env.example .envConfigure AWS_REGION and BEDROCK_MODEL_ID as needed. AWS credentials are intentionally not stored in the repository.
3. Run the API
python -m venv .venv
source .venv/bin/activate # Windows: .venv\\Scripts\\activate
pip install -r backend/requirements.txt
uvicorn app.main:app --app-dir backend --reloadAPI docs: http://localhost:8000/docs
4. Run the frontend
cd frontend
npm install
npm run devOpen http://localhost:5173.
5. Run the MCP server
pip install -r mcp_server/requirements.txt
python mcp_server/server.pyAvailable MCP tools:
search_documents(query, top_k=4)answer_from_documents(question, top_k=4)
API examples
Upload a PDF:
curl -X POST http://localhost:8000/documents \
-F "file=@example.pdf"Ask a question:
curl -X POST http://localhost:8000/query \
-H "Content-Type: application/json" \
-d '{"question":"What are the main conclusions?"}'Docker
cp .env.example .env
docker compose up --buildThe API container mounts your local
~/.awsdirectory read-only for development. For cloud deployment, use an IAM role instead of static credentials.
Resume / portfolio highlights
RAG pipeline for natural-language PDF querying
PyMuPDF document parsing and overlapping chunking
Hugging Face MiniLM embeddings + ChromaDB semantic retrieval
Citation-aware answer generation with AWS Bedrock
FastAPI service and React chat UI
MCP tools for reusable document search and answer generation
Security notes
Never commit
.env, AWS keys, uploaded PDFs, or persisted Chroma data.Use IAM roles with least-privilege Bedrock permissions for deployments.
Add authentication, tenant isolation, malware scanning, and object storage before using this as a public multi-user service.
License
MIT
QLoRA fine-tuning (optional)
The repository includes a separate training entry point matching the project's domain-adaptation workflow. Training is kept outside the API dependencies so the normal application remains lightweight.
pip install -r scripts/requirements-training.txt
python scripts/fine_tune_qlora.py --data training.jsonlTraining data uses JSONL records containing instruction, context, and answer fields. Run QLoRA on a compatible CUDA GPU; do not attempt 4-bit training on a CPU-only production API instance.
Load testing
A Locust scenario is included for reproducing concurrent query tests:
pip install locust
locust -f locustfile.py --host http://localhost:8000Use the Locust UI to run 15–20 concurrent users and record latency/throughput for your deployment. Performance varies by Bedrock model, region, retrieval corpus, and infrastructure, so benchmark numbers should be reported from an actual run rather than assumed from source code alone.
This server cannot be deployed
Maintenance
Related MCP Connectors
- docs2mcpOAuthcom.docs2mcp
Query your own PDFs and documents from any MCP client. Every answer cites the page it came from.
Remote ChromaDB vector database MCP server with streamable HTTP transport
Read-only hosted MCP over CanonicAI's cited Answers corpus on canonicai.com.
Agentic search over your Dewey document collections from any MCP-compatible client.
Related MCP Servers
- AlicenseAqualityDmaintenanceEnables MCP clients to list indexed PDF document collections and perform semantic search queries on them using locally extracted text and embeddings.2AGPL 3.0
- FlicenseNot gradedqualityCmaintenanceEnables local document question-answering and retrieval via MCP, supporting multi-turn conversation, intent recognition, and tools for document search, Q&A, and summarization.5-
- FlicenseNot gradedqualityDmaintenanceEnables natural language search over PDF documents using vector search, allowing MCP clients like Claude Desktop to query PDF content.1-
- FlicenseNot gradedqualityCmaintenanceProvides read-only, citation-backed semantic search and retrieval-augmented generation over enterprise documents via standardized MCP tools, with local embeddings for privacy.-