Documind MCP Server
README.md
# DocuMind — Intelligent Document Q&A System
DocuMind is a production-style Retrieval-Augmented Generation (RAG) application for asking natural-language questions over PDF documents. It parses PDFs with PyMuPDF, creates overlapping semantic chunks, embeds them with `sentence-transformers/all-MiniLM-L6-v2`, stores vectors in ChromaDB, retrieves the most relevant passages, and generates grounded answers with AWS Bedrock.
The project also exposes the retrieval layer through an MCP server so MCP-compatible clients can search and answer from the same document store without going through the REST API.
## Architecture
```text
PDF upload
|
v
PyMuPDF extraction
|
v
Overlapping text chunks
|
v
MiniLM embeddings
|
v
ChromaDB vector store
|
+---------------------> MCP tools
|
User question
|
v
Semantic top-k retrieval
|
v
AWS Bedrock LLM
|
v
Grounded answer + page citations
```
## Tech stack
- Python, FastAPI
- React + Vite
- PyMuPDF
- Hugging Face Sentence Transformers (`all-MiniLM-L6-v2`)
- ChromaDB
- AWS Bedrock
- MCP (Model Context Protocol)
- Docker / Docker Compose
## Repository layout
```text
backend/ FastAPI ingestion and Q&A API
frontend/ React interface
mcp_server/ MCP tools backed by the same vector store
tests/ Basic PDF parsing test
data/chroma/ Local persistent ChromaDB data (ignored by Git)
```
## Local setup
### 1. Prerequisites
- Python 3.11+
- Node.js 20+
- AWS credentials configured locally
- Access to the Bedrock model specified in `.env`
### 2. Environment
```bash
cp .env.example .env
```
Configure `AWS_REGION` and `BEDROCK_MODEL_ID` as needed. AWS credentials are intentionally not stored in the repository.
### 3. Run the API
```bash
python -m venv .venv
source .venv/bin/activate # Windows: .venv\\Scripts\\activate
pip install -r backend/requirements.txt
uvicorn app.main:app --app-dir backend --reload
```
API docs: `http://localhost:8000/docs`
### 4. Run the frontend
```bash
cd frontend
npm install
npm run dev
```
Open `http://localhost:5173`.
### 5. Run the MCP server
```bash
pip install -r mcp_server/requirements.txt
python mcp_server/server.py
```
Available MCP tools:
- `search_documents(query, top_k=4)`
- `answer_from_documents(question, top_k=4)`
## API examples
Upload a PDF:
```bash
curl -X POST http://localhost:8000/documents \
-F "file=@example.pdf"
```
Ask a question:
```bash
curl -X POST http://localhost:8000/query \
-H "Content-Type: application/json" \
-d '{"question":"What are the main conclusions?"}'
```
## Docker
```bash
cp .env.example .env
docker compose up --build
```
> The API container mounts your local `~/.aws` directory read-only for development. For cloud deployment, use an IAM role instead of static credentials.
## Resume / portfolio highlights
- RAG pipeline for natural-language PDF querying
- PyMuPDF document parsing and overlapping chunking
- Hugging Face MiniLM embeddings + ChromaDB semantic retrieval
- Citation-aware answer generation with AWS Bedrock
- FastAPI service and React chat UI
- MCP tools for reusable document search and answer generation
## Security notes
- Never commit `.env`, AWS keys, uploaded PDFs, or persisted Chroma data.
- Use IAM roles with least-privilege Bedrock permissions for deployments.
- Add authentication, tenant isolation, malware scanning, and object storage before using this as a public multi-user service.
## License
MIT
## QLoRA fine-tuning (optional)
The repository includes a separate training entry point matching the project's domain-adaptation workflow. Training is kept outside the API dependencies so the normal application remains lightweight.
```bash
pip install -r scripts/requirements-training.txt
python scripts/fine_tune_qlora.py --data training.jsonl
```
Training data uses JSONL records containing `instruction`, `context`, and `answer` fields. Run QLoRA on a compatible CUDA GPU; do not attempt 4-bit training on a CPU-only production API instance.
## Load testing
A Locust scenario is included for reproducing concurrent query tests:
```bash
pip install locust
locust -f locustfile.py --host http://localhost:8000
```
Use the Locust UI to run 15–20 concurrent users and record latency/throughput for your deployment. Performance varies by Bedrock model, region, retrieval corpus, and infrastructure, so benchmark numbers should be reported from an actual run rather than assumed from source code alone.
This server cannot be deployed
Maintenance
ActivityMaintained
ResponsivenessNo issues