Skip to main content
Glama

DocuMind — Intelligent Document Q&A System

DocuMind is a production-style Retrieval-Augmented Generation (RAG) application for asking natural-language questions over PDF documents. It parses PDFs with PyMuPDF, creates overlapping semantic chunks, embeds them with sentence-transformers/all-MiniLM-L6-v2, stores vectors in ChromaDB, retrieves the most relevant passages, and generates grounded answers with AWS Bedrock.

The project also exposes the retrieval layer through an MCP server so MCP-compatible clients can search and answer from the same document store without going through the REST API.

Architecture

PDF upload
   |
   v
PyMuPDF extraction
   |
   v
Overlapping text chunks
   |
   v
MiniLM embeddings
   |
   v
ChromaDB vector store
   |
   +---------------------> MCP tools
   |
User question
   |
   v
Semantic top-k retrieval
   |
   v
AWS Bedrock LLM
   |
   v
Grounded answer + page citations

Related MCP server: DocAgent-MCP

Tech stack

  • Python, FastAPI

  • React + Vite

  • PyMuPDF

  • Hugging Face Sentence Transformers (all-MiniLM-L6-v2)

  • ChromaDB

  • AWS Bedrock

  • MCP (Model Context Protocol)

  • Docker / Docker Compose

Repository layout

backend/       FastAPI ingestion and Q&A API
frontend/      React interface
mcp_server/    MCP tools backed by the same vector store
tests/         Basic PDF parsing test
data/chroma/   Local persistent ChromaDB data (ignored by Git)

Local setup

1. Prerequisites

  • Python 3.11+

  • Node.js 20+

  • AWS credentials configured locally

  • Access to the Bedrock model specified in .env

2. Environment

cp .env.example .env

Configure AWS_REGION and BEDROCK_MODEL_ID as needed. AWS credentials are intentionally not stored in the repository.

3. Run the API

python -m venv .venv
source .venv/bin/activate   # Windows: .venv\\Scripts\\activate
pip install -r backend/requirements.txt
uvicorn app.main:app --app-dir backend --reload

API docs: http://localhost:8000/docs

4. Run the frontend

cd frontend
npm install
npm run dev

Open http://localhost:5173.

5. Run the MCP server

pip install -r mcp_server/requirements.txt
python mcp_server/server.py

Available MCP tools:

  • search_documents(query, top_k=4)

  • answer_from_documents(question, top_k=4)

API examples

Upload a PDF:

curl -X POST http://localhost:8000/documents \
  -F "file=@example.pdf"

Ask a question:

curl -X POST http://localhost:8000/query \
  -H "Content-Type: application/json" \
  -d '{"question":"What are the main conclusions?"}'

Docker

cp .env.example .env
docker compose up --build

The API container mounts your local ~/.aws directory read-only for development. For cloud deployment, use an IAM role instead of static credentials.

Resume / portfolio highlights

  • RAG pipeline for natural-language PDF querying

  • PyMuPDF document parsing and overlapping chunking

  • Hugging Face MiniLM embeddings + ChromaDB semantic retrieval

  • Citation-aware answer generation with AWS Bedrock

  • FastAPI service and React chat UI

  • MCP tools for reusable document search and answer generation

Security notes

  • Never commit .env, AWS keys, uploaded PDFs, or persisted Chroma data.

  • Use IAM roles with least-privilege Bedrock permissions for deployments.

  • Add authentication, tenant isolation, malware scanning, and object storage before using this as a public multi-user service.

License

MIT

QLoRA fine-tuning (optional)

The repository includes a separate training entry point matching the project's domain-adaptation workflow. Training is kept outside the API dependencies so the normal application remains lightweight.

pip install -r scripts/requirements-training.txt
python scripts/fine_tune_qlora.py --data training.jsonl

Training data uses JSONL records containing instruction, context, and answer fields. Run QLoRA on a compatible CUDA GPU; do not attempt 4-bit training on a CPU-only production API instance.

Load testing

A Locust scenario is included for reproducing concurrent query tests:

pip install locust
locust -f locustfile.py --host http://localhost:8000

Use the Locust UI to run 15–20 concurrent users and record latency/throughput for your deployment. Performance varies by Bedrock model, region, retrieval corpus, and infrastructure, so benchmark numbers should be reported from an actual run rather than assumed from source code alone.

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    D
    maintenance
    Enables MCP clients to list indexed PDF document collections and perform semantic search queries on them using locally extracted text and embeddings.
    2
    AGPL 3.0
  • F
    license
    Not graded
    quality
    C
    maintenance
    Enables local document question-answering and retrieval via MCP, supporting multi-turn conversation, intent recognition, and tools for document search, Q&A, and summarization.
    5
    -
  • F
    license
    Not graded
    quality
    D
    maintenance
    Enables natural language search over PDF documents using vector search, allowing MCP clients like Claude Desktop to query PDF content.
    1
    -