DocAgent FastMCP
Provides local embedding models for the retrieve_documents tool, enabling local-first RAG over ingested documents.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@DocAgent FastMCPsearch my uploaded documents for the refund policy and cite the relevant pages"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
DocAgent
Overview
DocAgent is a document intelligence platform built by Hardik Kaurani around three ideas:
RAG for grounded document answers
LangGraph for tool-using agent workflows
MCP for exposing the same tools to external clients
Documents are loaded page-by-page, scanned pages can fall back to OCR, the text is split into overlapping chunks, embeddings are stored in Qdrant, retrieved candidates are reranked with a cross-encoder, and an Ollama-hosted LLM generates the final grounded answer.
The agent can additionally call a calculator, Tavily web search, and on-demand OCR.
Related MCP server: jina-mcp-server
Why DocAgent?
A basic document chatbot is usually:
Question
↓
Vector Search
↓
LLM
↓
AnswerDocAgent adds a decision-making layer:
Question
↓
LangGraph Agent
↓
Tool Selection
├── Document Retrieval
├── Calculator
├── Web Search
└── OCR
↓
Grounded Response + Citations + TraceThis separation lets deterministic operations such as retrieval and calculation remain tools instead of being delegated to the language model itself.
Architecture
1. System Architecture
flowchart TB
U[User]
U --> UI[Streamlit UI]
U --> API[FastAPI API]
UI --> API
API --> ING[Ingestion]
API --> RAG[RAG Query Pipeline]
API --> AGENT[LangGraph Agent]
ING --> LOAD[Load PDF / Image / TXT / MD]
LOAD --> OCR[Tesseract OCR when needed]
OCR --> CHUNK[Recursive Chunking]
CHUNK --> EMB[Ollama Embeddings]
EMB --> Q[(Qdrant)]
RAG --> RET[Vector Retrieval]
RET --> Q
RET --> RR[Cross-Encoder Reranker]
RR --> GEN[Ollama Generation]
GEN --> RESP[Grounded Answer + Citations]
AGENT --> TOOLS[Shared Tool Layer]
TOOLS --> RET
TOOLS --> CALC[Calculator]
TOOLS --> WEB[Tavily Search]
TOOLS --> AOCR[On-demand OCR]
MCP[MCP Client] --> SERVER[FastMCP Server]
SERVER --> TOOLS
RESP --> TRACE[JSONL Run Trace]
AGENT --> TRACE2. Document Ingestion & Retrieval
flowchart LR
D[Uploaded File] --> TYPE{Format}
TYPE -->|PDF| PDF[pypdf]
TYPE -->|Image| IMG[Tesseract OCR]
TYPE -->|TXT / MD| TXT[Text Loader]
PDF --> CHECK{Extracted text < 20 chars?}
CHECK -->|Yes| POCR[OCR PDF Page]
CHECK -->|No| PLAIN[Extracted Page Text]
POCR --> CHUNK[800-char chunks<br/>120-char overlap]
PLAIN --> CHUNK
IMG --> CHUNK
TXT --> CHUNK
CHUNK --> EMB[Ollama Embedding]
EMB --> STORE[(Qdrant)]
Q[User Query] --> QEMB[Query Embedding]
QEMB --> SEARCH[Top-K Vector Search]
STORE --> SEARCH
SEARCH --> CAND[Up to 20 candidates]
CAND --> RANK[Cross-Encoder Reranking]
RANK --> FINAL[Top 5 chunks]
FINAL --> CONTEXT[Grounding Context]
CONTEXT --> LLM[Ollama LLM]
LLM --> ANSWER[Answer + Citation Metadata]3. Agent Tool-Calling Workflow
flowchart TD
START[User Message] --> AGENT[LangGraph ReAct Agent]
AGENT --> DECIDE{Need a tool?}
DECIDE -->|Document question| RET[retrieve_documents]
DECIDE -->|Arithmetic| CALC[calculator]
DECIDE -->|External information| WEB[web_search]
DECIDE -->|Scanned file| OCR[ocr_scan]
DECIDE -->|No| FINAL[Final Answer]
RET --> AGENT
CALC --> AGENT
WEB --> AGENT
OCR --> AGENT
AGENT --> CHECK{Enough evidence?}
CHECK -->|No| DECIDE
CHECK -->|Yes| FINAL
RET --> CITE[Chunk IDs + metadata]
CITE --> FINAL
FINAL --> TRACE[Run ID + tool-call trace]4. MCP Tool Exposure
sequenceDiagram
participant C as MCP Client
participant S as DocAgent FastMCP
participant T as Shared Tools
participant Q as Qdrant
participant W as Tavily
participant F as Server Files
C->>S: Call tool
alt retrieve_documents
S->>T: retrieve_documents(query, doc_id?)
T->>Q: Vector search + filtering
Q-->>T: Candidate chunks
T-->>S: Ranked document evidence
else calculator
S->>T: calculator(expression)
T-->>S: Safe arithmetic result
else web_search
S->>T: web_search(query)
T->>W: Tavily search
W-->>T: Search results
T-->>S: Web evidence
else ocr_scan
S->>T: ocr_scan(file_path, page)
T->>F: Read file
T-->>T: Tesseract OCR
T-->>S: Extracted text
end
S-->>C: Structured tool resultCore Features
Agentic RAG
LangGraph orchestrates a ReAct-style loop where the model can call the appropriate tool before producing a final response.
Grounded Document QA
The direct RAG path retrieves relevant chunks, reranks them, and asks the LLM to answer only from the supplied context.
Responses expose citation metadata and a grounded flag.
Multi-format Ingestion
Supported uploads:
PDF
PNG
JPG / JPEG
TIFF
BMP
TXT
Markdown
OCR Fallback
PDF pages with less than the configured minimum extractable text are treated as scanned pages and passed through Tesseract OCR.
Two-stage Retrieval
Query
↓
Ollama embedding
↓
Qdrant similarity search
↓
Up to 20 candidates
↓
Cross-encoder reranking
↓
Top 5 passagesLocal-first Inference
Ollama provides the chat model and embedding model, allowing the primary RAG workflow to run locally.
MCP Integration
The same tools used by the LangGraph agent are exposed through FastMCP so compatible MCP clients can call DocAgent capabilities directly.
Traceable Runs
Queries and agent executions are written to daily JSONL trace files with run IDs, events, results, and latency information.
Optional LangSmith tracing can be enabled through environment variables.
Dockerized Stack
Docker Compose provides:
FastAPI backend
Streamlit frontend
Qdrant
Ollama
Tooling
Tool | Purpose |
| Search and rerank ingested document chunks |
| Safely evaluate basic arithmetic |
| Search the public web through Tavily |
| OCR an image or a specific PDF page |
The calculator uses Python's AST parser and permits only arithmetic operations:
+ - * / ** %
()Function calls, variable references, and arbitrary Python execution are not allowed.
API
DocAgent exposes a FastAPI REST API.
Method | Endpoint | Description |
GET |
| Check API, Qdrant, and Ollama availability |
POST |
| Upload, parse, chunk, embed, and index a document |
POST |
| Run the direct grounded RAG pipeline |
POST |
| Run the LangGraph agent |
Interactive API documentation:
http://localhost:8000/docsExample Requests
Ingest
curl -X POST "http://localhost:8000/ingest" \
-F "file=@document.pdf"Direct RAG query
curl -X POST "http://localhost:8000/query" \
-H "Content-Type: application/json" \
-d '{
"question": "What does the document say about deployment?"
}'Document-scoped query
curl -X POST "http://localhost:8000/query" \
-H "Content-Type: application/json" \
-d '{
"question": "Summarize the deployment process.",
"doc_id": "DOCUMENT_ID"
}'Agent chat
curl -X POST "http://localhost:8000/agent/chat" \
-H "Content-Type: application/json" \
-d '{
"message": "Find the relevant information and calculate the total."
}'Quick Start
1. Clone
git clone https://github.com/hardikkaurani/DocAgent.git
cd DocAgent2. Configure environment
cp .env.example .envExample configuration:
OLLAMA_BASE_URL=http://localhost:11434
OLLAMA_LLM_MODEL=llama3.1
OLLAMA_EMBED_MODEL=nomic-embed-text
QDRANT_URL=http://localhost:6333
QDRANT_COLLECTION=docagent_chunks
RERANKER_MODEL=cross-encoder/ms-marco-MiniLM-L-6-v2
TOP_K_RETRIEVE=20
TOP_K_RERANK=5
TAVILY_API_KEY=
LANGCHAIN_API_KEY=
LANGCHAIN_PROJECT=docagent
UPLOAD_DIR=./data/uploads
RUNS_DIR=./runsTAVILY_API_KEY is optional. Without it, live web search is unavailable to the agent.
3. Docker
Start the full stack:
docker compose up -dPull the required Ollama models:
./scripts/pull_ollama_models.shServices:
Service | Address |
Streamlit |
|
FastAPI |
|
Swagger |
|
Qdrant |
|
Ollama |
|
4. Local development
Create a virtual environment:
python3 -m venv .venv
source .venv/bin/activateInstall dependencies:
pip install -r requirements-dev.txtStart infrastructure:
docker compose up -d qdrant ollamaStart FastAPI:
uvicorn app.main:app --reloadStart Streamlit:
streamlit run streamlit_app/app.pyFor OCR support, install Tesseract and Poppler and make sure both are available on the system PATH.
MCP Server
DocAgent exposes the shared tool layer through FastMCP.
Start the MCP server:
python -m app.mcp_serverAvailable MCP operations:
calculator_toolretrieve_documents_toolweb_search_toolocr_scan_tool
The MCP server does not reimplement the underlying functionality. It reuses the same tool functions used by the LangGraph agent.
Testing
Run:
pytestThe repository includes tests for:
schema validation
chunking behavior
citation extraction
ingest/query API behavior
Tests requiring external infrastructure can skip when the relevant services are unavailable.
Observability
Every query and agent execution receives a unique run ID.
Local trace files are written as:
runs/
└── YYYY-MM-DD.jsonlA trace can contain:
input payload
retrieval events
reranking events
tool calls
generated result
latency
run ID
Optional LangSmith integration:
LANGCHAIN_API_KEY=...
LANGCHAIN_PROJECT=docagentProject Structure
DocAgent/
├── app/
│ ├── agent/
│ │ ├── graph.py
│ │ └── tools.py
│ ├── api/
│ │ ├── routes_agent.py
│ │ ├── routes_health.py
│ │ ├── routes_ingest.py
│ │ └── routes_query.py
│ ├── generation/
│ │ ├── llm.py
│ │ └── qa_chain.py
│ ├── ingestion/
│ │ ├── chunking.py
│ │ ├── loaders.py
│ │ └── ocr.py
│ ├── retrieval/
│ │ ├── reranker.py
│ │ └── retriever.py
│ ├── tracing/
│ │ └── tracer.py
│ ├── vectorstore/
│ │ ├── embeddings.py
│ │ └── qdrant_store.py
│ ├── config.py
│ ├── main.py
│ ├── mcp_server.py
│ └── schemas.py
├── streamlit_app/
│ └── app.py
├── tests/
├── scripts/
│ └── pull_ollama_models.sh
├── data/
├── Dockerfile
├── Dockerfile.streamlit
├── docker-compose.yml
├── requirements.txt
├── requirements-dev.txt
├── .env.example
├── .gitignore
├── LICENSE
└── README.mdTechnology Stack
Layer | Technology |
Language | Python |
API | FastAPI |
Validation | Pydantic v2 |
Agent orchestration | LangGraph |
LLM integration | LangChain + Ollama |
LLM runtime | Ollama |
Embeddings | Ollama |
Vector database | Qdrant |
Reranking | Sentence Transformers Cross-Encoder |
PDF parsing | pypdf |
OCR | Tesseract + pdf2image |
Web search | Tavily |
Tool protocol | MCP / FastMCP |
Frontend | Streamlit |
HTTP client | HTTPX |
Testing | Pytest |
Containers | Docker + Docker Compose |
CI/CD | GitHub Actions |
Design Principles
Ground first, generate second
Direct RAG answers are generated from retrieved context. When the context is insufficient, the system can return an ungrounded response rather than pretending the evidence exists.
Explicit tool boundaries
Retrieval, arithmetic, web search, and OCR are explicit tool calls, making agent behavior easier to inspect and trace.
One tool layer, multiple interfaces
The LangGraph agent and MCP server reuse the same underlying implementations.
Local-first by default
Ollama handles the primary LLM and embedding workloads while Qdrant provides vector storage and retrieval.
Evidence stays attached
Chunks retain:
chunk_id
doc_id
filename
page
text
scoreThis allows the application to expose the evidence associated with generated answers.
License
MIT License. See LICENSE for the complete license text.
Author
Hardik Kaurani
Software Developer focused on AI, backend engineering, RAG systems, and open source.
GitHub: https://github.com/hardikkaurani
Repository: https://github.com/hardikkaurani/DocAgent
This server cannot be deployed
Maintenance
Related MCP Connectors
Find, vet, and run MCP tools through a secure audited gateway with prompt-injection risk scoring
Pay-per-use tool marketplace for AI agents. Search, price-check, and call APIs via MCP.
The OpenRouter for tools. One MCP connection gives any AI agent 254 hosted tools, pay per call.
OCR, transcription, file extraction, and image generation for AI agents via MCP.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceEnables file system operations, web scraping, and AI-powered search through MCP tools for use by LLM agents.1-
- AlicenseNot gradedqualityCmaintenanceEnables remote MCP access to tools for URL-to-markdown conversion, web and image search, academic search, embeddings, reranking, classification, deduplication, and PDF extraction.Apache 2.0
- AlicenseNot gradedqualityCmaintenanceEnables models to query enterprise systems—structured data, live GitHub REST API, and unstructured document embeddings—through a single, per-caller scoped MCP tool surface with task-shaped tools and recovery-aware errors.138 npmISC
- AlicenseAqualityCmaintenanceProvides a document store, web search, and safe calculator as MCP tools, enabling both offline and LLM-driven clients to discover and call them over the Model Context Protocol.6MIT