Skip to main content
Glama
README.md
# DocAgent

<p align="center">
  <strong>Agentic Document Intelligence with RAG, LangGraph & MCP</strong>
</p>

<p align="center">
  A local-first document AI system that ingests your files, retrieves grounded evidence, reranks it, and lets an agent reason across retrieval, calculation, web search, and OCR.
</p>

<p align="center">
  <a href="https://github.com/hardikkaurani/DocAgent">Repository</a>
  ·
  <a href="https://github.com/hardikkaurani/DocAgent/issues">Issues</a>
  ·
  <a href="https://github.com/hardikkaurani">Hardik Kaurani</a>
</p>

<p align="center">
  <img src="https://img.shields.io/badge/Python-3.x-3776AB?style=flat-square&logo=python&logoColor=white" />
  <img src="https://img.shields.io/badge/FastAPI-005571?style=flat-square&logo=fastapi&logoColor=white" />
  <img src="https://img.shields.io/badge/LangGraph-Agentic%20RAG-1C3C3C?style=flat-square" />
  <img src="https://img.shields.io/badge/Ollama-Local%20LLM-000000?style=flat-square" />
  <img src="https://img.shields.io/badge/Qdrant-Vector%20DB-DC244C?style=flat-square" />
  <img src="https://img.shields.io/badge/MCP-FastMCP-000000?style=flat-square" />
  <img src="https://img.shields.io/badge/Streamlit-UI-FF4B4B?style=flat-square&logo=streamlit&logoColor=white" />
  <img src="https://img.shields.io/badge/Docker-Containerized-2496ED?style=flat-square&logo=docker&logoColor=white" />
</p>

---

## Overview

**DocAgent** is a document intelligence platform built by **Hardik Kaurani** around three ideas:

1. **RAG for grounded document answers**
2. **LangGraph for tool-using agent workflows**
3. **MCP for exposing the same tools to external clients**

Documents are loaded page-by-page, scanned pages can fall back to OCR, the text is split into overlapping chunks, embeddings are stored in Qdrant, retrieved candidates are reranked with a cross-encoder, and an Ollama-hosted LLM generates the final grounded answer.

The agent can additionally call a calculator, Tavily web search, and on-demand OCR.

## Why DocAgent?

A basic document chatbot is usually:

```text
Question
   ↓
Vector Search
   ↓
LLM
   ↓
Answer
```

DocAgent adds a decision-making layer:

```text
Question
   ↓
LangGraph Agent
   ↓
Tool Selection
   ├── Document Retrieval
   ├── Calculator
   ├── Web Search
   └── OCR
   ↓
Grounded Response + Citations + Trace
```

This separation lets deterministic operations such as retrieval and calculation remain tools instead of being delegated to the language model itself.

---

# Architecture

## 1. System Architecture

```mermaid
flowchart TB
    U[User]

    U --> UI[Streamlit UI]
    U --> API[FastAPI API]
    UI --> API

    API --> ING[Ingestion]
    API --> RAG[RAG Query Pipeline]
    API --> AGENT[LangGraph Agent]

    ING --> LOAD[Load PDF / Image / TXT / MD]
    LOAD --> OCR[Tesseract OCR when needed]
    OCR --> CHUNK[Recursive Chunking]
    CHUNK --> EMB[Ollama Embeddings]
    EMB --> Q[(Qdrant)]

    RAG --> RET[Vector Retrieval]
    RET --> Q
    RET --> RR[Cross-Encoder Reranker]
    RR --> GEN[Ollama Generation]
    GEN --> RESP[Grounded Answer + Citations]

    AGENT --> TOOLS[Shared Tool Layer]
    TOOLS --> RET
    TOOLS --> CALC[Calculator]
    TOOLS --> WEB[Tavily Search]
    TOOLS --> AOCR[On-demand OCR]

    MCP[MCP Client] --> SERVER[FastMCP Server]
    SERVER --> TOOLS

    RESP --> TRACE[JSONL Run Trace]
    AGENT --> TRACE
```

## 2. Document Ingestion & Retrieval

```mermaid
flowchart LR
    D[Uploaded File] --> TYPE{Format}

    TYPE -->|PDF| PDF[pypdf]
    TYPE -->|Image| IMG[Tesseract OCR]
    TYPE -->|TXT / MD| TXT[Text Loader]

    PDF --> CHECK{Extracted text < 20 chars?}
    CHECK -->|Yes| POCR[OCR PDF Page]
    CHECK -->|No| PLAIN[Extracted Page Text]

    POCR --> CHUNK[800-char chunks<br/>120-char overlap]
    PLAIN --> CHUNK
    IMG --> CHUNK
    TXT --> CHUNK

    CHUNK --> EMB[Ollama Embedding]
    EMB --> STORE[(Qdrant)]

    Q[User Query] --> QEMB[Query Embedding]
    QEMB --> SEARCH[Top-K Vector Search]
    STORE --> SEARCH
    SEARCH --> CAND[Up to 20 candidates]
    CAND --> RANK[Cross-Encoder Reranking]
    RANK --> FINAL[Top 5 chunks]
    FINAL --> CONTEXT[Grounding Context]
    CONTEXT --> LLM[Ollama LLM]
    LLM --> ANSWER[Answer + Citation Metadata]
```

## 3. Agent Tool-Calling Workflow

```mermaid
flowchart TD
    START[User Message] --> AGENT[LangGraph ReAct Agent]
    AGENT --> DECIDE{Need a tool?}

    DECIDE -->|Document question| RET[retrieve_documents]
    DECIDE -->|Arithmetic| CALC[calculator]
    DECIDE -->|External information| WEB[web_search]
    DECIDE -->|Scanned file| OCR[ocr_scan]
    DECIDE -->|No| FINAL[Final Answer]

    RET --> AGENT
    CALC --> AGENT
    WEB --> AGENT
    OCR --> AGENT

    AGENT --> CHECK{Enough evidence?}
    CHECK -->|No| DECIDE
    CHECK -->|Yes| FINAL

    RET --> CITE[Chunk IDs + metadata]
    CITE --> FINAL
    FINAL --> TRACE[Run ID + tool-call trace]
```

## 4. MCP Tool Exposure

```mermaid
sequenceDiagram
    participant C as MCP Client
    participant S as DocAgent FastMCP
    participant T as Shared Tools
    participant Q as Qdrant
    participant W as Tavily
    participant F as Server Files

    C->>S: Call tool

    alt retrieve_documents
        S->>T: retrieve_documents(query, doc_id?)
        T->>Q: Vector search + filtering
        Q-->>T: Candidate chunks
        T-->>S: Ranked document evidence
    else calculator
        S->>T: calculator(expression)
        T-->>S: Safe arithmetic result
    else web_search
        S->>T: web_search(query)
        T->>W: Tavily search
        W-->>T: Search results
        T-->>S: Web evidence
    else ocr_scan
        S->>T: ocr_scan(file_path, page)
        T->>F: Read file
        T-->>T: Tesseract OCR
        T-->>S: Extracted text
    end

    S-->>C: Structured tool result
```

---

# Core Features

### Agentic RAG

LangGraph orchestrates a ReAct-style loop where the model can call the appropriate tool before producing a final response.

### Grounded Document QA

The direct RAG path retrieves relevant chunks, reranks them, and asks the LLM to answer **only from the supplied context**.

Responses expose citation metadata and a `grounded` flag.

### Multi-format Ingestion

Supported uploads:

* PDF
* PNG
* JPG / JPEG
* TIFF
* BMP
* TXT
* Markdown

### OCR Fallback

PDF pages with less than the configured minimum extractable text are treated as scanned pages and passed through Tesseract OCR.

### Two-stage Retrieval

```text
Query
  ↓
Ollama embedding
  ↓
Qdrant similarity search
  ↓
Up to 20 candidates
  ↓
Cross-encoder reranking
  ↓
Top 5 passages
```

### Local-first Inference

Ollama provides the chat model and embedding model, allowing the primary RAG workflow to run locally.

### MCP Integration

The same tools used by the LangGraph agent are exposed through FastMCP so compatible MCP clients can call DocAgent capabilities directly.

### Traceable Runs

Queries and agent executions are written to daily JSONL trace files with run IDs, events, results, and latency information.

Optional LangSmith tracing can be enabled through environment variables.

### Dockerized Stack

Docker Compose provides:

* FastAPI backend
* Streamlit frontend
* Qdrant
* Ollama

---

# Tooling

| Tool                 | Purpose                                    |
| -------------------- | ------------------------------------------ |
| `retrieve_documents` | Search and rerank ingested document chunks |
| `calculator`         | Safely evaluate basic arithmetic           |
| `web_search`         | Search the public web through Tavily       |
| `ocr_scan`           | OCR an image or a specific PDF page        |

The calculator uses Python's AST parser and permits only arithmetic operations:

```text
+  -  *  /  **  %
()
```

Function calls, variable references, and arbitrary Python execution are not allowed.

---

# API

DocAgent exposes a FastAPI REST API.

| Method | Endpoint      | Description                                       |
| ------ | ------------- | ------------------------------------------------- |
| GET    | `/health`     | Check API, Qdrant, and Ollama availability        |
| POST   | `/ingest`     | Upload, parse, chunk, embed, and index a document |
| POST   | `/query`      | Run the direct grounded RAG pipeline              |
| POST   | `/agent/chat` | Run the LangGraph agent                           |

Interactive API documentation:

```text
http://localhost:8000/docs
```

## Example Requests

### Ingest

```bash
curl -X POST "http://localhost:8000/ingest" \
  -F "file=@document.pdf"
```

### Direct RAG query

```bash
curl -X POST "http://localhost:8000/query" \
  -H "Content-Type: application/json" \
  -d '{
    "question": "What does the document say about deployment?"
  }'
```

### Document-scoped query

```bash
curl -X POST "http://localhost:8000/query" \
  -H "Content-Type: application/json" \
  -d '{
    "question": "Summarize the deployment process.",
    "doc_id": "DOCUMENT_ID"
  }'
```

### Agent chat

```bash
curl -X POST "http://localhost:8000/agent/chat" \
  -H "Content-Type: application/json" \
  -d '{
    "message": "Find the relevant information and calculate the total."
  }'
```

---

# Quick Start

## 1. Clone

```bash
git clone https://github.com/hardikkaurani/DocAgent.git
cd DocAgent
```

## 2. Configure environment

```bash
cp .env.example .env
```

Example configuration:

```env
OLLAMA_BASE_URL=http://localhost:11434
OLLAMA_LLM_MODEL=llama3.1
OLLAMA_EMBED_MODEL=nomic-embed-text

QDRANT_URL=http://localhost:6333
QDRANT_COLLECTION=docagent_chunks

RERANKER_MODEL=cross-encoder/ms-marco-MiniLM-L-6-v2

TOP_K_RETRIEVE=20
TOP_K_RERANK=5

TAVILY_API_KEY=
LANGCHAIN_API_KEY=
LANGCHAIN_PROJECT=docagent

UPLOAD_DIR=./data/uploads
RUNS_DIR=./runs
```

`TAVILY_API_KEY` is optional. Without it, live web search is unavailable to the agent.

## 3. Docker

Start the full stack:

```bash
docker compose up -d
```

Pull the required Ollama models:

```bash
./scripts/pull_ollama_models.sh
```

Services:

| Service   | Address                      |
| --------- | ---------------------------- |
| Streamlit | `http://localhost:8501`      |
| FastAPI   | `http://localhost:8000`      |
| Swagger   | `http://localhost:8000/docs` |
| Qdrant    | `http://localhost:6333`      |
| Ollama    | `http://localhost:11434`     |

## 4. Local development

Create a virtual environment:

```bash
python3 -m venv .venv
source .venv/bin/activate
```

Install dependencies:

```bash
pip install -r requirements-dev.txt
```

Start infrastructure:

```bash
docker compose up -d qdrant ollama
```

Start FastAPI:

```bash
uvicorn app.main:app --reload
```

Start Streamlit:

```bash
streamlit run streamlit_app/app.py
```

For OCR support, install **Tesseract** and **Poppler** and make sure both are available on the system `PATH`.

---

# MCP Server

DocAgent exposes the shared tool layer through FastMCP.

Start the MCP server:

```bash
python -m app.mcp_server
```

Available MCP operations:

* `calculator_tool`
* `retrieve_documents_tool`
* `web_search_tool`
* `ocr_scan_tool`

The MCP server does not reimplement the underlying functionality. It reuses the same tool functions used by the LangGraph agent.

---

# Testing

Run:

```bash
pytest
```

The repository includes tests for:

* schema validation
* chunking behavior
* citation extraction
* ingest/query API behavior

Tests requiring external infrastructure can skip when the relevant services are unavailable.

---

# Observability

Every query and agent execution receives a unique run ID.

Local trace files are written as:

```text
runs/
└── YYYY-MM-DD.jsonl
```

A trace can contain:

* input payload
* retrieval events
* reranking events
* tool calls
* generated result
* latency
* run ID

Optional LangSmith integration:

```env
LANGCHAIN_API_KEY=...
LANGCHAIN_PROJECT=docagent
```

---

# Project Structure

```text
DocAgent/
├── app/
│   ├── agent/
│   │   ├── graph.py
│   │   └── tools.py
│   ├── api/
│   │   ├── routes_agent.py
│   │   ├── routes_health.py
│   │   ├── routes_ingest.py
│   │   └── routes_query.py
│   ├── generation/
│   │   ├── llm.py
│   │   └── qa_chain.py
│   ├── ingestion/
│   │   ├── chunking.py
│   │   ├── loaders.py
│   │   └── ocr.py
│   ├── retrieval/
│   │   ├── reranker.py
│   │   └── retriever.py
│   ├── tracing/
│   │   └── tracer.py
│   ├── vectorstore/
│   │   ├── embeddings.py
│   │   └── qdrant_store.py
│   ├── config.py
│   ├── main.py
│   ├── mcp_server.py
│   └── schemas.py
├── streamlit_app/
│   └── app.py
├── tests/
├── scripts/
│   └── pull_ollama_models.sh
├── data/
├── Dockerfile
├── Dockerfile.streamlit
├── docker-compose.yml
├── requirements.txt
├── requirements-dev.txt
├── .env.example
├── .gitignore
├── LICENSE
└── README.md
```

---

# Technology Stack

| Layer               | Technology                          |
| ------------------- | ----------------------------------- |
| Language            | Python                              |
| API                 | FastAPI                             |
| Validation          | Pydantic v2                         |
| Agent orchestration | LangGraph                           |
| LLM integration     | LangChain + Ollama                  |
| LLM runtime         | Ollama                              |
| Embeddings          | Ollama                              |
| Vector database     | Qdrant                              |
| Reranking           | Sentence Transformers Cross-Encoder |
| PDF parsing         | pypdf                               |
| OCR                 | Tesseract + pdf2image               |
| Web search          | Tavily                              |
| Tool protocol       | MCP / FastMCP                       |
| Frontend            | Streamlit                           |
| HTTP client         | HTTPX                               |
| Testing             | Pytest                              |
| Containers          | Docker + Docker Compose             |
| CI/CD               | GitHub Actions                      |

---

# Design Principles

### Ground first, generate second

Direct RAG answers are generated from retrieved context. When the context is insufficient, the system can return an ungrounded response rather than pretending the evidence exists.

### Explicit tool boundaries

Retrieval, arithmetic, web search, and OCR are explicit tool calls, making agent behavior easier to inspect and trace.

### One tool layer, multiple interfaces

The LangGraph agent and MCP server reuse the same underlying implementations.

### Local-first by default

Ollama handles the primary LLM and embedding workloads while Qdrant provides vector storage and retrieval.

### Evidence stays attached

Chunks retain:

```text
chunk_id
doc_id
filename
page
text
score
```

This allows the application to expose the evidence associated with generated answers.

---

# License

MIT License. See [`LICENSE`](LICENSE) for the complete license text.

---

# Author

## Hardik Kaurani

Software Developer focused on AI, backend engineering, RAG systems, and open source.

GitHub: https://github.com/hardikkaurani

Repository: https://github.com/hardikkaurani/DocAgent

---

<p align="center">
  <strong>DocAgent</strong>
  <br />
  Built and maintained by <a href="https://github.com/hardikkaurani">Hardik Kaurani</a>
</p>