Skip to main content
Glama
Sharvary-HH

DocAgent

by Sharvary-HH
README.md
# DocAgent — RAG + Agent System with Document AI

An end-to-end retrieval-augmented generation (RAG) pipeline plus a tool-using
LangGraph agent over your own documents. Upload PDFs, scanned images, or text
files; DocAgent chunks and embeds them into Qdrant, retrieves and reranks
relevant context, and answers with citations — either via a simple grounded
QA endpoint, or via an agent that can also call a calculator, search the web,
and OCR documents on demand.

Built for ML/GenAI engineer portfolios: RAG, agents, Document AI, local
inference, and observability, all running fully offline via Ollama.

## Architecture

```
                 ┌─────────────┐        ┌──────────────┐
 PDF/image/text →│  Ingestion  │→ chunks→│   Qdrant     │
                 │ (OCR via    │        │ (vector DB)  │
                 │  Tesseract) │        └──────┬───────┘
                 └─────────────┘               │ similarity search
                                                ▼
                                        ┌───────────────┐
                                        │ Cross-encoder │
                                        │   reranker    │
                                        └──────┬────────┘
                                               ▼
        FastAPI  ── /query ─────────►  Grounded QA (Ollama, JSON + Pydantic,
                                        citations)
                 ── /agent/chat ────►  LangGraph ReAct agent
                                        tools: retrieve_documents, calculator,
                                        web_search (Tavily), ocr_scan
                                        └─ also exposed via MCP (mcp_server.py)
                                        └─ every run logged to runs/*.jsonl

        Streamlit UI ── talks to FastAPI, shows answers + citations + tool calls
```

## Tech stack
Python, FastAPI, Pydantic v2, LangChain + LangGraph, Ollama (local LLM +
embeddings), Qdrant (vector DB), a small HF `sentence-transformers`
cross-encoder for reranking, Tesseract OCR, MCP (Model Context Protocol),
Streamlit, Docker Compose.

## Quickstart (Docker Compose — easiest)

```bash
cp .env.example .env          # edit if you want a different model or a Tavily key
docker compose up -d qdrant ollama backend streamlit
./scripts/pull_ollama_models.sh   # pulls the LLM + embedding model into Ollama
```

Then open:
- Streamlit UI: http://localhost:8501
- API docs: http://localhost:8000/docs

> **macOS tip:** running Ollama natively (`brew install ollama`, then `ollama
> serve`) is noticeably faster than the containerized `ollama` service since
> it gets Metal acceleration. If you do that, stop the `ollama` service in
> compose and set `OLLAMA_BASE_URL=http://host.docker.internal:11434` in `.env`.

## Local dev (no Docker for the backend)

```bash
python3 -m venv .venv && source .venv/bin/activate
pip install -r requirements-dev.txt
docker compose up -d qdrant ollama   # or run both natively
./scripts/pull_ollama_models.sh
uvicorn app.main:app --reload
# in another terminal:
streamlit run streamlit_app/app.py
```

Requires `tesseract` and `poppler` on PATH for OCR (`brew install tesseract
poppler` on macOS).

## API

- `POST /ingest` — multipart file upload (`pdf`, `png/jpg/tiff/bmp`, `txt/md`).
  Scanned PDF pages and images are OCR'd automatically. Returns `doc_id`.
- `POST /query` — `{"question": "...", "doc_id": "optional"}` → grounded,
  Pydantic-validated answer with citations (`AnswerResponse`). No tool use.
- `POST /agent/chat` — `{"message": "...", "doc_id": "optional"}` → LangGraph
  agent response, may call `retrieve_documents`, `calculator`, `web_search`,
  `ocr_scan` across multiple steps. Returns citations + a tool-call trace.
- `GET /health` — reports whether Qdrant and Ollama are reachable.

## MCP server

The same four tools are exposed as an MCP server for external MCP clients
(e.g. Claude Desktop):

```bash
python -m app.mcp_server
```

## Run tracing

Every `/query` and `/agent/chat` call appends a structured JSON record
(retrieved chunks, tool calls, latencies, final answer) to `runs/<date>.jsonl`
— no external service required. Set `LANGCHAIN_API_KEY` in `.env` to
additionally send traces to LangSmith.

## Tests

```bash
pytest                 # unit tests always run; live-service API test
                        # auto-skips unless Qdrant + Ollama are reachable
```

## Notes on deviations from a pure Ollama stack

- **Reranking** uses a small local HF cross-encoder
  (`cross-encoder/ms-marco-MiniLM-L-6-v2`) since Ollama has no reranking model
  type — it's a separate, tiny CPU model, not a generation backend.
- **Web search** (`web_search` tool) uses Tavily and requires `TAVILY_API_KEY`
  in `.env`; without it the tool reports itself unavailable and the agent
  falls back to documents/general knowledge.

## License

MIT