Skip to main content
Glama
Sharvary-HH

DocAgent

by Sharvary-HH

DocAgent — RAG + Agent System with Document AI

An end-to-end retrieval-augmented generation (RAG) pipeline plus a tool-using LangGraph agent over your own documents. Upload PDFs, scanned images, or text files; DocAgent chunks and embeds them into Qdrant, retrieves and reranks relevant context, and answers with citations — either via a simple grounded QA endpoint, or via an agent that can also call a calculator, search the web, and OCR documents on demand.

Built for ML/GenAI engineer portfolios: RAG, agents, Document AI, local inference, and observability, all running fully offline via Ollama.

Architecture

                 ┌─────────────┐        ┌──────────────┐
 PDF/image/text →│  Ingestion  │→ chunks→│   Qdrant     │
                 │ (OCR via    │        │ (vector DB)  │
                 │  Tesseract) │        └──────┬───────┘
                 └─────────────┘               │ similarity search
                                                ▼
                                        ┌───────────────┐
                                        │ Cross-encoder │
                                        │   reranker    │
                                        └──────┬────────┘
                                               ▼
        FastAPI  ── /query ─────────►  Grounded QA (Ollama, JSON + Pydantic,
                                        citations)
                 ── /agent/chat ────►  LangGraph ReAct agent
                                        tools: retrieve_documents, calculator,
                                        web_search (Tavily), ocr_scan
                                        └─ also exposed via MCP (mcp_server.py)
                                        └─ every run logged to runs/*.jsonl

        Streamlit UI ── talks to FastAPI, shows answers + citations + tool calls

Related MCP server: py-mcp-qdrant-rag

Tech stack

Python, FastAPI, Pydantic v2, LangChain + LangGraph, Ollama (local LLM + embeddings), Qdrant (vector DB), a small HF sentence-transformers cross-encoder for reranking, Tesseract OCR, MCP (Model Context Protocol), Streamlit, Docker Compose.

Quickstart (Docker Compose — easiest)

cp .env.example .env          # edit if you want a different model or a Tavily key
docker compose up -d qdrant ollama backend streamlit
./scripts/pull_ollama_models.sh   # pulls the LLM + embedding model into Ollama

Then open:

macOS tip: running Ollama natively (brew install ollama, then ollama serve) is noticeably faster than the containerized ollama service since it gets Metal acceleration. If you do that, stop the ollama service in compose and set OLLAMA_BASE_URL=http://host.docker.internal:11434 in .env.

Local dev (no Docker for the backend)

python3 -m venv .venv && source .venv/bin/activate
pip install -r requirements-dev.txt
docker compose up -d qdrant ollama   # or run both natively
./scripts/pull_ollama_models.sh
uvicorn app.main:app --reload
# in another terminal:
streamlit run streamlit_app/app.py

Requires tesseract and poppler on PATH for OCR (brew install tesseract poppler on macOS).

API

  • POST /ingest — multipart file upload (pdf, png/jpg/tiff/bmp, txt/md). Scanned PDF pages and images are OCR'd automatically. Returns doc_id.

  • POST /query{"question": "...", "doc_id": "optional"} → grounded, Pydantic-validated answer with citations (AnswerResponse). No tool use.

  • POST /agent/chat{"message": "...", "doc_id": "optional"} → LangGraph agent response, may call retrieve_documents, calculator, web_search, ocr_scan across multiple steps. Returns citations + a tool-call trace.

  • GET /health — reports whether Qdrant and Ollama are reachable.

MCP server

The same four tools are exposed as an MCP server for external MCP clients (e.g. Claude Desktop):

python -m app.mcp_server

Run tracing

Every /query and /agent/chat call appends a structured JSON record (retrieved chunks, tool calls, latencies, final answer) to runs/<date>.jsonl — no external service required. Set LANGCHAIN_API_KEY in .env to additionally send traces to LangSmith.

Tests

pytest                 # unit tests always run; live-service API test
                        # auto-skips unless Qdrant + Ollama are reachable

Notes on deviations from a pure Ollama stack

  • Reranking uses a small local HF cross-encoder (cross-encoder/ms-marco-MiniLM-L-6-v2) since Ollama has no reranking model type — it's a separate, tiny CPU model, not a generation backend.

  • Web search (web_search tool) uses Tavily and requires TAVILY_API_KEY in .env; without it the tool reports itself unavailable and the agent falls back to documents/general knowledge.

License

MIT

Related MCP Connectors

Related MCP Servers

  • F
    license
    Not graded
    quality
    D
    maintenance
    Enables semantic search and retrieval of information from technical documentation PDFs using RAG-powered natural language queries with Ollama embeddings and LLMs.
    6
    -
  • F
    license
    Not graded
    quality
    D
    maintenance
    Enables semantic search and retrieval-augmented generation (RAG) using Qdrant vector database. Supports indexing documents from URLs and local directories, with flexible embedding options using Ollama or OpenAI.
    2
    -
  • F
    license
    Not graded
    quality
    B
    maintenance
    Provides retrieval-augmented generation for insurance claims, enabling search, clause retrieval, and governed tool-calling over policy documents using local LLM (Ollama).
    -