Skip to main content
Glama
Sharvary-HH

DocAgent

by Sharvary-HH

DocAgent — RAG + Agent System with Document AI

An end-to-end retrieval-augmented generation (RAG) pipeline plus a tool-using LangGraph agent over your own documents. Upload PDFs, scanned images, or text files; DocAgent chunks and embeds them into Qdrant, retrieves and reranks relevant context, and answers with citations — either via a simple grounded QA endpoint, or via an agent that can also call a calculator, search the web, and OCR documents on demand.

Built for ML/GenAI engineer portfolios: RAG, agents, Document AI, local inference, and observability, all running fully offline via Ollama.

Architecture

                 ┌─────────────┐        ┌──────────────┐
 PDF/image/text →│  Ingestion  │→ chunks→│   Qdrant     │
                 │ (OCR via    │        │ (vector DB)  │
                 │  Tesseract) │        └──────┬───────┘
                 └─────────────┘               │ similarity search
                                                ▼
                                        ┌───────────────┐
                                        │ Cross-encoder │
                                        │   reranker    │
                                        └──────┬────────┘
                                               ▼
        FastAPI  ── /query ─────────►  Grounded QA (Ollama, JSON + Pydantic,
                                        citations)
                 ── /agent/chat ────►  LangGraph ReAct agent
                                        tools: retrieve_documents, calculator,
                                        web_search (Tavily), ocr_scan
                                        └─ also exposed via MCP (mcp_server.py)
                                        └─ every run logged to runs/*.jsonl

        Streamlit UI ── talks to FastAPI, shows answers + citations + tool calls

Related MCP server: py-mcp-qdrant-rag

Tech stack

Python, FastAPI, Pydantic v2, LangChain + LangGraph, Ollama (local LLM + embeddings), Qdrant (vector DB), a small HF sentence-transformers cross-encoder for reranking, Tesseract OCR, MCP (Model Context Protocol), Streamlit, Docker Compose.

Quickstart (Docker Compose — easiest)

cp .env.example .env          # edit if you want a different model or a Tavily key
docker compose up -d qdrant ollama backend streamlit
./scripts/pull_ollama_models.sh   # pulls the LLM + embedding model into Ollama

Then open:

macOS tip: running Ollama natively (brew install ollama, then ollama serve) is noticeably faster than the containerized ollama service since it gets Metal acceleration. If you do that, stop the ollama service in compose and set OLLAMA_BASE_URL=http://host.docker.internal:11434 in .env.

Local dev (no Docker for the backend)

python3 -m venv .venv && source .venv/bin/activate
pip install -r requirements-dev.txt
docker compose up -d qdrant ollama   # or run both natively
./scripts/pull_ollama_models.sh
uvicorn app.main:app --reload
# in another terminal:
streamlit run streamlit_app/app.py

Requires tesseract and poppler on PATH for OCR (brew install tesseract poppler on macOS).

API

  • POST /ingest — multipart file upload (pdf, png/jpg/tiff/bmp, txt/md). Scanned PDF pages and images are OCR'd automatically. Returns doc_id.

  • POST /query{"question": "...", "doc_id": "optional"} → grounded, Pydantic-validated answer with citations (AnswerResponse). No tool use.

  • POST /agent/chat{"message": "...", "doc_id": "optional"} → LangGraph agent response, may call retrieve_documents, calculator, web_search, ocr_scan across multiple steps. Returns citations + a tool-call trace.

  • GET /health — reports whether Qdrant and Ollama are reachable.

MCP server

The same four tools are exposed as an MCP server for external MCP clients (e.g. Claude Desktop):

python -m app.mcp_server

Run tracing

Every /query and /agent/chat call appends a structured JSON record (retrieved chunks, tool calls, latencies, final answer) to runs/<date>.jsonl — no external service required. Set LANGCHAIN_API_KEY in .env to additionally send traces to LangSmith.

Tests

pytest                 # unit tests always run; live-service API test
                        # auto-skips unless Qdrant + Ollama are reachable

Notes on deviations from a pure Ollama stack

  • Reranking uses a small local HF cross-encoder (cross-encoder/ms-marco-MiniLM-L-6-v2) since Ollama has no reranking model type — it's a separate, tiny CPU model, not a generation backend.

  • Web search (web_search tool) uses Tavily and requires TAVILY_API_KEY in .env; without it the tool reports itself unavailable and the agent falls back to documents/general knowledge.

License

MIT

A
license - permissive license
Not graded
quality - not tested
C
maintenance

Maintenance

Maintainers
Response time
Release cycle
Releases (12mo)
Commit activity

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Servers

  • F
    license
    Not graded
    quality
    D
    maintenance
    Enables semantic search and retrieval of information from technical documentation PDFs using RAG-powered natural language queries with Ollama embeddings and LLMs.
    6
  • F
    license
    Not graded
    quality
    D
    maintenance
    Enables semantic search and retrieval-augmented generation (RAG) using Qdrant vector database. Supports indexing documents from URLs and local directories, with flexible embedding options using Ollama or OpenAI.
    2
  • F
    license
    Not graded
    quality
    D
    maintenance
    A local document intelligence and knowledge management server for Claude Desktop that provides RAG-powered Q\&A, media transcription, and URL crawling. It features 11 tools for processing various file types and managing a persistent local vector store with zero infrastructure costs.
    1

View all related MCP servers

Related MCP Connectors

  • Persistent memory and knowledge management for AI agents with semantic search and 50+ tools.

  • OCR, transcription, file extraction, and image generation for AI agents via MCP.

  • Persistent memory and knowledge graphs for AI agents. Hybrid search, context checkpoints, and more.

View all MCP Connectors

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/Sharvary-HH/DocAgent'

If you have feedback or need assistance with the MCP directory API, please join our Discord server