Skip to main content
Glama
README.md
# Multimodal RAG Service

Production-minded retrieval for text, tables, images, and PDFs. Sentence-transformer text vectors and CLIP image vectors are stored in separate FAISS indexes; BM25 sparse results are combined using weighted reciprocal-rank fusion.

## Run locally

```powershell
python -m venv .venv
.venv\\Scripts\\Activate.ps1
pip install -e ".[dev]"
Copy-Item .env.example .env
uvicorn app.main:app --reload
```

The API documentation is available at `http://localhost:8000/docs`. Models download automatically at their first embedding request.

## API

```bash
curl -F "files=@report.pdf" http://localhost:8000/v1/documents
curl -X POST http://localhost:8000/v1/search -H "Content-Type: application/json" -d "{\"query\":\"revenue trend\",\"top_k\":5}"
curl -N -X POST http://localhost:8000/v1/chat/stream -H "Content-Type: application/json" -d "{\"query\":\"Summarize the report\"}"
```

## MCP integration

Run the REST API, then start the stdio MCP bridge:

```bash
rag-mcp
```

Set `RAG_API_URL` to point at a non-default API address. The bridge provides `search_knowledge_base`, `ask_knowledge_base`, `list_documents`, and `service_health` tools to MCP clients.

## Evaluation

Supply JSONL rows containing `query`, `relevant_document_ids`, and optionally `reference_answer`:

```bash
python -m app.evaluation --dataset eval/golden.jsonl
```

Reported metrics include Recall@K, Precision@K, MRR, p50/p95 retrieval latency, ROUGE-L, and BLEU-1.

## Container deployment

```bash
docker compose up --build
```

The compose configuration mounts a persistent named volume at `/service/data`.

TDQS

A3.8/5.0

Scored across 4 tools

Disambiguation5/5

Each tool has a clearly distinct purpose: searching returns chunks, asking returns an answer with citations, listing documents provides an index, and health checks API availability. No overlap or ambiguity.

Naming Consistency4/5

Most tools follow a verb_noun pattern (search_knowledge_base, ask_knowledge_base, list_documents), but service_health breaks the pattern by using a noun phrase. The naming is still clear and consistent in snake_case.

Tool Count5/5

With only 4 tools, the server is tightly scoped for a RAG query service. Each tool is essential and there is no bloat or redundancy.

Completeness4/5

The core query workflow (search, ask, list) is covered, and health monitoring is included. However, there is no tool to add or remove documents, which may be a gap if the server is expected to manage the index, but it might be intentionally read-only.

Maintenance

ActivitySlowing
ResponsivenessNo issues