Skip to main content
Glama
ufigiabhi

RAG MCP Adapter

by ufigiabhi
README.md
# RAG MCP Adapter

A small, read-only **Model Context Protocol (MCP)** adapter for an existing RAG API.

It exposes the RAG system as standard MCP tools so MCP-compatible clients and agents can:

- check the health of the RAG service
- ask questions against indexed documents
- optionally restrict retrieval to a specific file
- receive answers, sources, evidence scores, and model metadata

## Architecture

```text
MCP Client
    |
    v
rag_mcp_server.py
    |
    | HTTP
    v
RAG API: http://localhost:8003
    |
    v
Qdrant + Elasticsearch + reranker + local vLLM
```

The adapter does not contain the RAG pipeline itself. It calls the existing RAG API through the `RAG_API_URL` environment variable.

## MCP Tools

### `get_rag_health()`

Checks whether the RAG API, Qdrant, and Elasticsearch are available.

### `ask_documents(...)`

Asks a question against the indexed document collection.

Main inputs:

- `question`: required question text
- `filename`: optional file filter
- `limit`: maximum retrieval candidates
- `min_score`: retrieval threshold
- `evidence_min_score`: minimum evidence threshold
- `max_new_tokens`: answer generation limit

The result includes the answer, sources, evidence score, generator model, and orchestrator.

## Project Files

```text
rag-mcp-adapter/
├── rag_mcp_server.py
├── test_mcp_client.py
├── test_mcp_client_basic.py
├── requirements.txt
├── .env.example
└── README.md
```

## Setup

```bash
cd ~/Downloads/rag-mcp-adapter

python3 -m venv .venv
source .venv/bin/activate

python -m pip install --upgrade pip
python -m pip install -r requirements.txt
```

Set the RAG API address:

```bash
export RAG_API_URL=http://localhost:8003
```

## Running the Complete System

### Terminal 1 — start vLLM

```bash
cd ~/vllm-qwen
source .venv/bin/activate

export VLLM_USE_FLASHINFER_SAMPLER=0

vllm serve "$HOME/models/Qwen2.5-7B-Instruct-AWQ" \
  --served-model-name qwen-local \
  --host 0.0.0.0 \
  --port 8001 \
  --max-model-len 10000 \
  --gpu-memory-utilization 0.88
```

Leave this terminal running.

### Terminal 2 — start the RAG API

```bash
cd ~/Downloads/rag-solution
docker compose up -d

export RAG_API_URL=http://localhost:8003

curl -s "$RAG_API_URL/health" | python3 -m json.tool
```

### Terminal 3 — run the MCP tests

```bash
cd ~/Downloads/rag-mcp-adapter
source .venv/bin/activate

export RAG_API_URL=http://localhost:8003

python test_mcp_client.py
```

The expanded suite currently checks:

- MCP connection and tool discovery
- generated tool schema
- RAG service health
- filtered DOCX retrieval
- filtered TXT retrieval
- full-corpus retrieval
- unsupported-question refusal
- additional document facts
- response metadata
- source provenance
- empty-input validation

Expected summary:

```text
Passed: 12/12
Failed: 0/12
```

A JSON report is written to:

```text
mcp_test_report.json
```

## Notes

- The adapter is intentionally read-only.
- The RAG API must be running at `http://localhost:8003`.
- The local vLLM/Qwen server must be running at
  `http://localhost:8001/v1`.
- Important: the RAG `/health` endpoint can still report that Qdrant
  and Elasticsearch are healthy when vLLM is offline. In that case,
  MCP discovery and health checks may pass, but supported
  `ask_documents` requests can return HTTP `500 Internal Server Error`.
- With both services running, the complete integration test should
  report `12/12` tests passed.
- The adapter can be moved or deployed separately because it only depends on the RAG API URL.
- Do not expose destructive tools such as deleting documents or clearing indexes unless there is a clear requirement and appropriate access control.