rag-mcp
by johnylira
README.md
# RAG-MCP Server
**Model Context Protocol (MCP)** server built with **Flask**, **LangGraph**, and **LlamaIndex** for RAG. Designed to be minimal, functional, and easy to extend.
## Overview
- **MCP over HTTP**: `POST /mcp` with JSON-RPC 2.0.
- **LangGraph**: orchestrates the agent with function calling.
- **LlamaIndex**: indexes local documents and answers via RAG.
- **Docker**: `docker compose up --build` and you are done.
---
## Architecture
```
MCP Client → Flask /mcp → LangGraph Agent → Tools
├─ rag_search (LlamaIndex RAG)
├─ agora (UTC datetime)
└─ calcular (arithmetic expression)
```
- The client lists tools (`tools/list`) and calls them (`tools/call`).
- The `ask_agent` tool triggers the LangGraph graph; the LLM decides when to use function calling.
- Other tools are called directly by MCP.
---
## Prerequisites
- **Docker** and **Docker Compose** (or Python 3.11+ locally).
- **OPENAI_API_KEY** (for LLM and embeddings).
---
## Project Structure
```text
.
├── app.py # MCP server + LangGraph + LlamaIndex
├── data/ # RAG corpus (.md, .txt, .pdf, etc.)
│ └── kb.md
├── requirements.txt # Python dependencies
├── Dockerfile
├── docker-compose.yml
├── .env # OPENAI_API_KEY and variables
└── README.md
```
---
## Configuration
### 1. Clone / Create the Project
Create an empty directory and paste the project files (see **Files** section).
### 2. Environment Variables
Create `.env` in the root:
```bash
OPENAI_API_KEY=sk-...
LLM_MODEL=gpt-4o-mini
```
Supported variables:
| Variable | Default | Description |
|-----------------|---------------|------------------------------------------|
| `OPENAI_API_KEY`| (required) | OpenAI API key. |
| `LLM_MODEL` | `gpt-4o-mini` | Model for LLM and embeddings. |
| `DATA_DIR` | `/app/data` | Directory with documents for RAG. |
| `PORT` | `8080` | Server port. |
### 3. RAG Data
Place documents in `data/` (e.g., `kb.md`, `policies.md`, `manuals/`). The server indexes everything on startup.
Minimal example (`data/kb.md`):
```markdown
# Acme Corp
Support SLA: 4 hours during business hours (UTC-3).
Pro Plan costs USD 49/month and includes 10k RAG queries/day.
P1 incidents must be opened in #sre channel.
```
---
## Running
### Docker (recommended)
```bash
docker compose up --build
```
The server runs at `http://127.0.0.1:8080`.
### Local (without Docker)
```bash
pip install -r requirements.txt
export OPENAI_API_KEY=sk-...
export DATA_DIR=./data
python app.py
```
---
## Endpoints
### `POST /mcp` (JSON-RPC 2.0)
MCP uses three main methods:
#### `initialize`
```bash
curl -s http://127.0.0.1:8080/mcp -H 'content-type: application/json' -d '{
"jsonrpc":"2.0","id":1,"method":"initialize","params":{}
}'
```
Response:
```json
{
"jsonrpc":"2.0",
"id":1,
"result":{
"protocolVersion":"2024-11-05",
"capabilities":{"tools":{}},
"serverInfo":{"name":"rag-mcp","version":"1.0.0"}
}
}
```
#### `tools/list`
Lists all available tools (including `ask_agent`):
```bash
curl -s http://127.0.0.1:8080/mcp -H 'content-type: application/json' -d '{
"jsonrpc":"2.0","id":2,"method":"tools/list","params":{}
}'
```
#### `tools/call`
Calls a tool:
```bash
curl -s http://127.0.0.1:8080/mcp -H 'content-type: application/json' -d '{
"jsonrpc":"2.0","id":3,"method":"tools/call",
"params":{
"name":"ask_agent",
"arguments":{"question":"What is the SLA and how much does Pro cost?"}
}
}'
```
Response:
```json
{
"jsonrpc":"2.0",
"id":3,
"result":{
"content":[{"type":"text","text":"The SLA is 4 hours during business hours (UTC-3). The Pro Plan costs USD 49/month..."}]
}
}
```
### `GET /health`
Simple health check:
```bash
curl -s http://127.0.0.1:8080/health
# {"ok": true}
```
---
## Available Tools
| Tool | Description | InputSchema |
|---------------|------------------------------------------------------|------------------------------------------|
| `ask_agent` | LangGraph agent with RAG + function calling. | `{"question": "string"}` |
| `rag_search` | Searches facts in the base via RAG (LlamaIndex). | `{"query": "string"}` |
| `agora` | Returns current UTC datetime (ISO-8601). | `{}` |
| `calcular` | Evaluates safe arithmetic expression. | `{"expressao": "string"}` |
### Direct Usage Examples
```bash
# RAG direct
curl -s http://127.0.0.1:8080/mcp -H 'content-type: application/json' -d '{
"jsonrpc":"2.0","id":4,"method":"tools/call",
"params":{"name":"rag_search","arguments":{"query":"support SLA"}}
}'
# Datetime
curl -s http://127.0.0.1:8080/mcp -H 'content-type: application/json' -d '{
"jsonrpc":"2.0","id":5,"method":"tools/call",
"params":{"name":"agora","arguments":{}}
}'
# Calculation
curl -s http://127.0.0.1:8080/mcp -H 'content-type: application/json' -d '{
"jsonrpc":"2.0","id":6,"method":"tools/call",
"params":{"name":"calcular","arguments":{"expressao":"(2+3)*4"}}
}'
```
---
## Integration with Cursor / Zed / Other MCP Clients
Add to your editor config (e.g., `~/.cursor/settings.json`):
```json
{
"mcpServers": {
"rag-mcp": {
"url": "http://127.0.0.1:8080/mcp"
}
}
}
```
The client will:
1. Call `initialize`.
2. List tools (`tools/list`).
3. Use `ask_agent` or call tools directly as needed.
---
## How RAG Works
1. **Indexing**: On startup, `SimpleDirectoryReader` reads `data/` and `VectorStoreIndex` creates embeddings with `text-embedding-3-small`.
2. **Query**: `as_query_engine` retrieves the 3 most similar chunks and the LLM synthesizes the answer.
3. **Update**: To reindex, add/remove files in `data/` and restart the container.
---
## How the Agent Works (LangGraph)
- The `agent` node invokes the LLM with `bind_tools(TOOLS)`.
- `tools_condition` decides: if the model requests function calling, it goes to the `tools` node; otherwise, it ends.
- The `tools` node executes the tool and returns to `agent`, which generates the final response.
Flow:
```
START → agent → (tools?) → tools → agent → END
```
---
## Extending
### Adding a New Tool
In `app.py`, add:
```python
@tool
def my_tool(param1: str, param2: int = 0) -> str:
"""Clear description of what the tool does."""
# logic
return "result"
```
Then:
```python
TOOLS.append(my_tool)
```
Restart the server. The tool automatically appears in `tools/list`.
### Changing the Model
Change `LLM_MODEL` in `.env`:
```bash
LLM_MODEL=gpt-4o
```
Or use another provider (e.g., Anthropic, Groq) by replacing `ChatOpenAI` and embeddings in `app.py`.
### Changing the Vector Store
Replace `VectorStoreIndex` with a persistent store (Chroma, Pinecone, Weaviate, etc.):
```python
from llama_index.vector_stores.chroma import ChromaVectorStore
import chromadb
client = chromadb.PersistentClient(path="./chroma")
collection = client.get_or_create_collection("rag")
vector_store = ChromaVectorStore(chroma_collection=collection)
_index = VectorStoreIndex.from_documents(docs, vector_store=vector_store)
```
---
## Security and Best Practices
- **Do not expose** the server directly to the internet without authentication.
- Use **internal network** (Docker) if the MCP client is on the same host.
- Validate inputs in custom tools (especially if accessing DBs or external APIs).
- For production, add:
- Rate limiting.
- Structured logging.
- Metrics (Prometheus, OpenTelemetry).
---
## Troubleshooting
### `ModuleNotFoundError`
- Verify you installed `requirements.txt`.
- In Docker, run `docker compose build --no-cache`.
### Invalid `OPENAI_API_KEY`
- Confirm the key in `docker compose exec mcp env | grep OPENAI`.
- Test locally: `curl https://api.openai.com/v1/models -H "Authorization: Bearer $OPENAI_API_KEY"`.
### RAG Does Not Find Documents
- Verify `data/` has valid files (`.md`, `.txt`, etc.).
- Check logs: `docker compose logs mcp`.
- Reindex by restarting the container.
### Error in `calcular`
- The tool only accepts simple arithmetic expressions.
- Avoid variables, functions, or complex Python syntax.
---
## Files
### `app.py`
```python
"""MCP Server (JSON-RPC) + LangGraph + LlamaIndex RAG."""
from __future__ import annotations
import ast
import operator as op
import os
from datetime import datetime, timezone
from typing import Annotated, TypedDict
from flask import Flask, jsonify, request
from langchain_core.messages import AnyMessage, HumanMessage, SystemMessage
from langchain_core.tools import tool
from langchain_openai import ChatOpenAI
from langgraph.graph import START, StateGraph
from langgraph.graph.message import add_messages
from langgraph.prebuilt import ToolNode, tools_condition
from llama_index.core import Settings, SimpleDirectoryReader, VectorStoreIndex
from llama_index.embeddings.openai import OpenAIEmbedding
from llama_index.llms.openai import OpenAI as LlamaLLM
DATA_DIR = os.getenv("DATA_DIR", "./data")
MODEL = os.getenv("LLM_MODEL", "gpt-4o-mini")
# --- RAG: index ./data once on process startup ---
Settings.llm = LlamaLLM(model=MODEL)
Settings.embed_model = OpenAIEmbedding(model="text-embedding-3-small")
_index = VectorStoreIndex.from_documents(SimpleDirectoryReader(DATA_DIR).load_data())
_qe = _index.as_query_engine(similarity_top_k=3)
# --- Function calling: each @tool becomes JSON schema for LLM and MCP ---
@tool
def rag_search(query: str) -> str:
"""Search facts in local base via RAG (LlamaIndex). Use for policies, products, and docs."""
return str(_qe.query(query))
@tool
def agora() -> str:
"""Returns current UTC datetime (ISO-8601)."""
return datetime.now(timezone.utc).isoformat()
@tool
def calcular(expressao: str) -> str:
"""Evaluates safe arithmetic. Examples: (2+3)*4, 10/2, 2**8."""
ops = {
ast.Add: op.add, ast.Sub: op.sub, ast.Mult: op.mul, ast.Div: op.truediv,
ast.Mod: op.mod, ast.Pow: op.pow, ast.USub: op.neg,
}
def _eval(n):
if isinstance(n, ast.Expression):
return _eval(n.body)
if isinstance(n, ast.Constant) and isinstance(n.value, (int, float)):
return n.value
if isinstance(n, ast.BinOp) and type(n.op) in ops:
return ops[type(n.op)](_eval(n.left), _eval(n.right))
if isinstance(n, ast.UnaryOp) and type(n.op) in ops:
return ops[type(n.op)](_eval(n.operand))
raise ValueError("invalid expression")
return str(_eval(ast.parse(expressao, mode="eval")))
TOOLS = [rag_search, agora, calcular]
# --- LangGraph: agent ↔ tools until model stops requesting function calls ---
class State(TypedDict):
messages: Annotated[list[AnyMessage], add_messages]
llm = ChatOpenAI(model=MODEL, temperature=0).bind_tools(TOOLS)
def agent_node(state: State) -> dict:
sys = SystemMessage(content="MCP assistant. Use tools when needed. Respond in English.")
return {"messages": [llm.invoke([sys, *state["messages"]])]}
_g = StateGraph(State)
_g.add_node("agent", agent_node)
_g.add_node("tools", ToolNode(TOOLS))
_g.add_edge(START, "agent")
_g.add_conditional_edges("agent", tools_condition) # tools or END
_g.add_edge("tools", "agent")
GRAPH = _g.compile()
def _schema(t) -> dict:
"""Converts LangChain tool to MCP inputSchema."""
s = t.args_schema.model_json_schema() if t.args_schema else {"type": "object"}
s.pop("title", None)
return s
MCP_TOOLS = [
{"name": t.name, "description": t.description, "inputSchema": _schema(t)}
for t in TOOLS
] + [{
"name": "ask_agent",
"description": "LangGraph agent with RAG + function calling. Pass the user question.",
"inputSchema": {
"type": "object",
"properties": {"question": {"type": "string"}},
"required": ["question"],
},
}]
def _run(name: str, args: dict) -> str:
if name == "ask_agent":
out = GRAPH.invoke({"messages": [HumanMessage(content=args.get("question", ""))]})
return str(out["messages"][-1].content)
fn = {t.name: t for t in TOOLS}.get(name)
if not fn:
raise ValueError(f"unknown tool: {name}")
return str(fn.invoke(args or {}))
# --- Flask: HTTP transport for MCP (JSON-RPC 2.0) ---
app = Flask(__name__)
@app.post("/mcp")
def mcp():
body = request.get_json(force=True) or {}
method, rid, params = body.get("method"), body.get("id"), body.get("params") or {}
if method == "initialize":
return jsonify({"jsonrpc": "2.0", "id": rid, "result": {
"protocolVersion": "2024-11-05",
"capabilities": {"tools": {}},
"serverInfo": {"name": "rag-mcp", "version": "1.0.0"},
}})
if method == "tools/list":
return jsonify({"jsonrpc": "2.0", "id": rid, "result": {"tools": MCP_TOOLS}})
if method == "tools/call":
try:
text = _run(params.get("name"), params.get("arguments") or {})
result = {"content": [{"type": "text", "text": text}]}
except Exception as e:
result = {"content": [{"type": "text", "text": str(e)}], "isError": True}
return jsonify({"jsonrpc": "2.0", "id": rid, "result": result})
if method == "notifications/initialized" or rid is None:
return ("", 204)
return jsonify({"jsonrpc": "2.0", "id": rid, "error": {"code": -32601, "message": method}}), 400
@app.get("/health")
def health():
return {"ok": True}
if __name__ == "__main__":
app.run(host="0.0.0.0", port=int(os.getenv("PORT", 8080)))
```
### `requirements.txt`
```text
flask>=3.0
langgraph>=0.2
langchain-core>=0.3
langchain-openai>=0.2
llama-index>=0.12
llama-index-llms-openai
llama-index-embeddings-openai
```
### `Dockerfile`
```dockerfile
FROM python:3.12-slim
WORKDIR /app
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt
COPY app.py .
COPY data ./data
ENV PORT=8080 DATA_DIR=/app/data
EXPOSE 8080
CMD ["python", "app.py"]
```
### `docker-compose.yml`
```yaml
services:
mcp:
build: .
ports: ["8080:8080"]
env_file: .env
environment:
PORT: "8080"
DATA_DIR: /app/data
LLM_MODEL: gpt-4o-mini
volumes:
- ./data:/app/data:ro
restart: unless-stopped
healthcheck:
test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8080/health')"]
interval: 15s
retries: 5
```
### `data/kb.md`
```markdown
# Acme Corp
Support SLA: 4 hours during business hours (UTC-3).
Pro Plan costs USD 49/month and includes 10k RAG queries/day.
P1 incidents must be opened in #sre channel.
```
### `.env`
```bash
OPENAI_API_KEY=sk-...
LLM_MODEL=gpt-4o-mini
```
---
## License
MIT.This server cannot be deployed
Maintenance
ActivityMaintained
ResponsivenessNo issues