rag-mcp
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@rag-mcpWhat are the Pro Plan features and pricing?"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
RAG-MCP Server
Model Context Protocol (MCP) server built with Flask, LangGraph, and LlamaIndex for RAG. Designed to be minimal, functional, and easy to extend.
Overview
MCP over HTTP:
POST /mcpwith JSON-RPC 2.0.LangGraph: orchestrates the agent with function calling.
LlamaIndex: indexes local documents and answers via RAG.
Docker:
docker compose up --buildand you are done.
Related MCP server: DocAgent-MCP
Architecture
MCP Client → Flask /mcp → LangGraph Agent → Tools
├─ rag_search (LlamaIndex RAG)
├─ agora (UTC datetime)
└─ calcular (arithmetic expression)The client lists tools (
tools/list) and calls them (tools/call).The
ask_agenttool triggers the LangGraph graph; the LLM decides when to use function calling.Other tools are called directly by MCP.
Prerequisites
Docker and Docker Compose (or Python 3.11+ locally).
OPENAI_API_KEY (for LLM and embeddings).
Project Structure
.
├── app.py # MCP server + LangGraph + LlamaIndex
├── data/ # RAG corpus (.md, .txt, .pdf, etc.)
│ └── kb.md
├── requirements.txt # Python dependencies
├── Dockerfile
├── docker-compose.yml
├── .env # OPENAI_API_KEY and variables
└── README.mdConfiguration
1. Clone / Create the Project
Create an empty directory and paste the project files (see Files section).
2. Environment Variables
Create .env in the root:
OPENAI_API_KEY=sk-...
LLM_MODEL=gpt-4o-miniSupported variables:
Variable | Default | Description |
| (required) | OpenAI API key. |
|
| Model for LLM and embeddings. |
|
| Directory with documents for RAG. |
|
| Server port. |
3. RAG Data
Place documents in data/ (e.g., kb.md, policies.md, manuals/). The server indexes everything on startup.
Minimal example (data/kb.md):
# Acme Corp
Support SLA: 4 hours during business hours (UTC-3).
Pro Plan costs USD 49/month and includes 10k RAG queries/day.
P1 incidents must be opened in #sre channel.Running
Docker (recommended)
docker compose up --buildThe server runs at http://127.0.0.1:8080.
Local (without Docker)
pip install -r requirements.txt
export OPENAI_API_KEY=sk-...
export DATA_DIR=./data
python app.pyEndpoints
POST /mcp (JSON-RPC 2.0)
MCP uses three main methods:
initialize
curl -s http://127.0.0.1:8080/mcp -H 'content-type: application/json' -d '{
"jsonrpc":"2.0","id":1,"method":"initialize","params":{}
}'Response:
{
"jsonrpc":"2.0",
"id":1,
"result":{
"protocolVersion":"2024-11-05",
"capabilities":{"tools":{}},
"serverInfo":{"name":"rag-mcp","version":"1.0.0"}
}
}tools/list
Lists all available tools (including ask_agent):
curl -s http://127.0.0.1:8080/mcp -H 'content-type: application/json' -d '{
"jsonrpc":"2.0","id":2,"method":"tools/list","params":{}
}'tools/call
Calls a tool:
curl -s http://127.0.0.1:8080/mcp -H 'content-type: application/json' -d '{
"jsonrpc":"2.0","id":3,"method":"tools/call",
"params":{
"name":"ask_agent",
"arguments":{"question":"What is the SLA and how much does Pro cost?"}
}
}'Response:
{
"jsonrpc":"2.0",
"id":3,
"result":{
"content":[{"type":"text","text":"The SLA is 4 hours during business hours (UTC-3). The Pro Plan costs USD 49/month..."}]
}
}GET /health
Simple health check:
curl -s http://127.0.0.1:8080/health
# {"ok": true}Available Tools
Tool | Description | InputSchema |
| LangGraph agent with RAG + function calling. |
|
| Searches facts in the base via RAG (LlamaIndex). |
|
| Returns current UTC datetime (ISO-8601). |
|
| Evaluates safe arithmetic expression. |
|
Direct Usage Examples
# RAG direct
curl -s http://127.0.0.1:8080/mcp -H 'content-type: application/json' -d '{
"jsonrpc":"2.0","id":4,"method":"tools/call",
"params":{"name":"rag_search","arguments":{"query":"support SLA"}}
}'
# Datetime
curl -s http://127.0.0.1:8080/mcp -H 'content-type: application/json' -d '{
"jsonrpc":"2.0","id":5,"method":"tools/call",
"params":{"name":"agora","arguments":{}}
}'
# Calculation
curl -s http://127.0.0.1:8080/mcp -H 'content-type: application/json' -d '{
"jsonrpc":"2.0","id":6,"method":"tools/call",
"params":{"name":"calcular","arguments":{"expressao":"(2+3)*4"}}
}'Integration with Cursor / Zed / Other MCP Clients
Add to your editor config (e.g., ~/.cursor/settings.json):
{
"mcpServers": {
"rag-mcp": {
"url": "http://127.0.0.1:8080/mcp"
}
}
}The client will:
Call
initialize.List tools (
tools/list).Use
ask_agentor call tools directly as needed.
How RAG Works
Indexing: On startup,
SimpleDirectoryReaderreadsdata/andVectorStoreIndexcreates embeddings withtext-embedding-3-small.Query:
as_query_engineretrieves the 3 most similar chunks and the LLM synthesizes the answer.Update: To reindex, add/remove files in
data/and restart the container.
How the Agent Works (LangGraph)
The
agentnode invokes the LLM withbind_tools(TOOLS).tools_conditiondecides: if the model requests function calling, it goes to thetoolsnode; otherwise, it ends.The
toolsnode executes the tool and returns toagent, which generates the final response.
Flow:
START → agent → (tools?) → tools → agent → ENDExtending
Adding a New Tool
In app.py, add:
@tool
def my_tool(param1: str, param2: int = 0) -> str:
"""Clear description of what the tool does."""
# logic
return "result"Then:
TOOLS.append(my_tool)Restart the server. The tool automatically appears in tools/list.
Changing the Model
Change LLM_MODEL in .env:
LLM_MODEL=gpt-4oOr use another provider (e.g., Anthropic, Groq) by replacing ChatOpenAI and embeddings in app.py.
Changing the Vector Store
Replace VectorStoreIndex with a persistent store (Chroma, Pinecone, Weaviate, etc.):
from llama_index.vector_stores.chroma import ChromaVectorStore
import chromadb
client = chromadb.PersistentClient(path="./chroma")
collection = client.get_or_create_collection("rag")
vector_store = ChromaVectorStore(chroma_collection=collection)
_index = VectorStoreIndex.from_documents(docs, vector_store=vector_store)Security and Best Practices
Do not expose the server directly to the internet without authentication.
Use internal network (Docker) if the MCP client is on the same host.
Validate inputs in custom tools (especially if accessing DBs or external APIs).
For production, add:
Rate limiting.
Structured logging.
Metrics (Prometheus, OpenTelemetry).
Troubleshooting
ModuleNotFoundError
Verify you installed
requirements.txt.In Docker, run
docker compose build --no-cache.
Invalid OPENAI_API_KEY
Confirm the key in
docker compose exec mcp env | grep OPENAI.Test locally:
curl https://api.openai.com/v1/models -H "Authorization: Bearer $OPENAI_API_KEY".
RAG Does Not Find Documents
Verify
data/has valid files (.md,.txt, etc.).Check logs:
docker compose logs mcp.Reindex by restarting the container.
Error in calcular
The tool only accepts simple arithmetic expressions.
Avoid variables, functions, or complex Python syntax.
Files
app.py
"""MCP Server (JSON-RPC) + LangGraph + LlamaIndex RAG."""
from __future__ import annotations
import ast
import operator as op
import os
from datetime import datetime, timezone
from typing import Annotated, TypedDict
from flask import Flask, jsonify, request
from langchain_core.messages import AnyMessage, HumanMessage, SystemMessage
from langchain_core.tools import tool
from langchain_openai import ChatOpenAI
from langgraph.graph import START, StateGraph
from langgraph.graph.message import add_messages
from langgraph.prebuilt import ToolNode, tools_condition
from llama_index.core import Settings, SimpleDirectoryReader, VectorStoreIndex
from llama_index.embeddings.openai import OpenAIEmbedding
from llama_index.llms.openai import OpenAI as LlamaLLM
DATA_DIR = os.getenv("DATA_DIR", "./data")
MODEL = os.getenv("LLM_MODEL", "gpt-4o-mini")
# --- RAG: index ./data once on process startup ---
Settings.llm = LlamaLLM(model=MODEL)
Settings.embed_model = OpenAIEmbedding(model="text-embedding-3-small")
_index = VectorStoreIndex.from_documents(SimpleDirectoryReader(DATA_DIR).load_data())
_qe = _index.as_query_engine(similarity_top_k=3)
# --- Function calling: each @tool becomes JSON schema for LLM and MCP ---
@tool
def rag_search(query: str) -> str:
"""Search facts in local base via RAG (LlamaIndex). Use for policies, products, and docs."""
return str(_qe.query(query))
@tool
def agora() -> str:
"""Returns current UTC datetime (ISO-8601)."""
return datetime.now(timezone.utc).isoformat()
@tool
def calcular(expressao: str) -> str:
"""Evaluates safe arithmetic. Examples: (2+3)*4, 10/2, 2**8."""
ops = {
ast.Add: op.add, ast.Sub: op.sub, ast.Mult: op.mul, ast.Div: op.truediv,
ast.Mod: op.mod, ast.Pow: op.pow, ast.USub: op.neg,
}
def _eval(n):
if isinstance(n, ast.Expression):
return _eval(n.body)
if isinstance(n, ast.Constant) and isinstance(n.value, (int, float)):
return n.value
if isinstance(n, ast.BinOp) and type(n.op) in ops:
return ops[type(n.op)](_eval(n.left), _eval(n.right))
if isinstance(n, ast.UnaryOp) and type(n.op) in ops:
return ops[type(n.op)](_eval(n.operand))
raise ValueError("invalid expression")
return str(_eval(ast.parse(expressao, mode="eval")))
TOOLS = [rag_search, agora, calcular]
# --- LangGraph: agent ↔ tools until model stops requesting function calls ---
class State(TypedDict):
messages: Annotated[list[AnyMessage], add_messages]
llm = ChatOpenAI(model=MODEL, temperature=0).bind_tools(TOOLS)
def agent_node(state: State) -> dict:
sys = SystemMessage(content="MCP assistant. Use tools when needed. Respond in English.")
return {"messages": [llm.invoke([sys, *state["messages"]])]}
_g = StateGraph(State)
_g.add_node("agent", agent_node)
_g.add_node("tools", ToolNode(TOOLS))
_g.add_edge(START, "agent")
_g.add_conditional_edges("agent", tools_condition) # tools or END
_g.add_edge("tools", "agent")
GRAPH = _g.compile()
def _schema(t) -> dict:
"""Converts LangChain tool to MCP inputSchema."""
s = t.args_schema.model_json_schema() if t.args_schema else {"type": "object"}
s.pop("title", None)
return s
MCP_TOOLS = [
{"name": t.name, "description": t.description, "inputSchema": _schema(t)}
for t in TOOLS
] + [{
"name": "ask_agent",
"description": "LangGraph agent with RAG + function calling. Pass the user question.",
"inputSchema": {
"type": "object",
"properties": {"question": {"type": "string"}},
"required": ["question"],
},
}]
def _run(name: str, args: dict) -> str:
if name == "ask_agent":
out = GRAPH.invoke({"messages": [HumanMessage(content=args.get("question", ""))]})
return str(out["messages"][-1].content)
fn = {t.name: t for t in TOOLS}.get(name)
if not fn:
raise ValueError(f"unknown tool: {name}")
return str(fn.invoke(args or {}))
# --- Flask: HTTP transport for MCP (JSON-RPC 2.0) ---
app = Flask(__name__)
@app.post("/mcp")
def mcp():
body = request.get_json(force=True) or {}
method, rid, params = body.get("method"), body.get("id"), body.get("params") or {}
if method == "initialize":
return jsonify({"jsonrpc": "2.0", "id": rid, "result": {
"protocolVersion": "2024-11-05",
"capabilities": {"tools": {}},
"serverInfo": {"name": "rag-mcp", "version": "1.0.0"},
}})
if method == "tools/list":
return jsonify({"jsonrpc": "2.0", "id": rid, "result": {"tools": MCP_TOOLS}})
if method == "tools/call":
try:
text = _run(params.get("name"), params.get("arguments") or {})
result = {"content": [{"type": "text", "text": text}]}
except Exception as e:
result = {"content": [{"type": "text", "text": str(e)}], "isError": True}
return jsonify({"jsonrpc": "2.0", "id": rid, "result": result})
if method == "notifications/initialized" or rid is None:
return ("", 204)
return jsonify({"jsonrpc": "2.0", "id": rid, "error": {"code": -32601, "message": method}}), 400
@app.get("/health")
def health():
return {"ok": True}
if __name__ == "__main__":
app.run(host="0.0.0.0", port=int(os.getenv("PORT", 8080)))requirements.txt
flask>=3.0
langgraph>=0.2
langchain-core>=0.3
langchain-openai>=0.2
llama-index>=0.12
llama-index-llms-openai
llama-index-embeddings-openaiDockerfile
FROM python:3.12-slim
WORKDIR /app
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt
COPY app.py .
COPY data ./data
ENV PORT=8080 DATA_DIR=/app/data
EXPOSE 8080
CMD ["python", "app.py"]docker-compose.yml
services:
mcp:
build: .
ports: ["8080:8080"]
env_file: .env
environment:
PORT: "8080"
DATA_DIR: /app/data
LLM_MODEL: gpt-4o-mini
volumes:
- ./data:/app/data:ro
restart: unless-stopped
healthcheck:
test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8080/health')"]
interval: 15s
retries: 5data/kb.md
# Acme Corp
Support SLA: 4 hours during business hours (UTC-3).
Pro Plan costs USD 49/month and includes 10k RAG queries/day.
P1 incidents must be opened in #sre channel..env
OPENAI_API_KEY=sk-...
LLM_MODEL=gpt-4o-miniLicense
MIT.
This server cannot be deployed
Maintenance
Related MCP Connectors
- docs2mcpOAuthcom.docs2mcp
Query your own PDFs and documents from any MCP client. Every answer cites the page it came from.
Agent-native MCP server over the public saagarpatel.dev corpus. Read-only, stateless.
The CustomGPT.ai MCP server is a fully managed, RAG-powered endpoint that connects large language models with private knowledge bases and external data sources. It provides tools for retrieval-augmented generation queries (send_message), data ingestion (upload_file), and source listing, enabling AI agents to query private documents like PDFs with high accuracy and real-time citations.
Query any docs site via MCP. Submit a URL, ask questions, get cited answers.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceEnables grounding AI responses in a local document corpus by exposing MCP tools to list, search, and summarize documents, and generating answers using OpenAI.-
- FlicenseNot gradedqualityCmaintenanceEnables local document question-answering and retrieval via MCP, supporting multi-turn conversation, intent recognition, and tools for document search, Q&A, and summarization.5-
- FlicenseNot gradedqualityBmaintenanceEnables document Q&A, summarization, keyword extraction, and Wikipedia lookup through MCP tools, using RAG with FAISS and Ollama.-
- FlicenseNot gradedqualityCmaintenanceEnables querying internal documents via a FastAPI REST API and MCP server, using retrieval-augmented generation and an agentic loop that can invoke tools like document search and calculations.-