Local RAG MCP Server
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Local RAG MCP ServerIngest the file 'report.pdf' and ask: what are the main findings?"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
AI Orchestration MVP — Local-First RAG Pipeline
Bypass cloud API costs. Zero latency. Total privacy.
A production-ready, modular RAG (Retrieval-Augmented Generation) pipeline that runs entirely on your machine. Built with custom TCP socket framing and the Model Context Protocol (MCP) for modular, maintainable AI architectures.
Features
Local-First Architecture Every component runs locally. No API keys. No recurring cloud bills. No data leaves your machine.
Custom TCP Socket Framing Length-prefixed binary protocol for reliable message passing between components. Production-grade networking without HTTP overhead.
Model Context Protocol (MCP) Modular tool-calling interface. Register tools, discover them dynamically, and execute them over TCP — decoupling the LLM from your business logic.
RAG Pipeline
ChromaDB vector store for document embeddings
Sentence-transformers for local embedding generation
Ollama integration for local LLM inference (Llama 3.2, Mistral, etc.)
Ingest text, files, or entire directories
Streamlit Chat Interface Ready-to-use conversational UI. Start the server, ingest documents, and start asking questions.
Related MCP server: RAG MCP Server
Architecture
┌──────────────────────────────────────────────────┐
│ Streamlit UI (port 8501) │
│ ┌──────────────────────────────────────────────┐ │
│ │ Orchestrator Instance │ │
│ │ │ │
│ │ ┌─────────────┐ ┌──────────────────────┐ │ │
│ │ │ TCP Server │◄──┤ MCP Server │ │ │
│ │ │ (port 5555) │ │ - query_rag tool │ │ │
│ │ └──────┬──────┘ │ - ingest_file tool │ │ │
│ │ │ │ - ingest_text tool │ │ │
│ │ │ │ - status tool │ │ │
│ │ │ └──────────┬─────────────┘ │ │
│ │ │ │ │ │
│ │ │ ┌──────────┴─────────────┐ │ │
│ │ │ │ RAG Pipeline │ │ │
│ │ │ │ ┌────────────────────┐ │ │ │
│ │ └──────────┤ │ ChromaDB (vector) │ │ │ │
│ │ │ └────────────────────┘ │ │ │
│ │ │ ┌────────────────────┐ │ │ │
│ │ │ │ SentenceTransformer│ │ │ │
│ │ │ └────────────────────┘ │ │ │
│ │ │ ┌────────────────────┐ │ │ │
│ │ │ │ Ollama (local LLM)│ │ │ │
│ │ │ └────────────────────┘ │ │ │
│ │ └─────────────────────────┘ │ │
│ └──────────────────────────────────────────────────┘ │
└──────────────────────────────────────────────────────┘
External TCP clients can connect on port 5555
and use the MCP protocol to query tools remotely.Quick Start
1. Prerequisites
# Install Python dependencies
pip install -r requirements.txt
# Install and start Ollama (for local LLM)
# https://ollama.com
ollama pull llama3.2
ollama serve2. Start the Orchestrator
python orchestrator.py3. Launch the Chat UI
streamlit run streamlit_app.pyOpen http://localhost:8501 in your browser.
4. Ingest Documents
In the sidebar:
Paste text directly, or
Provide a file path
Then ask questions in the chat.
TCP Client Example
from tcp_network_module import TCPClient
from mcp_server import MCPClient
def send(msg):
client.send(msg)
client = TCPClient("127.0.0.1", 5555)
mcp_client = MCPClient(send)
def handle_action(msg):
mcp_client.handle_response(msg)
client.on("mcp", handle_action)
client.connect()
# Query the RAG pipeline
client.send({"action": "call_tool", "tool": "query_rag",
"arguments": {"question": "What is in my documents?"}})Project Structure
ai-orchestration-mvp/
├── tcp_network_module.py # TCP framing (server + client)
├── mcp_server.py # Model Context Protocol implementation
├── rag_pipeline.py # ChromaDB + Ollama RAG pipeline
├── orchestrator.py # Main coordinator (entry point)
├── streamlit_app.py # Chat UI
├── requirements.txt # Python dependencies
├── examples/
│ └── basic_usage.py # TCP client example
└── README.md # This fileWhy This Architecture?
Cloud APIs | This Architecture |
$50–$500+/month recurring | $0 recurring (one-time setup) |
500ms–3s latency | 50–200ms local latency |
Data sent to third parties | Data never leaves your machine |
Rate limits apply | No rate limits |
Requires internet | Works fully offline |
Requirements
Python 3.10+
Ollama (free, local LLM runner)
8GB+ RAM recommended
License
MIT
This server cannot be deployed
Maintenance
Related MCP Connectors
Nifty's MCP server — exposes tasks, projects, messages, and files as tools for AI agents.
MCP server for progressive tool usage at any scale (see https://klavis.ai)
MCP server exposing the Backtest360 engine API as tools for AI agents.
Related MCP Servers
- FlicenseNot gradedqualityCmaintenanceEnables local document question-answering and retrieval via MCP, supporting multi-turn conversation, intent recognition, and tools for document search, Q&A, and summarization.5-
- FlicenseNot gradedqualityBmaintenanceExposes a Retrieval-Augmented Generation pipeline as MCP tools, allowing users to index documents and query them through any MCP-compatible client like Claude or IDEs.-
- AlicenseNot gradedqualityCmaintenanceMCP server for a self-hosted RAG system that enables AI tools to search and retrieve grounded answers from locally ingested documents via MCP tools, with local embeddings and no API key required.MIT
- AlicenseNot gradedqualityBmaintenanceEnables fully local retrieval over a personal document corpus via hybrid search, cross-encoder reranking, RAPTOR summaries, and knowledge graph queries, served to AI agents over MCP.MIT