Skip to main content
Glama
Youssef-AMARZOU

Local RAG MCP Server

README.md
# AI Orchestration MVP — Local-First RAG Pipeline

> **Bypass cloud API costs. Zero latency. Total privacy.**

A production-ready, modular RAG (Retrieval-Augmented Generation) pipeline that runs entirely on your machine. Built with custom TCP socket framing and the Model Context Protocol (MCP) for modular, maintainable AI architectures.

---

## Features

**Local-First Architecture**
Every component runs locally. No API keys. No recurring cloud bills. No data leaves your machine.

**Custom TCP Socket Framing**
Length-prefixed binary protocol for reliable message passing between components. Production-grade networking without HTTP overhead.

**Model Context Protocol (MCP)**
Modular tool-calling interface. Register tools, discover them dynamically, and execute them over TCP — decoupling the LLM from your business logic.

**RAG Pipeline**
- ChromaDB vector store for document embeddings
- Sentence-transformers for local embedding generation
- Ollama integration for local LLM inference (Llama 3.2, Mistral, etc.)
- Ingest text, files, or entire directories

**Streamlit Chat Interface**
Ready-to-use conversational UI. Start the server, ingest documents, and start asking questions.

---

## Architecture

```
┌──────────────────────────────────────────────────┐
│                Streamlit UI (port 8501)            │
│  ┌──────────────────────────────────────────────┐ │
│  │            Orchestrator Instance               │ │
│  │                                                │ │
│  │  ┌─────────────┐   ┌──────────────────────┐   │ │
│  │  │  TCP Server  │◄──┤  MCP Server           │   │ │
│  │  │  (port 5555) │   │  - query_rag tool     │   │ │
│  │  └──────┬──────┘   │  - ingest_file tool    │   │ │
│  │         │          │  - ingest_text tool    │   │ │
│  │         │          │  - status tool         │   │ │
│  │         │          └──────────┬─────────────┘   │ │
│  │         │                     │                  │ │
│  │         │          ┌──────────┴─────────────┐   │ │
│  │         │          │  RAG Pipeline            │   │ │
│  │         │          │  ┌────────────────────┐ │   │ │
│  │         └──────────┤  │  ChromaDB (vector) │ │   │ │
│  │                    │  └────────────────────┘ │   │ │
│  │                    │  ┌────────────────────┐ │   │ │
│  │                    │  │  SentenceTransformer│ │   │ │
│  │                    │  └────────────────────┘ │   │ │
│  │                    │  ┌────────────────────┐ │   │ │
│  │                    │  │  Ollama (local LLM)│ │   │ │
│  │                    │  └────────────────────┘ │   │ │
│  │                    └─────────────────────────┘   │ │
│  └──────────────────────────────────────────────────┘ │
└──────────────────────────────────────────────────────┘

External TCP clients can connect on port 5555
and use the MCP protocol to query tools remotely.
```

---

## Quick Start

### 1. Prerequisites

```bash
# Install Python dependencies
pip install -r requirements.txt

# Install and start Ollama (for local LLM)
# https://ollama.com
ollama pull llama3.2
ollama serve
```

### 2. Start the Orchestrator

```bash
python orchestrator.py
```

### 3. Launch the Chat UI

```bash
streamlit run streamlit_app.py
```

Open http://localhost:8501 in your browser.

### 4. Ingest Documents

In the sidebar:
- Paste text directly, or
- Provide a file path

Then ask questions in the chat.

---

## TCP Client Example

```python
from tcp_network_module import TCPClient
from mcp_server import MCPClient

def send(msg):
    client.send(msg)

client = TCPClient("127.0.0.1", 5555)
mcp_client = MCPClient(send)

def handle_action(msg):
    mcp_client.handle_response(msg)

client.on("mcp", handle_action)
client.connect()

# Query the RAG pipeline
client.send({"action": "call_tool", "tool": "query_rag",
             "arguments": {"question": "What is in my documents?"}})
```

---

## Project Structure

```
ai-orchestration-mvp/
├── tcp_network_module.py   # TCP framing (server + client)
├── mcp_server.py           # Model Context Protocol implementation
├── rag_pipeline.py         # ChromaDB + Ollama RAG pipeline
├── orchestrator.py         # Main coordinator (entry point)
├── streamlit_app.py        # Chat UI
├── requirements.txt        # Python dependencies
├── examples/
│   └── basic_usage.py      # TCP client example
└── README.md               # This file
```

---

## Why This Architecture?

| Cloud APIs | This Architecture |
|---|---|
| $50–$500+/month recurring | $0 recurring (one-time setup) |
| 500ms–3s latency | 50–200ms local latency |
| Data sent to third parties | Data never leaves your machine |
| Rate limits apply | No rate limits |
| Requires internet | Works fully offline |

---

## Requirements

- Python 3.10+
- Ollama (free, local LLM runner)
- 8GB+ RAM recommended

---

## License

MIT