Skip to main content
Glama
Akashat-02-Dev

FastMCP Enterprise Server

README.md
# FastMCP Enterprise Server

![Python 3.11+](https://img.shields.io/badge/python-3.11+-blue.svg)
![FastAPI](https://img.shields.io/badge/FastAPI-005571?style=flat&logo=fastapi)
![PostgreSQL](https://img.shields.io/badge/PostgreSQL-316192?style=flat&logo=postgresql&logoColor=white)
![Redis](https://img.shields.io/badge/redis-%23DD0031.svg?style=flat&logo=redis&logoColor=white)
![Kubernetes](https://img.shields.io/badge/kubernetes-%23326ce5.svg?style=flat&logo=kubernetes&logoColor=white)

An enterprise-grade Model Context Protocol (MCP) server built with **FastMCP** and **FastAPI**. Designed for high concurrency, robust security, and scalable AI integrations, this server provides tools for semantic document search, SSRF-protected external context fetching, and cached profile retrieval.

## ๐Ÿš€ Tech Stack

- **Core Framework**: [FastAPI](https://fastapi.tiangolo.com/) & [FastMCP](https://github.com/jlowin/fastmcp)
- **Database**: PostgreSQL with `pgvector` (Async via `asyncpg` & SQLAlchemy)
- **Caching & Rate Limiting**: Redis, `fastapi-limiter`
- **Authentication**: Firebase Admin SDK (JWT Validation)
- **Testing**: `pytest-asyncio`, `httpx`, `locust` (Load Testing)
- **Infrastructure**: Docker, Kubernetes

## ๐Ÿ—๏ธ Architecture & Features

### 1. Robust Security Model
- **Firebase Authentication**: Custom FastAPI middleware validating Firebase JWT tokens for the `/sse` and `/messages` endpoints.
- **Confused Deputy Mitigation**: Scope validation (`db.read`) enforced at the middleware level before tool execution.
- **SSRF Protection**: Outbound external context fetches route through a hardened `safe_fetch` protocol blocklist, preventing internal metadata enumeration (e.g., AWS/GCP `169.254.169.254`).
- **Schema Drift Protection**: Dynamic execution-time SHA-256 hash validation ensures AI tools (like `search_docs`) have not had their schemas silently modified or manipulated.

### 2. High-Concurrency Optimizations
- **SSE Rate Limiting**: AI client loops are strictly rate-limited (e.g., 50 requests/min) on a per-Firebase-UID basis via Redis and `fastapi-limiter`.
- **HNSW Vector Search**: Document embeddings are indexed in PostgreSQL using Hierarchical Navigable Small World (`hnsw`) graphs for sub-millisecond semantic search retrieval.
- **Read-Through Caching**: Heavy SQL queries (like user profiles) are cached in Redis to offload PostgreSQL overhead.

### 3. Integrated Tools
- `search_docs`: Semantic search against the `Document` pgvector embeddings. Supports pagination (`limit`/`offset`).
- `fetch_external_context`: Safely retrieves and truncates external URL payloads.
- `get_user_profile`: Retrieves cached RBAC and department metadata for an authenticated user.

## ๐Ÿ› ๏ธ Setup & Local Development

### Prerequisites
- Python 3.11+
- PostgreSQL (with `pgvector` extension installed)
- Redis Server
- Docker & Kubernetes (for deployment)

### 1. Environment Configuration
Clone the repository and install the dependencies:
```bash
git clone <repository-url>
cd mcp-server
python -m venv venv
source venv/bin/activate  # On Windows: venv\Scripts\activate
pip install -r requirements.txt
```

Set your environment variables (or create a `.env` file):
```bash
DATABASE_URL=postgresql+asyncpg://user:password@localhost:5432/mcpdb
REDIS_URL=redis://localhost:6379/0
```

### 2. Running the Server
Launch the application using Uvicorn:
```bash
uvicorn src.main:app --host 0.0.0.0 --port 8000 --reload
```

## ๐Ÿงช Testing

The repository maintains a formal test suite and a load-testing infrastructure.

### Unit & Integration Testing
Run the `pytest` suite to validate health checks, SSE connections, and Auth middleware.
```bash
pytest tests/test_api.py -v
```

### Load Testing
Simulate high-concurrency AI clients maintaining SSE connections and dispatching `search_docs` queries.
```bash
locust -f locustfile.py
```
*Navigate to `http://localhost:8089` to start the Locust UI.*

## ๐Ÿšข Kubernetes Deployment

The application is containerized and ready for Kubernetes orchestration.

1. **Build the Docker Image**:
```bash
docker build -t fastmcp-app:latest .
```

2. **Apply Manifests**:
```bash
kubectl apply -f k8s/deployment.yaml
kubectl apply -f k8s/service.yaml
```

The Deployment is configured with:
- `3` Replicas
- Resource requests/limits configured for memory and CPU.
- Readiness and Liveness probes pointing to `/health`.

## ๐Ÿ“œ Workflow Protocol (FastMCP SSE)

This server natively exposes the MCP SSE protocol via FastAPI routes:
1. **Connect**: Client connects via `GET /sse` (requires `Authorization: Bearer <token>`).
2. **Handshake**: Server opens the stream and emits an `endpoint` event containing the POST URL.
3. **RPC Calls**: Client posts JSON-RPC payloads to `POST /messages` to invoke registered tools.