Skip to main content
Glama
vvramana007

MCP-LLM-Gateways

by vvramana007
README.md
# MCP Servers & LLM Gateways

Reference implementation of the four assessment tasks: a validated MCP server,
an MCP security gateway, a streaming PII-redaction LLM gateway, and a
rate-limiting model-fallback router. Each task is a small, self-contained module
so it can be run and evaluated independently. Small mock services are included so
everything runs locally with no external credentials.

## Requirements

- Python 3.10+
- Dependencies listed in `requirements.txt`

## Setup

```bash
python -m venv .venv
# Windows:        .venv\Scripts\Activate.ps1
# macOS / Linux:  source .venv/bin/activate
pip install -r requirements.txt
```

## Project structure

| File | Responsibility |
| --- | --- |
| `task1_mcp_server.py` | Task 1 — MCP server (stdio) exposing `get_customer_record` and `trigger_refund` with strict schema validation. |
| `test_task1_client.py` | MCP client that launches the Task 1 server and exercises its tools. |
| `task2_gateway.py` | Task 2 — MCP security gateway: Bearer-token role auth and `admin_*` tool blocking. |
| `mock_downstream.py` | Mock downstream MCP server for Task 2. |
| `pii.py` | Streaming, chunk-boundary-safe PII redactor (email / SSN / card). |
| `task3_stream_gateway.py` | Task 3 — LLM gateway that redacts PII on a live response stream. |
| `mock_llm_stream.py` | Mock streaming LLM provider for Task 3. |
| `rate_limiter.py` | Task 4 — token-aware sliding-window rate limiter on SQLite. |
| `task4_model_router.py` | Task 4 — model router: rate limiting, primary→secondary failover, sanitized errors. |
| `mock_models.py` | Mock primary/secondary model providers for Task 4. |
| `run_task2_test.py`, `run_task3_test.py`, `run_task4_test.py` | Self-contained test runners (start servers, run scenarios, report PASS/FAIL). |
| `_testutil.py` | Shared helpers for the test runners. |

## Task 1 — Validated MCP server (stdio)

Two tools with strict input validation via Pydantic type annotations:
`customer_id` must match `CUST-XXXXX`, refund `amount` must be a positive finite
number, and `reason` must be 10–500 characters. Invalid arguments are rejected
before the tool body runs. `stdout` is reserved exclusively for JSON-RPC; all
logging is directed to `stderr`.

```bash
python test_task1_client.py
```

To inspect the raw protocol and confirm `stdout` carries only JSON-RPC:

```bash
printf '%s\n' \
'{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2024-11-05","capabilities":{},"clientInfo":{"name":"t","version":"1"}}}' \
'{"jsonrpc":"2.0","method":"notifications/initialized"}' \
'{"jsonrpc":"2.0","id":2,"method":"tools/list"}' \
| python task1_mcp_server.py 2>/dev/null
```

## Task 2 — MCP security gateway

Reads the `Authorization: Bearer <token>` header, resolves the role, and inspects
the JSON-RPC payload. `tools/list` and normal tool calls are forwarded unchanged;
a `tools/call` for any tool whose name starts with `admin_` is rejected locally
with JSON-RPC error `-32001` (never forwarded) unless the caller is `admin`.

```bash
python run_task2_test.py
```

Manual run:

```bash
python -m uvicorn mock_downstream:app --port 9001 &
python -m uvicorn task2_gateway:app --port 8002
# then, in another shell:
curl -X POST http://127.0.0.1:8002/mcp -H "Authorization: Bearer viewer-token" \
  -H "Content-Type: application/json" \
  -d '{"jsonrpc":"2.0","id":1,"method":"tools/call","params":{"name":"admin_reset_key"}}'
```

## Task 3 — Streaming PII redaction gateway

Relays a streamed LLM response and redacts emails, SSNs, and card numbers in real
time. A small overlap buffer ensures PII split across network chunks is still
detected, while memory stays bounded and time-to-first-token stays low.

```bash
python run_task3_test.py
```

Manual run:

```bash
python -m uvicorn mock_llm_stream:app --port 9002 &
python -m uvicorn task3_stream_gateway:app --port 8003
curl -N -X POST http://127.0.0.1:8003/v1/completions/stream \
  -H "Content-Type: application/json" \
  -d '{"mock_text":"Email alice@example.com and SSN 123-45-6789"}'
```

## Task 4 — Rate limiter & model fallback router

A token-aware sliding-window limiter (default 50,000 tokens/minute per API key)
backed by on-disk SQLite. The check-and-charge step runs inside a
`BEGIN IMMEDIATE` transaction and is additionally serialized by an in-process
async lock, so concurrent requests cannot exceed the limit. The router calls the
primary provider with a 3-second timeout; a `429` or timeout fails over to the
secondary. If both fail, a sanitized error is returned with no internal details.

```bash
python run_task4_test.py
```

The database defaults to `./data/rate_limits.db`. On a local disk this works as
is; on a network or synced folder, set `RATE_LIMIT_DB` to a local path (SQLite
WAL mode requires real disk file locking).

## Running the full test suite

```bash
python test_task1_client.py
python run_task2_test.py
python run_task3_test.py
python run_task4_test.py
```

## Notes on production hardening

- Replace the static token map with signed JWT validation or IdP introspection.
- Use TLS/mTLS between the gateway and upstream services; store secrets in a
  managed secret store.
- Replace character-based token estimation with the model's real tokenizer.
- For multiple replicas, move the rate limiter to a shared store such as Redis;
  SQLite is used here to satisfy the single-node on-disk requirement.
- Add structured logging, metrics, tracing, request IDs, and health endpoints.