Skip to main content
Glama
README.md
# mlserve

Self-hosted, low-latency ML inference endpoints for a spare GPU machine.
One Python process exposes state-of-the-art open models through **two
authenticated surfaces** that share the same warm model pool:

- a **REST API** (`/v1/embeddings`, `/v1/summarize`, model management)
- an **MCP server** (`/mcp`, streamable HTTP) for agentic applications

Designed for small-GPU hardware (tested target: 6 GB VRAM / 16 GB RAM) on
Windows, macOS and Linux.

## Features

- **Embeddings** via sentence-transformers (MiniLM, BGE, ...) on CUDA/MPS/CPU
- **Summarization** via Ollama-served quantized LLMs (llama3.2, gemma3,
  deepseek-r1, ...) with prompt-guided instructions
- **Never dormant** — models preload at startup, a keep-warm loop pings idle
  models, and Ollama runs with `keep_alive: -1` so weights stay in VRAM
- **Async & parallel** — non-blocking endpoints with per-model concurrency
  limits sized for small GPUs
- **Authenticated** — API-key middleware (`X-API-Key` or `Bearer`) covering
  REST *and* MCP
- **Observable** — structured JSON logs with per-request trace ids, optional
  [Langfuse](https://langfuse.com) tracing (free tier)
- **Config-driven** — add/remove models by editing `config/models.yaml`;
  add new runtimes by registering a backend class

## Architecture

```
                ┌─────────────────────────────────────────────┐
   REST clients │  FastAPI app (single process, single worker)│
  ──────────────┤                                             │
   X-API-Key    │  RequestContext ─ ApiKey middleware         │
                │        │                                    │
   MCP clients  │   ┌────┴─────┐        ┌──────────────────┐  │
  ──────────────┤   │ /v1/*    │        │ /mcp (FastMCP)   │  │
   (agents)     │   └────┬─────┘        └───────┬──────────┘  │
                │        └───────┬──────────────┘             │
                │                ▼                            │
                │   EmbeddingService · SummarizationService   │
                │        (logging · tracing · limits)         │
                │                ▼                            │
                │        ModelRegistry (+ keep-warm loop)     │
                │      semaphores · lifecycle · lazy loads    │
                │        ▼                      ▼             │
                │  sentence-transformers      Ollama          │
                │  (CUDA / MPS / CPU)     (quantized LLMs)    │
                └─────────────────────────────────────────────┘
```

## Quickstart

```bash
git clone <this-repo> && cd ds-ml-mcp-service
python -m venv .venv && source .venv/bin/activate   # Windows: .venv\Scripts\activate
pip install -e ".[embeddings]"

# Ollama powers summarization: https://ollama.com/download
ollama pull llama3.2:3b

cp .env.example .env      # set MLSERVE_API_KEYS to a long random value
mlserve                   # serves on http://127.0.0.1:8000
```

Full platform guides (Windows/CUDA specifics, Apple Silicon, remote access):
**[docs/SETUP.md](docs/SETUP.md)**.

## REST API

```bash
export KEY="<your api key>"

# Embeddings — single text or batch
curl -s http://127.0.0.1:8000/v1/embeddings \
  -H "X-API-Key: $KEY" -H "Content-Type: application/json" \
  -d '{"input": ["first text", "second text"]}'

# Summarization — instruction-guided
curl -s http://127.0.0.1:8000/v1/summarize \
  -H "X-API-Key: $KEY" -H "Content-Type: application/json" \
  -d '{"text": "<long text>", "instruction": "Three bullet points.", "max_words": 80}'

# Ops
curl -s http://127.0.0.1:8000/health                      # unauthenticated
curl -s -H "X-API-Key: $KEY" http://127.0.0.1:8000/v1/models
curl -s -X POST -H "X-API-Key: $KEY" http://127.0.0.1:8000/v1/models/embed-bge-small/load
```

Interactive OpenAPI docs at `http://127.0.0.1:8000/docs`.

Every response carries a `request_id` (also in the `X-Request-ID` header);
grep it in `logs/mlserve.jsonl` to see the full trace of that call.

## MCP

The MCP endpoint lives at `http://<host>:8000/mcp` (streamable HTTP) and
exposes three tools: `list_models`, `embed_texts`, `summarize_text`.

Client configuration (any MCP client that supports HTTP servers + headers):

```json
{
  "mcpServers": {
    "mlserve": {
      "type": "http",
      "url": "http://127.0.0.1:8000/mcp",
      "headers": { "X-API-Key": "<your api key>" }
    }
  }
}
```

## Configuration

Process settings come from env vars / `.env` (see
[.env.example](.env.example)); the model catalog lives in
[config/models.yaml](config/models.yaml). Highlights:

| Setting | Default | Purpose |
| --- | --- | --- |
| `MLSERVE_API_KEYS` | — (required) | comma-separated accepted API keys |
| `MLSERVE_HOST` / `MLSERVE_PORT` | `127.0.0.1` / `8000` | bind address |
| `MLSERVE_KEEPALIVE_INTERVAL_SECONDS` | `240` | idle-model warm-ping cadence |
| `MLSERVE_LOG_PAYLOADS` | `true` | log truncated input/output previews |
| `MLSERVE_MAX_BATCH_SIZE` | `64` | embedding batch cap |
| `MLSERVE_CORS_ORIGINS` | off | browser origins allowed to call the API |
| `LANGFUSE_PUBLIC_KEY` / `LANGFUSE_SECRET_KEY` | off | enable Langfuse tracing |

## Adding / removing models

Edit `config/models.yaml` and restart — or load/unload at runtime via
`POST /v1/models/{name}/load|unload`. New runtimes (llama.cpp, vLLM, remote
APIs) are one subclass away; see
[docs/ARCHITECTURE.md](docs/ARCHITECTURE.md#adding-a-backend).

## Tests

```bash
pip install -e ".[dev]"
pytest                    # fast suite, fake backends, no downloads
pytest -m integration     # real-model tests (downloads MiniLM, ~90 MB)
```

Latest local run: [docs/TEST_REPORT.md](docs/TEST_REPORT.md).

## Deployment

Docker + Google Cloud (and generic VM) instructions:
[docs/DEPLOYMENT.md](docs/DEPLOYMENT.md).

## Docs

- [docs/SETUP.md](docs/SETUP.md) — install on Windows (GPU laptop), macOS, Linux
- [docs/ARCHITECTURE.md](docs/ARCHITECTURE.md) — components, request flow, extension points
- [docs/DEPLOYMENT.md](docs/DEPLOYMENT.md) — Docker, docker-compose, GCP
- [docs/TEST_REPORT.md](docs/TEST_REPORT.md) — latest verified test run