Skip to main content
Glama
README.md
# llmstudio-mcp

[Polski](README.pl.md)

MCP server exposing a local [LM Studio](https://lmstudio.ai) instance to LLM agents — list / load / unload models, run chat and text completions, generate embeddings, and reach any other LM Studio REST endpoint via a raw escape hatch.

## Table of contents

- [Tools](#tools)
- [Environment variables](#environment-variables)
- [Prerequisites](#prerequisites)
- [Wiring it up](#wiring-it-up)
- [Local run](#local-run)

## Tools

| Tool | Parameters | Description |
|------|-----------|-------------|
| `list_models` | `type: "all" \| "llm" \| "embedding" \| "vlm" = "all"` | Lists downloaded models with metadata (type, state, context length, quant) |
| `get_model` | `model_key: str` | Info about a single model |
| `get_loaded_models` | — | Models currently resident in memory |
| `load_model` | `model_key, context_length?, flash_attention?, offload_kv_cache_to_gpu?, num_experts?, eval_batch_size?` | Loads a model into memory |
| `unload_model` | `instance_id: str` | Frees a model from memory |
| `chat` | `model, messages, temperature?, max_tokens?, top_p?, top_k?, stop?, stream?, tools?` | OpenAI-compatible chat completion (`/v1/chat/completions`) |
| `complete` | `model, prompt, temperature?, max_tokens?, top_p?, top_k?, stop?, stream?` | OpenAI-compatible text completion (`/v1/completions`) |
| `embed` | `model, input, normalize=false` | OpenAI-compatible embeddings (`/v1/embeddings`) |
| `raw_request` | `method, path, json_body?` | Escape hatch for any LM Studio REST endpoint |

## Environment variables

| Variable | Required | Default | Description |
|---------|----------|---------|-------------|
| `LMSTUDIO_MCP_BASE_URL` | no | `http://localhost:1234` | LM Studio server URL |
| `LMSTUDIO_MCP_API_KEY` | no | empty | API token (only if you enabled token auth in LM Studio's Developer tab) |
| `LMSTUDIO_MCP_TIMEOUT` | no | `60.0` | HTTP timeout in seconds; raise for slow models or large generations |

## Prerequisites

LM Studio running with its API server enabled — either the desktop app with
the Developer tab toggle on, or `lms server start` from the CLI. At least one
model downloaded (`lms get <model>`).

## Wiring it up

Only requirement: `uv` (https://docs.astral.sh/uv/). Nothing else to install.

### Claude Code

```
claude mcp add llmstudio-mcp -- uvx --from git+https://github.com/dam2452/llmstudio-mcp.git llmstudio-mcp
```

If you enabled API token auth in LM Studio:

```
claude mcp add llmstudio-mcp -e LMSTUDIO_MCP_API_KEY=<token> -- uvx --from git+https://github.com/dam2452/llmstudio-mcp.git llmstudio-mcp
```

### Claude Desktop / other MCP client

```json
{
  "mcpServers": {
    "llmstudio-mcp": {
      "command": "uvx",
      "args": ["--from", "git+https://github.com/dam2452/llmstudio-mcp.git", "llmstudio-mcp"]
    }
  }
}
```

After pushing a new version: `uv cache clean` and restart the client.

## Local run

```
uv run --directory . llmstudio-mcp
```

Tests (manual):

```
uv run --directory . --with pytest pytest test/
```

TDQS

A4/5.0

Scored across 9 tools

Disambiguation5/5

Each tool has a distinct purpose: chat, complete, and embed for inference; list_models, get_model, load_model, unload_model for model management; raw_request as an escape hatch. No functional overlap.

Naming Consistency5/5

Consistent verb_noun pattern (e.g., get_loaded_models, list_models, load_model) with simple action verbs for inference (chat, complete, embed). All lowercase with underscores; clear and predictable.

Tool Count5/5

9 tools cover the core server management lifecycle: model discovery, loading/unloading, inference, and a raw_request for extensibility. Neither too few nor too many for the domain.

Completeness4/5

Covers all essential operations: model listing, loading, inference (chat/complete/embed), and unloading. Missing model download/delete is a minor gap, but raw_request can handle edge cases.

Maintenance

ActivityStale
ResponsivenessNo issues