Skip to main content
Glama

local-llm-mcp

An MCP server that lets Claude Code (or any MCP client) hand work to a local LLM behind an OpenAI-compatible API (llama.cpp, llama-swap, mlx-lm, vLLM, LM Studio, Ollama, ...).

Keep using your frontier model for thinking, and send the local model the things it is good for:

  • Private data that must not leave the machine. Pass file paths; this server reads the files and sends them only to the local LLM. The client model sees the local model's answer, or with output_path only a notice that the answer was written to a file.

  • Bulk, mechanical text work such as summarizing, classifying, translating, or reformatting many files.

Install

Requires uv and a local LLM server with an OpenAI-compatible API.

Claude Code (plugin)

claude plugin marketplace add koyakimu/local-llm-mcp
claude plugin install local-llm@koyakimu

When the plugin is enabled, Claude Code asks for the base URL (default http://127.0.0.1:8080/v1), the model name (empty = first model from /v1/models) and an optional API key, which is kept in secure storage. Change them later in /config.

Codex (plugin)

codex plugin marketplace add koyakimu/local-llm-mcp

Then run codex plugin add local-llm@koyakimu (or install it from /plugins). The plugin's server uses http://127.0.0.1:8080/v1 and the first model from /v1/models; the Codex plugin format has no user settings. To pick the model, point to another server, or set a timeout, register the server manually instead (below).

Do not add a [mcp_servers.local-llm] table to ~/.codex/config.toml for the plugin's server: Codex reads it as a separate server without a command and fails to load the config (invalid transport).

Any MCP client (manual)

# Claude Code
claude mcp add --scope user local-llm -e LOCAL_LLM_MODEL=my-model \
  -- uvx --from git+https://github.com/koyakimu/local-llm-mcp local-llm-mcp
# Codex
codex mcp add local-llm --env LOCAL_LLM_MODEL=my-model \
  -- uvx --from git+https://github.com/koyakimu/local-llm-mcp local-llm-mcp

Codex stops a tool call after 60 seconds by default (per its docs), and the first call to a local model can take longer (model loading, long prompts). For a manually registered server, raise it in ~/.codex/config.toml:

[mcp_servers.local-llm]
tool_timeout_sec = 600

Related MCP server: MCP LLM Integration Server

Use

Ask things like:

  • "Summarize each file in ~/Documents/minutes/ in three lines with local_llm and write the summaries to ~/Documents/summaries/"

  • "Use local_llm to group the errors in ~/logs/app.log by type"

Tool

local_llm(prompt, files=None, output_path=None, system=None, max_tokens=4096, thinking=False)

Argument

Meaning

prompt

Instruction for the local model. It sees nothing else, so make it self-contained

files

UTF-8 text files appended after the prompt

output_path

Write the answer to this new file and return only a notice. Existing files are never overwritten

system

Optional system prompt

max_tokens

Maximum tokens to generate

thinking

Sends chat_template_kwargs: {"enable_thinking": ...} (llama.cpp, vLLM, mlx-lm), plus reasoning_effort: "none" when off (Splash). Off by default because thinking makes every call much slower

Configuration

Variable

Default

LOCAL_LLM_BASE_URL

http://127.0.0.1:8080/v1

OpenAI-compatible base URL

LOCAL_LLM_MODEL

first model from /v1/models

Model name to request

LOCAL_LLM_API_KEY

none

Sent as Authorization: Bearer ... if set

LOCAL_LLM_MAX_INPUT_CHARS

60000

Limit for prompt + files

LOCAL_LLM_TIMEOUT

600

Seconds; the first call may include model loading

LOCAL_LLM_ALLOWED_ROOTS

~, /tmp, $TMPDIR, and the macOS per-user temp dir

Colon-separated directories that files may be read from and written to

Safety

The paths come from a model that may be reading untrusted documents, so the server limits what it touches:

  • Files are read and written only under LOCAL_LLM_ALLOWED_ROOTS, after resolving symlinks.

  • Any path with a hidden component (~/.ssh, .env, ~/.config, ...) is refused, so keys and settings cannot be pulled into an answer and dotfiles cannot be overwritten.

  • File sizes are checked before reading; output_path never overwrites an existing file.

The answer itself can still contain parts of the input. Use output_path when even the answer should stay local.

Development

uv run --group dev pytest               # no LLM needed; the API is mocked
uv run scripts/check_live.py            # against a running local LLM

License

MIT

Available Tools

1 tool
local_llmA

Run a task on a local LLM on this machine. Nothing is sent to the cloud.

Use it for (1) private data that must not leave the machine: pass file paths in files instead of reading them yourself; this server reads them and only the local model's answer comes back to you; and (2) bulk, mechanical text work (summarizing, classifying, translating, reformatting) where a smaller model is good enough. Do not use it for tasks that need strong reasoning or careful code changes.

Files must be under the home directory or /tmp and must not be hidden (no path component starting with ".").

Args: prompt: The instruction for the local model. Write it self-contained; the model sees nothing else. files: Text files to include after the prompt (UTF-8). Total input is capped (~60k characters). output_path: If set, the answer is written to this new file and only a short notice is returned, so even the answer stays out of the conversation. Existing files are never overwritten. system: Optional system prompt. max_tokens: Maximum tokens to generate. thinking: Enable the model's thinking mode, if it has one (slower, sometimes more accurate).

ParametersJSON Schema
NameRequiredDescriptionDefault
filesNo
promptYes
systemNo
thinkingNo
max_tokensNo
output_pathNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and delivers: privacy guarantee, path restrictions (under home or /tmp, no hidden components), an input size cap (~60k chars), non-overwrite safety for output_path, and the fact that output_path suppresses the answer from the conversation. These are substantive behavioral facts an agent could not infer.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose, then use cases, then constraints, then a clean Args block. Dense but every sentence adds a fact that affects invocation (privacy, path rules, overwrite behavior, thinking tradeoff). No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described. Combined with full parameter semantics, safety constraints, and usage routing, an agent has everything needed to call this correctly and predict the side effects.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate, and it documents all six parameters: prompt is self-contained instruction, files are UTF-8 text appended after the prompt, output_path redirects the answer and never overwrites, system is an optional system prompt, max_tokens and thinking are explained including the slower/sometimes-more-accurate tradeoff. No parameter is left to guesswork.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Run a task on a local LLM on this machine') and immediately distinguishes the scope with 'Nothing is sent to the cloud', which is the defining characteristic versus a remote LLM tool. An agent knows exactly what this does and what makes it different.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly enumerates two use cases (private data that must not leave the machine; bulk mechanical text work) and explicitly states when not to use it ('tasks that need strong reasoning or careful code changes'). It even explains the alternative pattern for private data (pass file paths rather than reading them yourself). This is textbook when/when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool updatev0.1.1
    • First observedlocal_llm

TDQS

A4.8/5.0

Scored across 1 tool

Disambiguation5/5

There is only one tool, so there is no possibility of confusion or misselection between tools. Its purpose (run a task on a local LLM) is unambiguous and clearly stated.

Naming Consistency4/5

With a single tool named `local_llm`, there is no convention to break, so consistency is trivially satisfied. It uses a clear snake_case noun rather than a verb_noun pattern, which is a minor deviation but perfectly readable.

Tool Count4/5

A single, argument-rich tool is a reasonable fit for a narrowly scoped domain like invoking a local LLM, and nothing obvious is bolted on unnecessarily. It sits at the low end of the typical range, but the domain genuinely appears to need only one entry point.

Completeness4/5

The tool's arguments cover the core lifecycle for its purpose: prompt, context files, output redirection, system prompt, token limit, and thinking mode. Minor gaps exist (e.g. no way to list or select available local models or check readiness), but agents can work around these.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers