local-llm-mcp
Allows sending prompts and files to a local Ollama server via an OpenAI-compatible API, enabling private and bulk text processing tasks such as summarizing, classifying, translating, or reformatting.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@local-llm-mcpSummarize each file in ~/Documents/meeting-notes and write summaries to ~/Documents/summaries"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
local-llm-mcp
An MCP server that lets Claude Code (or any MCP client) hand work to a local LLM behind an OpenAI-compatible API (llama.cpp, llama-swap, mlx-lm, vLLM, LM Studio, Ollama, ...).
Keep using your frontier model for thinking, and send the local model the things it is good for:
Private data that must not leave the machine. Pass file paths; this server reads the files and sends them only to the local LLM. The client model sees the local model's answer, or with
output_pathonly a notice that the answer was written to a file.Bulk, mechanical text work such as summarizing, classifying, translating, or reformatting many files.
Install
Requires uv and a local LLM server with an OpenAI-compatible API.
Claude Code (plugin)
claude plugin marketplace add koyakimu/local-llm-mcp
claude plugin install local-llm@koyakimuWhen the plugin is enabled, Claude Code asks for the base URL (default http://127.0.0.1:8080/v1), the model name
(empty = first model from /v1/models) and an optional API key, which is kept in secure storage.
Change them later in /config.
Codex (plugin)
codex plugin marketplace add koyakimu/local-llm-mcpThen run codex plugin add local-llm@koyakimu (or install it from /plugins). The plugin's server uses
http://127.0.0.1:8080/v1 and the first model from /v1/models; the Codex plugin format has no user settings.
To pick the model, point to another server, or set a timeout, register the server manually instead (below).
Do not add a [mcp_servers.local-llm] table to ~/.codex/config.toml for the plugin's server: Codex reads it as a
separate server without a command and fails to load the config (invalid transport).
Any MCP client (manual)
# Claude Code
claude mcp add --scope user local-llm -e LOCAL_LLM_MODEL=my-model \
-- uvx --from git+https://github.com/koyakimu/local-llm-mcp local-llm-mcp
# Codex
codex mcp add local-llm --env LOCAL_LLM_MODEL=my-model \
-- uvx --from git+https://github.com/koyakimu/local-llm-mcp local-llm-mcpCodex stops a tool call after 60 seconds by default (per its docs), and the first call to a local model can take
longer (model loading, long prompts). For a manually registered server, raise it in ~/.codex/config.toml:
[mcp_servers.local-llm]
tool_timeout_sec = 600Related MCP server: MCP LLM Integration Server
Use
Ask things like:
"Summarize each file in ~/Documents/minutes/ in three lines with local_llm and write the summaries to ~/Documents/summaries/"
"Use local_llm to group the errors in ~/logs/app.log by type"
Tool
local_llm(prompt, files=None, output_path=None, system=None, max_tokens=4096, thinking=False)
Argument | Meaning |
| Instruction for the local model. It sees nothing else, so make it self-contained |
| UTF-8 text files appended after the prompt |
| Write the answer to this new file and return only a notice. Existing files are never overwritten |
| Optional system prompt |
| Maximum tokens to generate |
| Sends |
Configuration
Variable | Default | |
|
| OpenAI-compatible base URL |
| first model from | Model name to request |
| none | Sent as |
|
| Limit for prompt + files |
|
| Seconds; the first call may include model loading |
|
| Colon-separated directories that files may be read from and written to |
Safety
The paths come from a model that may be reading untrusted documents, so the server limits what it touches:
Files are read and written only under
LOCAL_LLM_ALLOWED_ROOTS, after resolving symlinks.Any path with a hidden component (
~/.ssh,.env,~/.config, ...) is refused, so keys and settings cannot be pulled into an answer and dotfiles cannot be overwritten.File sizes are checked before reading;
output_pathnever overwrites an existing file.
The answer itself can still contain parts of the input. Use output_path when even the answer should stay local.
Development
uv run --group dev pytest # no LLM needed; the API is mocked
uv run scripts/check_live.py # against a running local LLMLicense
MIT
This server cannot be deployed
Maintenance
Related MCP Connectors
Text generation over MCP: prose, emails, blog outlines, SQL, humanizing, text diffs, fake data.
OCR, transcription, file extraction, and image generation for AI agents via MCP.
MCP server that lets AI assistants use all OneSchema features exposed via the public API.
- QuallaaOAuthcom.quallaa
Talk to your public-facing AI from any MCP client — Claude, ChatGPT, Cursor, Cline, Windsurf.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceIntegrates local language models (like Qwen3-8B) with MCP clients, providing tools for chat, code analysis, text generation, translation, and content summarization using your own hardware.-
- FlicenseBqualityDmaintenanceEnables integration of local LLM capabilities with MCP-compatible clients like Claude Desktop, Continue.dev, and Cline. Provides tools for processing text prompts through local language models using a customizable inference function.21-
- AlicenseAqualityCmaintenanceProxies LLM completion requests to OpenAI-compatible providers via MCP tools.2MIT
- AlicenseAqualityBmaintenanceEnables any MCP client to perform image understanding and OCR via any OpenAI-compatible vision-language model. Supports local, private inference without images leaving the machine.234 npmMIT