localagents
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@localagentsUse the local agent to add a --json flag to the CLI and cover it in tests."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
localagents
Hand Claude Code's grunt work to a model running on your own hardware.
localagents is an MCP server that gives Claude Code a run_agent tool. Each call
starts a full headless Claude Code session — same tools, same CLAUDE.md, same
working tree — except its API traffic goes to a llama.cpp or vLLM server you run
instead of to Anthropic. Claude writes the brief, the local model does the work,
Claude reviews the result. Your Anthropic token budget goes on the parts that
need it.
A 27B Qwen on one GPU is perfectly capable of "add a CLI for this module and tests to match"; Opus is better spent on the design conversation than on watching pytest run. Two local agents in parallel can build the two halves of a package against a pinned interface.
Status: early. It works, I use it daily, and the interface will move. It targets llama.cpp and vLLM specifically; Ollama isn't a goal.
How it works
Claude Code (your session)
│ MCP: run_agent(task, model=...)
▼
localagents ── spawns ──▶ headless `claude` (Agent SDK)
│ │ ANTHROPIC_BASE_URL
│ ▼
└──── in-process shim ◀────┘ normalises requests, logs them,
│ translates backend errors
▼
llama-server / vllm (/v1/messages, on your machine or your LAN)Three things make this more than an environment variable:
A registry that is probed live.
models.yamllists where servers are and a menu of model names. What each server is actually serving right now, its real context window, and how many slots are busy are discovered on every call. You bring models up and down by hand — the server never launches anything — and when Claude needs a model that isn't running it asks you for it by name.A shim between Claude Code and the backend. Claude Code sends things local chat templates reject, and local servers fail in ways Claude Code doesn't recognise. The shim fixes both directions (details below) and writes a
requests.jsonlper job so you can see exactly what went over the wire.The same isolation model as Claude's own subagents. By default a job works in your tree, like the Agent tool does.
isolation: worktreegives it a fresh git worktree on alocal-agent/<job>branch, kept only if it changed something, with a diffstat in the job record so Claude can review it as a diff.
Related MCP server: Ollama MCP Server
Requirements
Python 3.12+ and uv
Claude Code. The Agent SDK bundles its own
claudebinary, so nothing else to install.A server that speaks Anthropic's
/v1/messages:llama.cpp
llama-server— start it with--jinja; add--slots --metricsto get occupancy and cache stats in the tool output.vLLM with
--enable-auto-tool-choice --tool-call-parser <parser>.
A model that can actually drive Claude Code: solid native tool calling and a context window of 128k or more per request. Qwen3.8-27B works well. Smaller windows work but compact constantly; see Context windows.
Install
git clone https://github.com/ccebelenski/localagents.git && cd localagents
uv tool install -e . # `localagents` on PATH; editable, so repo edits apply
cp models.example.yaml models.yaml # edit for your servers (gitignored)
claude mcp add --scope user local -- localagents --config "$PWD/models.yaml"User scope means every project gets the local server. It inherits the cwd of the
Claude Code session that launched it, so run_agent defaults to that project's
tree. A project can carry its own ./models.yaml to override the registry.
If you'd rather keep it to one project, put this in that project's .mcp.json:
{"mcpServers": {"local": {"command": "localagents", "args": ["--config", "/path/to/models.yaml"]}}}Restart Claude Code (or /mcp → reconnect) after adding it; MCP servers load at startup.
Using it
Claude picks it up like any tool. Ask for it by name and it'll do the right thing:
Use the local agent to add a
--jsonflag to the CLI and cover it in the tests.
What Claude does behind that: list_models to see what's up, run_agent(task=…)
which returns a job id, then wait_job / job_status / job_log until it's done,
then reads files_touched (or the worktree diff) and checks the work. Jobs that
outlive Claude Code's 2-minute tool timeout get backgrounded and picked up later;
you don't have to do anything.
If nothing suitable is running you'll be asked to start one:
qwen3.8-27bis not running anywhere. Ask the user to bring it up. Notes: default mid-size coder on llama.cpp; run with --reasoning on
Start it however you normally do, say "it's up", and Claude retries.
Tools
tool | what it does |
| endpoints with live health, served ids, context window, slot occupancy; the pool with |
| start a job: |
| follow and control jobs |
| what to tell the user to bring a pool model up |
| add to the pool from inside a session (written to |
| one-shot generation with no tools — summaries, drafts, classification |
Job records live in ~/.local/state/localagents/jobs/<job>/: transcript.txt
(what the agent said and did), events.jsonl (every SDK message), requests.jsonl
(every backend request with timing, size and usage), and requests_full.jsonl if
you turn on request dumping.
Config: models.yaml
Start from models.example.yaml. It's re-read on every call, so edits take effect
immediately, and the server never rewrites it — register_* write to a sidecar
models.local.yaml that is merged on top.
endpoints:
llamacpp:
base_url: http://127.0.0.1:8080
backend: llama.cpp
gpu-server:
base_url: http://gpu-server.lan:8000
backend: vllm
host: gpu-server
models:
qwen3.8-27b:
notes: default mid-size coder on llama.cpp; run with --reasoning on
deepseek-v4-flash:
host: gpu-server
notes: vllm needs --enable-auto-tool-choice --tool-call-parser deepseek_v3endpoints are places that serve
/v1/messages. What they serve is probed.models are just names. A name is fuzzy-matched against served ids (
qwen3.8-27bfindsunsloth/Qwen3.8-27B-GGUF:UD-Q4_K_XL), so an entry needs nothing butnotesand maybe ahostto relay when asking you to start it.served_name(exact id or glob),endpoint,context(fallback window) andbring_up(a start command) exist as overrides if you want them. Start commands go stale quickly; a name and a note usually age better.defaults cover the default model,
permission_mode(acceptEdits), allowed and disallowed tools (subagents can't spawn subagents), which Claude settings to load,max_turns,timeout_s, and a system-prompt suffix telling the agent it's a delegate and how to report back.
What the shim does
Both llama.cpp and vLLM speak /v1/messages natively, so pointing
ANTHROPIC_BASE_URL at them almost works. The shim closes the gaps:
System messages mid-conversation. Claude Code puts role: system entries
inside messages — the skills listing, a token-budget marker, and one more per
turn. Qwen's chat template refuses: "System message must be at the beginning".
The shim folds each one into the adjacent user message as a <system>…</system>
text block, in place. Hoisting them into the top-level system field instead
changes the start of the prompt every turn, which invalidates the server's
KV-cache prefix and re-evaluates the whole ~35k-token prompt each time (21–47 s
per turn on a 27B). Folding in place keeps the prompt append-only: f_sim_best
0.88–0.99 in llama-server's log, 2.5–14 s per turn.
Context overflow. Claude Code assumes a 200k window for any model it doesn't
recognise; with a smaller slot it runs into llama.cpp's
exceed_context_size_error, which it doesn't understand, and the job dies. See
the next section.
Everything the shim does is a no-op when it isn't needed, and every request is logged with its timing, message count, byte size and reported usage.
Context windows
Two layers keep a session inside the real window:
The probe reads it — llama.cpp
/propsn_ctx(per slot:-cdivided by--parallelwhen unified KV is off), vLLMmax_model_len— and the session getsCLAUDE_CODE_MAX_CONTEXT_TOKENS. Claude Code's own auto-compact then fires at the right point. Below 128k the output budget is also shrunk ton_ctx/8, because the compact threshold iswindow − max_outputand would otherwise sit at zero.If a request still overflows, the shim rewrites the backend's error into Anthropic's
prompt is too long: N tokens > M maximum, which Claude Code answers by compacting and retrying.
At 64k this works but thrashes: Claude Code's ~20k of fixed prompt and tool schemas, plus a ~7k-token compaction summary (~55 s on a 27B) and the files it re-attaches, refill the window within a few turns and its thrash guard ends the job. Give each slot 128k or more.
Backend notes
llama.cpp:
llama-server -hf <gguf> --jinja -fa on --slots --metrics, plus--reasoning onfor thinking models and--parallel Nfor concurrent jobs. With/slotson,list_modelsshows{total, busy, free}so Claude knows whether a second agent will run now or queue. With/metricson, each job records prompt tokens processed vs. cached, cache hit ratio, prompt and generation tok/s, and speculative-decode acceptance — the counters are server-wide, so overlapping jobs share the delta.vLLM:
vllm serve <model> --served-model-name <alias> --enable-auto-tool-choice --tool-call-parser <parser>. The context window comes frommax_model_lenin/v1/models; occupancy (requests_running/waiting,kv_cache_usage) and the per-job cache stats come from its always-on/metrics. vLLM reports context overflow on the Anthropic endpoint as HTTP 500internal_error; the shim translates that too. The session is started withCLAUDE_CODE_ATTRIBUTION_HEADER=0because the per-request attribution hash defeats prefix caching. Verified with DeepSeek-V4-Flash: thinking blocks replay with their signatures, prefix cache hits on every turn after the first.The first turn of a job costs about 20k tokens of prompt on a cold slot (system prompt plus tool schemas), ~10 s on a 27B. Everything after that is a cache hit plus the delta.
Development
uv sync --dev
uv run pytest -qSee CONTRIBUTING.md for the layout and how to test changes
against a real server.
License
MIT. See LICENSE.
Copyright © 2026 Chris Cebelenski
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Servers
- AlicenseCqualityDmaintenanceBridges Claude Desktop with local LLM instances running via llama-server, enabling full conversation support with complete parameter control and health monitoring. Allows users to chat with their local models directly through Claude Desktop with configurable sampling parameters.399Creative Commons Zero v1.0 Universal
- AlicenseNot gradedqualityDmaintenanceEnables Claude to delegate coding tasks to local Ollama models, reducing API token usage by up to 98.75% while leveraging local compute resources. Supports code generation, review, refactoring, and file analysis with Claude providing oversight and quality assurance.48824AGPL 3.0
- AlicenseNot gradedqualityDmaintenanceExposes local Ollama instances as tools for Claude Code, allowing users to offload code generation, text drafting, and embedding tasks to local GPUs. It supports multi-turn conversations and model management through the Model Context Protocol.MIT
- AlicenseNot gradedqualityCmaintenanceEnables Claude Code to delegate mechanical tasks (summaries, boilerplate, reformatting) to local models running in LM Studio.1MIT
Related MCP Connectors
Let ChatGPT, Claude & Cursor use your Mac: email, calendar, iMessage, Teams, files. Local, free.
Hosted MCP server connecting claude.ai, ChatGPT and other AI apps to your own computer
Real-time chat hub for AI agents — Claude Code, Cursor, Cline, Codex over MCP or REST.
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/ccebelenski/localagents'
If you have feedback or need assistance with the MCP directory API, please join our Discord server