Skip to main content
Glama
rasvan-ghiliciu

local-executor-mcp

README.md
# local-executor-mcp

An MCP server that lets a frontier planner (Claude Code) **delegate mechanical
code-generation subtasks to a local LLM** served by `llama-swap`, to save
frontier-model tokens.

**Mental model:** Claude Code plans and verifies; the local model does the
high-volume, mechanical writing. The split saves tokens because the *expensive*
part — generating large amounts of output — happens on a free local model.

## Design choices (and why)

- **Generation-only, stateless.** The server never writes files and never runs
  tests. Claude Code keeps its native file + bash tools for that. This keeps the
  server portable (no repo coupling, no Windows path pain) and trivial to share.
- **Compact output.** The tool returns just the generated artifact plus a short
  usage footer. (Claude Code warns when an MCP tool output exceeds ~10k tokens,
  and re-reading huge blobs would eat the token savings.)
- **Model routing built in.** A `tier` argument routes between a `quality` model
  and a `fast` model.
- **Fallback is explicit.** On failure the tool returns `DELEGATION_FAILED` and
  tells the planner to do the task itself — so a weak local result never blocks
  you. This is your guardrail against compounding errors.
- **Metrics.** Every call is appended to `~/.local-executor-mcp/usage.jsonl`
  (model, tokens, latency, ok/fail) so you can quantify savings later.

## Model recommendations (from a typical llama-swap list)

| Use | Model | Why |
| --- | --- | --- |
| `quality` (default) | `qwen3.6-35b-a3b-q8_0` | 35B MoE, ~3B active → fast *and* capable, high-fidelity quant |
| `fast` | `qwen3.5-4b-q8_0` | small, quick, for trivial / high-volume tasks |
| A/B alternative for `quality` | `qwopus3.6-35b-a3b-v1-q8_0` | if code-tuned, may beat Qwen — measure it |

Avoid for code: translation models (`hy-mt2-*`, `tower-plus`), the vision model
(`minicpm-v`), and ultra-low quants (`iq2_xxs` / `iq3_xxs`).

## Build

Requires Node.js 18+.

```bash
npm install
npm run build
```

## Quick test (no Claude Code needed)

Confirm the server starts and your backend is reachable:

```bash
LOCAL_LLM_BASE_URL=http://10.0.0.42:8080/v1 npx @modelcontextprotocol/inspector node build/index.js
```

In the Inspector UI, call `list_local_models` first, then try `delegate_code`.

## Use it from Claude Code

### Solo (personal)

```bash
claude mcp add local-executor \
  --env LOCAL_LLM_BASE_URL=http://10.0.0.42:8080/v1 \
  --env LOCAL_LLM_MODEL_QUALITY=qwen3.6-35b-a3b-q8_0 \
  --env LOCAL_LLM_MODEL_FAST=qwen3.5-4b-q8_0 \
  -- node /absolute/path/to/local-executor-mcp/build/index.js
```

Restart your Claude Code session so the tools are discovered.

### Team rollout (the point of this project)

Use **project scope** so the config is checked into git and every teammate gets
it automatically. Copy `.mcp.json.example` to `.mcp.json` at the repo root and
commit it. The first time each teammate opens the repo, Claude Code asks them to
approve the server (workspace trust), then it just works.

- `.mcp.json.example` — assumes the package is published to npm and runs via
  `npx`. Lowest friction; teammates need nothing installed.
- `.mcp.json.vendored.example` — no publishing: drop the built server in
  `tools/local-executor-mcp/` inside the repo and point `node` at it.

The backend URL uses `${LOCAL_LLM_BASE_URL:-...default...}` expansion, so each
machine can override it via an environment variable without editing the file.

### Windows notes

- Node + `npx` work fine on Windows for stdio servers like this one.
- If a teammate hits issues with `npx`, pre-install once
  (`npm i -g @yourorg/local-executor-mcp`) and set `"command": "local-executor-mcp"`.
- Nothing here touches the filesystem, so there are no path-separator issues.

## How to actually drive it (prompt Claude Code like this)

> "Plan the module. For the boilerplate parts — the Odoo model fields, the XML
> views, and the security CSV — use the `delegate_code` tool with a precise spec
> and the relevant context, then review what comes back and run the tests."

Good delegation targets: boilerplate, scaffolding, CRUD, test stubs, format
conversions, Odoo `models`/`views`/`manifest`/`ir.model.access.csv`, bulk
rule-based edits. Keep logic, debugging, and security-sensitive code on Claude.

## Inspect your savings

```bash
cat ~/.local-executor-mcp/usage.jsonl
```

Each line records the model, token counts, latency and success. Aggregate it to
get "X tokens / Y% of a representative workload handled locally at Z% fallback
rate" — a quantified, defensible impact claim.

TDQS

A4.3/5.0

Scored across 2 tools

Disambiguation5/5

The two tools have completely distinct purposes: delegate_code handles code generation subtasks, while list_local_models checks available models. There is zero overlap or ambiguity.

Naming Consistency5/5

Both tool names follow a clean verb_noun pattern (delegate_code, list_local_models), using snake_case consistently. The naming is predictable and easy to understand.

Tool Count3/5

With only 2 tools, the server feels minimal. However, the narrow scope of 'local execution' justifies a small surface; it is borderline but not excessive.

Completeness4/5

For its stated purpose, the server covers the core delegation flow and model discovery. A status or cancel tool would be a minor enhancement, but no critical gaps exist.

Maintenance

ActivityInactive
ResponsivenessNo issues