Skip to main content
Glama

openai-mcp — a local subagent for Devin

An MCP server that lets Devin offload low-complexity, monotonous work to a cheap local model (Ollama by default, or any OpenAI-compatible endpoint) instead of burning frontier tokens on it.

The headline feature is run_local_agent: the small model runs its own tool loop server-side (read/search files, run whitelisted commands, stage writes). File contents and intermediate steps never enter Devin's context — only a compact final answer comes back. Think of it as a local, free-tier version of Devin's cloud subagents.

Why it actually saves tokens

  • Server-side file I/O. summarize, extract, map_files, and run_local_agent take paths/globs and read files inside the server process. A 200KB file summarized this way costs Devin ~500 tokens instead of ~60k.

  • Compact results. Tool results, diffs, and agent steps are truncated/capped before returning.

  • Visible savings. Every response reports usage and delegated_bytes — bytes processed that never entered Devin's context. get_usage_stats totals it for the session.

Related MCP server: Local Worker MCP

Requirements

  • Node >= 22.18 (runs TypeScript natively — no build step). Verified on Node 26.

  • Ollama (ollama serve) for the default provider, or any OpenAI-compatible /v1 endpoint.

  • rg (ripgrep) is optional — search_files uses it when installed and falls back to a built-in walker (gitignore-aware, hidden/binary skipped) when it isn't.

  • nvtop (optional) for the best GPU detection — all vendors + per-process usage; falls back to nvidia-smi/rocm-smi.

Setup

1. Server config

npm install           # required once — the server refuses to start without node_modules
npm run configure     # interactive wizard: runs `npm install` itself if node_modules is missing,
                      # probes ollama, writes ~/.config/openai-mcp/config.json, prefills from an
                      # existing config (update mode), then covers steps 2 & 3 below — installs the
                      # server into ~/.config/devin/mcp_config.json AND merges recommended
                      # permissions into ~/.config/devin/config.json (backs both up first)
# or: npm install && cp config.example.json ~/.config/openai-mcp/config.json  (edit by hand)

File roots are auto-detected. Effective roots = allowed_roots (config) ∪ workspace roots the MCP client advertises (the spec's roots capability) ∪ the directory Devin spawned the server from (= your project dir, controlled by trust_cwd: true). So file tools work out of the box inside the project, and allowed_roots is for extra directories outside it. deny_globs blocks secrets everywhere. get_server_info shows each root with its source.

2. Register with Devin

npm run configure does this for you — shown here for manual setup.

Devin settings → MCP → Add custom MCP (or edit ~/.config/devin/mcp_config.json):

{
  "mcpServers": {
    "local_llm": {
      "command": "node",
      "args": ["/path/to/openai-mcp/src/index.ts"],
      "env": {
        "LOCAL_LLM_CONFIG": "~/.config/openai-mcp/config.json",
        "LOCAL_LLM_RUN_COMMAND": "1",
        "LOCAL_LLM_DYNAMIC_TOOLS": "1"
      }
    }
  }
}

Provider entries also accept think (bool or "low"/"medium"/"high"), options (provider-native request overrides), temperature, num_ctx, and keep_alive — defaults applied to every call; per-request args override them. The configure wizard asks about think since thinking-capable models can eat the whole max_tokens budget.

CLI equivalent: devin mcp add -s user local_llm -- node /path/to/openai-mcp/src/index.ts

Omit the two LOCAL_LLM_* env flags to disable shell access and dynamic tools entirely.

npm run configure merges these for you (with a backup) — shown here for manual setup.

In ~/.config/devin/config.json:

{
  "permissions": {
    "allow": ["mcp__local_llm__chat", "mcp__local_llm__complete", "mcp__local_llm__summarize",
              "mcp__local_llm__extract", "mcp__local_llm__classify", "mcp__local_llm__map_files",
              "mcp__local_llm__run_local_agent", "mcp__local_llm__list_models",
              "mcp__local_llm__get_system_resources", "mcp__local_llm__list_staged",
              "mcp__local_llm__get_diff", "mcp__local_llm__discard_write",
              "mcp__local_llm__get_usage_stats", "mcp__local_llm__list_dynamic_tools",
              "mcp__local_llm__get_command_whitelist", "mcp__local_llm__get_job",
              "mcp__local_llm__list_jobs", "mcp__local_llm__control_job",
              "mcp__local_llm__get_server_info", "mcp__local_llm__propose_write",
              "mcp__local_llm__propose_edit",
              "mcp__local_llm__recommend_model", "mcp__local_llm__list_commits",
              "mcp__local_llm__revert_write", "mcp__local_llm__search_models",
              "mcp__local_llm__list_model_tags", "mcp__local_llm__wait"],
    "ask":   ["mcp__local_llm__commit_write", "mcp__local_llm__set_write_mode",
              "mcp__local_llm__pull_model", "mcp__local_llm__download_model",
              "mcp__local_llm__delete_model", "mcp__local_llm__prune_models",
              "mcp__local_llm__register_tool", "mcp__local_llm__unregister_tool",
              "mcp__local_llm__update_command_whitelist", "mcp__local_llm__unload_model",
              "mcp__local_llm__uncommit_write", "mcp__local_llm__verify_staged"]
  }
}

ask = you get an approve button before writes hit disk, models download, or new tools register.

Tools

Tool

What it does

run_local_agent

Local subagent. Small model runs its own tool loop; returns final answer + staged op ids. verify:"self" adds a critique pass. Sync calls wait up to agent_sync_grace_ms (default 45s), then return a job_id instead of losing the result; run_async:true skips the wait and returns a job immediately (poll get_job).

chat / complete

Raw completions passthrough.

summarize

Summarize files/globs server-side; only summaries return.

extract

Structured JSON extraction, optional schema validation + retry.

classify

Label text/files into candidate labels.

map_files

Apply one instruction across a glob — report results or stage_writes for review.

propose_write

Local model rewrites/creates one file → staged diff.

propose_edit

Patch-style edit: the model returns <<<<<<< SEARCH / ======= / >>>>>>> REPLACE hunks instead of the whole file; the server applies them to current content and stages the result. Cheaper + safer than propose_write for localized changes — a non-matching hunk is a clean error, not a truncated file.

list_staged / get_diff / commit_write / discard_write

Staged-write lifecycle. Nothing the local model writes reaches disk without commit_write. Commit refuses if the file drifted since staging (force=true overrides) — and entirely while write_mode is propose. commit_write requires path + summary args so the approval prompt shows what is being written, not just an opaque op id; the path is verified against the staged op.

set_write_mode

propose (default): server stages diffs only; Devin applies via its own edit tools (get_diff include_content for full content). write: commit_write writes to disk. persist:true saves to config.

verify_staged

Lint/typecheck a staged op without touching the real file: content is materialized to a temp file in the same directory (imports resolve), {file} in argv substitutes the temp path, runs under the command whitelist, then cleans up. restore_paths snapshots/restores side files the checker may rewrite. The local agent gets this too — run_command on the real path only sees old disk content.

list_commits / revert_write / uncommit_write

Undo: revert_write(commit_id) stages a revert op (reviewed path); uncommit_write(commit_id) restores pre-commit content immediately in write mode (deletes files the commit created; drift-checked).

list_models

Models per provider (sizes, capabilities, last-modified).

get_system_resources

RAM + VRAM + loaded models (local ollama only). ollama_gpu reports which GPU(s) the server is pinned to (env vars) or observed using (runner process device fds).

recommend_model

Rank installed models vs detected memory budget (pinned/observed GPU > discrete GPU > RAM) with three-way fit verdicts (yes/marginal/no). apply:true writes default_model to config and applies live; model:"x:y" + apply:true sets an explicit default.

search_models

Search the public ollama registry for models — name, description, capability chips (tools/thinking/vision), parameter sizes, installed flag.

list_model_tags

Pullable tags for a registry model (gemma4 → e4b, 12b, 26b, …); include_sizes:true fetches real download sizes via the manifest API.

pull_model

Download a model on ollama — background job, poll get_job.

control_job

Steer a running job: pause/resume/cancel/inject (inject = operator instruction the local agent sees next step).

wait

Sleep N seconds (max 600); with job_id returns early when the job finishes — poll long pulls/agent runs without busy-looping get_job.

download_model

Download from Hugging Face via hf CLI (for llama.cpp/TabbyAPI/vLLM servers). Background job.

unload_model

Free a model's VRAM.

delete_model

Permanently delete a model (irreversible). Refuses to delete the provider default or a loaded model without force:true.

prune_models

Bulk-delete old models. dry_run:true (default) reports candidates + reclaimable bytes; always keeps keep[], the provider default, loaded models, keep_recent newest, max_age_days recent pulls, and unused_days models used recently (per-model last-used is persisted to model-usage.json next to the config — models with no record are kept).

register_tool / unregister_tool / list_dynamic_tools

Give the local agent new tools at runtime (shell templates or js snippets).

get_command_whitelist / update_command_whitelist

Manage what run_command may execute.

get_usage_stats / get_server_info

Session savings + effective config.

Slash commands (MCP prompts): /mcp__local_llm__delegate <task> [files] and /mcp__local_llm__cheap_summary <path> [focus].

Generation params

Every generation tool accepts these on top of model/provider/temperature/num_ctx/max_tokens:

Param

Effect

think

Chain-of-thought control: true/false, or "low"/"medium"/"high" on models with effort levels. Maps to ollama's think; on openai-compatible servers false → chat_template_kwargs.enable_thinking (vLLM/llama.cpp convention), a level → reasoning_effort.

top_p, seed, stop

Standard sampling controls.

options

Provider-native override bag — merged into ollama options (top_k, repeat_penalty, …) or the openai-compat request body. Wins over the named params.

Set defaults per provider in config: "think": false, "options": {"top_k": 20}. Request args override them.

Thinking models eat max_tokens. On thinking-capable models (gemma4, qwen3, gpt-oss) a tight max_tokens can be consumed entirely by reasoning, returning empty text. Fix per call with think:false, or globally per provider with "think": false. chat/complete return the model's thinking trace when the provider reports it, so the budget usage is observable. (run_local_agent accepts think/options too — disabling think leaves more context and steps for tool calls.)

Progress

Two free/cheap progress channels:

  • MCP notifications/progress: run_local_agent reports each agent step and map_files reports per-file completion — if the client attaches a progressToken. These go to client UI plumbing, not into the model context — zero token cost.

  • Job polling: every run_local_agent call runs under a job. run_async:true returns the job_id immediately; sync calls return the inline result if they finish within agent_sync_grace_ms, otherwise the job_id. get_job shows live progress + the final result. Polling costs a few tokens per call, but only when Devin chooses to check. (pull_model/download_model are always jobs.)

Safety model

  • Staged writes + write modes: local-model output lands in a review queue, never directly on disk. Default write_mode: "propose" means the server never writes — Devin reviews get_diff and applies through its own edit tools, so changes go through Devin's native file-review flow (at the cost of the diff/content tokens). set_write_mode("write") (or write_mode: "write" in config) lets commit_write apply server-side — those bypass Devin's edit-review UI but stay git-visible. The required path/summary args keep the ask-permission prompt meaningful, and commit_write/set_write_mode belong in ask permissions either way.

  • run_command: whitelist-only, executed via execFile argv — no shell, so ;, &&, |, > are literal arguments and can't inject. allow_args constrains subcommands (e.g. git → read-only verbs).

  • File scope: union of configured allowed_roots + MCP-advertised workspace roots + cwd (disable with trust_cwd: false), minus deny_globs for secrets.

  • VRAM: ollama calls set num_ctx explicitly (default 16384) — an unbounded context blew up the KV-cache allocation on a 16GB GPU. get_system_resources + unload_model help manage memory.

  • Honesty: the instructions tell Devin not to delegate anything where a plausible-but-wrong answer is worse than no answer. 8B models produce garbage sometimes — verify:"self" and staged diffs are the mitigations, not a guarantee.

Provider config

Two types:

  • ollama → native /api (chat with num_ctx/keep_alive/tools, pull, ps, unload, delete). Local or remote.

  • openai → generic /v1 (chat completions, list models). Works with LM Studio, vLLM, llama.cpp, OpenAI proper (api_key_env references an env var — never store keys in the file).

Multi-GPU hosts

GPU detection prefers nvtop -s (all vendors + per-process usage), falling back to nvidia-smi/rocm-smi. GPUs are identified by PCI bus id — tool enumeration order is not trusted (nvtop order ≠ DRM cardN ≠ lspci order on multi-vendor hosts). Each GPU is bound to its /sys/class/drm/cardN entry via PCI match or device-name matching, so pciSlot/driver are always correct; missing VRAM is filled from sysfs mem_info_vram_* (i915/xe/amdgpu) or ollama's own discovery log.

get_system_resources returns an ollama_gpu block describing where the server runs:

  • env + env_source — GPU pinning vars discovered from the ollama serve process env (env_source: "proc"), falling back to the systemd unit (Environment=/EnvironmentFile, "systemd") or the journal's server config env=map[...] line ("journal"). Cross-user /proc denial is reported in env_note rather than silently yielding an empty env.

  • pinned_indices/pinned_slots — numeric selectors resolved to PCI slots through the journal's inference compute device table (e.g. GGML_VK_VISIBLE_DEVICES=2 → the Vulkan device whose pci_id matches index 2).

  • runners/gpus_in_use — loaded runner processes attributed to GPUs via /proc/<pid>/fd device nodes, keyed on PCI slot.

  • inference_devices — devices ollama discovered (name, pci_id, total/available VRAM) — authoritative for backends nvtop can't measure (e.g. Intel Arc via Vulkan).

To pin ollama to a specific card — e.g. an A380 while a P100 stays reserved for a vLLM server — set the backend-appropriate selector on ollama serve (GGML_VK_VISIBLE_DEVICES=N for llama.cpp Vulkan/Intel Arc, CUDA_VISIBLE_DEVICES=N NVIDIA, HIP_VISIBLE_DEVICES/ROCR_VISIBLE_DEVICES AMD, ONEAPI_DEVICE_SELECTOR=level_zero:N Intel SYCL) in its unit's EnvironmentFile, restart ollama, then get_system_resources shows the pin and recommend_model sizes its budget against that GPU — never a roomier GPU ollama can't use.

Quick single-provider env mode (no config file): LOCAL_LLM_BASE_URL, LOCAL_LLM_API_KEY_ENV, LOCAL_LLM_MODEL, LOCAL_LLM_PROVIDER_TYPE.

llama.cpp / exllamav3 / vLLM

These already work — their servers (llama-server, TabbyAPI, vLLM) all expose an OpenAI-compatible /v1, so just add type: "openai" provider entries:

"llamacpp": { "type": "openai", "base_url": "http://localhost:8080/v1", "default_model": "local-model" },
"tabby":    { "type": "openai", "base_url": "http://localhost:5000/v1", "api_key_env": "TABBY_API_KEY" }

What you lose vs. ollama-type providers: no num_ctx control (llama.cpp uses -c at launch), no pull/ps/unload management, and no system-resources introspection for remote hosts. download_model fills the acquisition gap — it runs hf download as a job (requires the hf CLI from huggingface_hub). Note llama-server can't hot-load a downloaded GGUF without a restart; TabbyAPI can load via its own admin endpoints.

Dev

npm install
npm test            # unit tests (node:test)
npm run smoke       # end-to-end: spawns server, hits real ollama
npm run build       # optional: emit dist/ for node<22.18 or distribution
node src/index.ts   # run the server standalone (stdin MCP)

Related MCP Connectors

Related MCP Servers