openai-mcp
Downloads models from Hugging Face via the hf CLI as a background job, so they can be served by local OpenAI-compatible inference servers such as llama.cpp, TabbyAPI, or vLLM.
Default local model provider integration. Lists installed and registry models (with sizes, capabilities, and tags), pulls, deletes, and prunes models, unloads models to free VRAM, and detects which GPU(s) loaded models are using. All generation tools (chat, complete, summarize, extract, classify, map_files, propose_write, run_local_agent) can be routed to an Ollama model with provider-specific controls like think, num_ctx, keep_alive, and native options.
Provider integration for any OpenAI-compatible /v1 endpoint (e.g. vLLM, llama.cpp, TabbyAPI) alongside Ollama. Supports the same generation and agent tooling, mapping thinking controls to reasoning_effort or chat_template_kwargs, and allows provider-native request overrides through an options bag for per-call or per-provider defaults.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@openai-mcpsummarize src/index.ts with the local model"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
openai-mcp — a local subagent for Devin
An MCP server that lets Devin offload low-complexity, monotonous work to a cheap local model (Ollama by default, or any OpenAI-compatible endpoint) instead of burning frontier tokens on it.
The headline feature is run_local_agent: the small model runs its own tool loop server-side
(read/search files, run whitelisted commands, stage writes). File contents and intermediate steps
never enter Devin's context — only a compact final answer comes back. Think of it as a local,
free-tier version of Devin's cloud subagents.
Why it actually saves tokens
Server-side file I/O.
summarize,extract,map_files, andrun_local_agenttake paths/globs and read files inside the server process. A 200KB file summarized this way costs Devin ~500 tokens instead of ~60k.Compact results. Tool results, diffs, and agent steps are truncated/capped before returning.
Visible savings. Every response reports
usageanddelegated_bytes— bytes processed that never entered Devin's context.get_usage_statstotals it for the session.
Related MCP server: Local Worker MCP
Requirements
Node >= 22.18 (runs TypeScript natively — no build step). Verified on Node 26.
Ollama (
ollama serve) for the default provider, or any OpenAI-compatible/v1endpoint.rg(ripgrep) is optional —search_filesuses it when installed and falls back to a built-in walker (gitignore-aware, hidden/binary skipped) when it isn't.nvtop(optional) for the best GPU detection — all vendors + per-process usage; falls back tonvidia-smi/rocm-smi.
Setup
1. Server config
npm install # required once — the server refuses to start without node_modules
npm run configure # interactive wizard: runs `npm install` itself if node_modules is missing,
# probes ollama, writes ~/.config/openai-mcp/config.json, prefills from an
# existing config (update mode), then covers steps 2 & 3 below — installs the
# server into ~/.config/devin/mcp_config.json AND merges recommended
# permissions into ~/.config/devin/config.json (backs both up first)
# or: npm install && cp config.example.json ~/.config/openai-mcp/config.json (edit by hand)File roots are auto-detected. Effective roots = allowed_roots (config) ∪ workspace roots the
MCP client advertises (the spec's roots capability) ∪ the directory Devin spawned the server from
(= your project dir, controlled by trust_cwd: true). So file tools work out of the box inside the
project, and allowed_roots is for extra directories outside it. deny_globs blocks secrets
everywhere. get_server_info shows each root with its source.
2. Register with Devin
npm run configure does this for you — shown here for manual setup.
Devin settings → MCP → Add custom MCP (or edit ~/.config/devin/mcp_config.json):
{
"mcpServers": {
"local_llm": {
"command": "node",
"args": ["/path/to/openai-mcp/src/index.ts"],
"env": {
"LOCAL_LLM_CONFIG": "~/.config/openai-mcp/config.json",
"LOCAL_LLM_RUN_COMMAND": "1",
"LOCAL_LLM_DYNAMIC_TOOLS": "1"
}
}
}
}Provider entries also accept think (bool or "low"/"medium"/"high"), options
(provider-native request overrides), temperature, num_ctx, and keep_alive — defaults applied
to every call; per-request args override them. The configure wizard asks about think since
thinking-capable models can eat the whole max_tokens budget.
CLI equivalent: devin mcp add -s user local_llm -- node /path/to/openai-mcp/src/index.ts
Omit the two LOCAL_LLM_* env flags to disable shell access and dynamic tools entirely.
3. Recommended Devin permissions
npm run configure merges these for you (with a backup) — shown here for manual setup.
In ~/.config/devin/config.json:
{
"permissions": {
"allow": ["mcp__local_llm__chat", "mcp__local_llm__complete", "mcp__local_llm__summarize",
"mcp__local_llm__extract", "mcp__local_llm__classify", "mcp__local_llm__map_files",
"mcp__local_llm__run_local_agent", "mcp__local_llm__list_models",
"mcp__local_llm__get_system_resources", "mcp__local_llm__list_staged",
"mcp__local_llm__get_diff", "mcp__local_llm__discard_write",
"mcp__local_llm__get_usage_stats", "mcp__local_llm__list_dynamic_tools",
"mcp__local_llm__get_command_whitelist", "mcp__local_llm__get_job",
"mcp__local_llm__list_jobs", "mcp__local_llm__control_job",
"mcp__local_llm__get_server_info", "mcp__local_llm__propose_write",
"mcp__local_llm__propose_edit",
"mcp__local_llm__recommend_model", "mcp__local_llm__list_commits",
"mcp__local_llm__revert_write", "mcp__local_llm__search_models",
"mcp__local_llm__list_model_tags", "mcp__local_llm__wait"],
"ask": ["mcp__local_llm__commit_write", "mcp__local_llm__set_write_mode",
"mcp__local_llm__pull_model", "mcp__local_llm__download_model",
"mcp__local_llm__delete_model", "mcp__local_llm__prune_models",
"mcp__local_llm__register_tool", "mcp__local_llm__unregister_tool",
"mcp__local_llm__update_command_whitelist", "mcp__local_llm__unload_model",
"mcp__local_llm__uncommit_write", "mcp__local_llm__verify_staged"]
}
}ask = you get an approve button before writes hit disk, models download, or new tools register.
Tools
Tool | What it does |
| Local subagent. Small model runs its own tool loop; returns final answer + staged op ids. |
| Raw completions passthrough. |
| Summarize files/globs server-side; only summaries return. |
| Structured JSON extraction, optional schema validation + retry. |
| Label text/files into candidate labels. |
| Apply one instruction across a glob — |
| Local model rewrites/creates one file → staged diff. |
| Patch-style edit: the model returns |
| Staged-write lifecycle. Nothing the local model writes reaches disk without |
|
|
| Lint/typecheck a staged op without touching the real file: content is materialized to a temp file in the same directory (imports resolve), |
| Undo: |
| Models per provider (sizes, capabilities, last-modified). |
| RAM + VRAM + loaded models (local ollama only). |
| Rank installed models vs detected memory budget (pinned/observed GPU > discrete GPU > RAM) with three-way |
| Search the public ollama registry for models — name, description, capability chips (tools/thinking/vision), parameter sizes, installed flag. |
| Pullable tags for a registry model ( |
| Download a model on ollama — background job, poll |
| Steer a running job: |
| Sleep N seconds (max 600); with |
| Download from Hugging Face via |
| Free a model's VRAM. |
| Permanently delete a model (irreversible). Refuses to delete the provider default or a loaded model without |
| Bulk-delete old models. |
| Give the local agent new tools at runtime ( |
| Manage what |
| Session savings + effective config. |
Slash commands (MCP prompts): /mcp__local_llm__delegate <task> [files] and /mcp__local_llm__cheap_summary <path> [focus].
Generation params
Every generation tool accepts these on top of model/provider/temperature/num_ctx/max_tokens:
Param | Effect |
| Chain-of-thought control: |
| Standard sampling controls. |
| Provider-native override bag — merged into ollama |
Set defaults per provider in config: "think": false, "options": {"top_k": 20}. Request args
override them.
Thinking models eat max_tokens. On thinking-capable models (gemma4, qwen3, gpt-oss) a tight
max_tokens can be consumed entirely by reasoning, returning empty text. Fix per call with
think:false, or globally per provider with "think": false. chat/complete return the model's
thinking trace when the provider reports it, so the budget usage is observable.
(run_local_agent accepts think/options too — disabling think leaves more context and steps
for tool calls.)
Progress
Two free/cheap progress channels:
MCP
notifications/progress:run_local_agentreports each agent step andmap_filesreports per-file completion — if the client attaches aprogressToken. These go to client UI plumbing, not into the model context — zero token cost.Job polling: every
run_local_agentcall runs under a job.run_async:truereturns thejob_idimmediately; sync calls return the inline result if they finish withinagent_sync_grace_ms, otherwise thejob_id.get_jobshows live progress + the finalresult. Polling costs a few tokens per call, but only when Devin chooses to check. (pull_model/download_modelare always jobs.)
Safety model
Staged writes + write modes: local-model output lands in a review queue, never directly on disk. Default
write_mode: "propose"means the server never writes — Devin reviewsget_diffand applies through its own edit tools, so changes go through Devin's native file-review flow (at the cost of the diff/content tokens).set_write_mode("write")(orwrite_mode: "write"in config) letscommit_writeapply server-side — those bypass Devin's edit-review UI but stay git-visible. The requiredpath/summaryargs keep theask-permission prompt meaningful, andcommit_write/set_write_modebelong inaskpermissions either way.run_command: whitelist-only, executed viaexecFileargv — no shell, so;,&&,|,>are literal arguments and can't inject.allow_argsconstrains subcommands (e.g. git → read-only verbs).File scope: union of configured
allowed_roots+ MCP-advertised workspace roots + cwd (disable withtrust_cwd: false), minusdeny_globsfor secrets.VRAM: ollama calls set
num_ctxexplicitly (default 16384) — an unbounded context blew up the KV-cache allocation on a 16GB GPU.get_system_resources+unload_modelhelp manage memory.Honesty: the instructions tell Devin not to delegate anything where a plausible-but-wrong answer is worse than no answer. 8B models produce garbage sometimes —
verify:"self"and staged diffs are the mitigations, not a guarantee.
Provider config
Two types:
ollama→ native/api(chat withnum_ctx/keep_alive/tools, pull, ps, unload, delete). Local or remote.openai→ generic/v1(chat completions, list models). Works with LM Studio, vLLM, llama.cpp, OpenAI proper (api_key_envreferences an env var — never store keys in the file).
Multi-GPU hosts
GPU detection prefers nvtop -s (all vendors + per-process usage), falling back to
nvidia-smi/rocm-smi. GPUs are identified by PCI bus id — tool enumeration order is not
trusted (nvtop order ≠ DRM cardN ≠ lspci order on multi-vendor hosts). Each GPU is bound to its
/sys/class/drm/cardN entry via PCI match or device-name matching, so pciSlot/driver are always
correct; missing VRAM is filled from sysfs mem_info_vram_* (i915/xe/amdgpu) or ollama's own
discovery log.
get_system_resources returns an ollama_gpu block describing where the server runs:
env+env_source— GPU pinning vars discovered from theollama serveprocess env (env_source: "proc"), falling back to the systemd unit (Environment=/EnvironmentFile,"systemd") or the journal'sserver config env=map[...]line ("journal"). Cross-user/procdenial is reported inenv_noterather than silently yielding an empty env.pinned_indices/pinned_slots— numeric selectors resolved to PCI slots through the journal'sinference computedevice table (e.g.GGML_VK_VISIBLE_DEVICES=2→ the Vulkan device whosepci_idmatches index 2).runners/gpus_in_use— loaded runner processes attributed to GPUs via/proc/<pid>/fddevice nodes, keyed on PCI slot.inference_devices— devices ollama discovered (name,pci_id,total/availableVRAM) — authoritative for backends nvtop can't measure (e.g. Intel Arc via Vulkan).
To pin ollama to a specific card — e.g. an A380 while a P100 stays reserved for a vLLM server —
set the backend-appropriate selector on ollama serve (GGML_VK_VISIBLE_DEVICES=N for llama.cpp
Vulkan/Intel Arc, CUDA_VISIBLE_DEVICES=N NVIDIA, HIP_VISIBLE_DEVICES/ROCR_VISIBLE_DEVICES
AMD, ONEAPI_DEVICE_SELECTOR=level_zero:N Intel SYCL) in its unit's EnvironmentFile, restart
ollama, then get_system_resources shows the pin and recommend_model sizes its budget against
that GPU — never a roomier GPU ollama can't use.
Quick single-provider env mode (no config file): LOCAL_LLM_BASE_URL, LOCAL_LLM_API_KEY_ENV,
LOCAL_LLM_MODEL, LOCAL_LLM_PROVIDER_TYPE.
llama.cpp / exllamav3 / vLLM
These already work — their servers (llama-server, TabbyAPI, vLLM) all expose an
OpenAI-compatible /v1, so just add type: "openai" provider entries:
"llamacpp": { "type": "openai", "base_url": "http://localhost:8080/v1", "default_model": "local-model" },
"tabby": { "type": "openai", "base_url": "http://localhost:5000/v1", "api_key_env": "TABBY_API_KEY" }What you lose vs. ollama-type providers: no num_ctx control (llama.cpp uses -c at launch),
no pull/ps/unload management, and no system-resources introspection for remote hosts.
download_model fills the acquisition gap — it runs hf download as a job (requires the hf CLI
from huggingface_hub). Note llama-server can't hot-load a downloaded GGUF without a restart;
TabbyAPI can load via its own admin endpoints.
Dev
npm install
npm test # unit tests (node:test)
npm run smoke # end-to-end: spawns server, hits real ollama
npm run build # optional: emit dist/ for node<22.18 or distribution
node src/index.ts # run the server standalone (stdin MCP)This server cannot be deployed
Maintenance
Related MCP Connectors
Hand tasks, bugs and finished work to AI coding agents, and get back a write-up with evidence.
Deterministic AI agent microtools, no accounts/API keys. fetch_extract: 98% token cut. 38 tools.
Shared control plane for AI coding agents — tasks, memory, decisions, file locks. 12 tools.
LLM chat, text tools, image generation, editing, batch image jobs, and asynchronous video generation
Related MCP Servers
- AlicenseBqualityDmaintenanceEnables coding agents like Claude Code and Codex to offload boilerplate generation, summarization, and other bounded text tasks to local or cheap cloud LLMs, keeping the frontier agent in charge of judgment and code edits.93MIT
- AlicenseBqualityCmaintenanceDelegates heavy, repetitive, and verifiable tasks like PDF extraction, code analysis, and log processing to a local LLM to reduce token consumption for frontier AI models, while keeping decision-making with the main AI.8MIT
- AlicenseAqualityBmaintenanceLets a frontier coding agent delegate research, cataloguing, and long-running computation to a local LLM with guarded filesystem, web, and Python execution tools, preserving the agent's context and tokens.6MIT
- AlicenseNot gradedqualityCmaintenanceTurns bounded coding tasks into reviewable Git diffs using isolated worktrees, configurable worker models, and test execution.MIT