codex-offload-mcp
Integrates with OpenAI's Codex CLI to execute coding tasks as background jobs, providing tools for starting, monitoring, and retrieving results.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@codex-offload-mcpfix the login bug in auth.js"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
codex-offload-mcp
An MCP server for running Claude Code and the Codex CLI as a pair: Claude drives, and hands self-contained tasks to Codex as background jobs.
codex_start returns a job id immediately instead of blocking, so the driving model keeps working
while Codex works alongside it — two agents on the problem at once, rather than one waiting on the
other.
Why this exists
To use two coding models together rather than one at a time. Claude Code holds the conversation, the context and the judgment about what to do next; Codex is a capable second agent that can be given a well-specified piece of work and left to get on with it. Neither replaces the other, and the interesting part is what they do concurrently.
Blocking would defeat the entire point. If handing work to Codex meant waiting for Codex, there would be no pair — just one model parked while the other runs, which is strictly worse than doing the work yourself. So dispatch returns in about a second and the job continues in a detached process that outlives even this server. The driving model carries on reasoning, reading and answering, and collects the result when it is ready.
That concurrency is also what makes the delegation worth specifying carefully. Codex cannot see the conversation, so a task has to travel as a complete brief — and the work that comes back is checked against git rather than taken on trust, because handing work between models is precisely where "I did the thing" and "the thing got done" come apart.
Related MCP server: codex-cli-mcp-tool
Tools
Tool | Returns |
|
|
|
|
| State, elapsed time, commands run, files touched, recent activity |
| Structured handoff report + git-verified changes; progress if still running |
| Continues a job's Codex thread with a follow-up; returns a new jobId |
| Available models and the reasoning efforts each one accepts |
| Kills the job and its child processes |
| Known jobs, newest first, optionally filtered by state |
None of them block.
The handoff
Getting work back from Codex matters as much as sending it. Three things make the return trip trustworthy:
1. A typed report. Jobs run with --output-schema, so the final message is not prose but a
fixed shape:
{
"summary": "...",
"status": "complete | partial | blocked",
"filesChanged": [{ "path": "...", "change": "..." }],
"verification": [{ "command": "...", "passed": false, "details": "..." }],
"followUps": ["..."],
"blockers": ["..."],
"documentation": "which docs were updated, or why none were warranted",
"confidence": "high | medium | low"
}This makes Codex commit to whether it actually finished and whether its checks passed, instead of
burying a failed test in a paragraph. Pass structured: false for a job where you want prose.
2. Changes verified against git, not self-reported. codex_result also returns
actualChanges, computed by diffing the repo against the commit and dirty-file set recorded when
the job started. Files already modified beforehand are flagged preexisting so they are not
mistaken for Codex's work. When the report and actualChanges disagree, actualChanges is the
one to trust.
3. A way to push back. codex_reply continues the original Codex thread, so corrections
("you missed the error path", "now update the tests") land with all of Codex's context intact
rather than starting a cold job that has to rediscover everything.
Install
Prerequisites
Node.js 20+
The Codex CLI, installed and authenticated — run
codex loginand confirmcodex --versionworks. This server drives that binary; without it every job fails immediately.
Build
git clone https://github.com/jgt87/codex-offload-mcp.git
cd codex-offload-mcp
npm install
npm run buildThis produces dist/index.js. Note its absolute path — every step below needs it.
Add to VS Code
MCP support is built into current VS Code; if the Command Palette lists MCP: commands, you have
it. Pick either route:
Guided. Command Palette (Ctrl+Shift+P) → MCP: Add Server → Command (stdio). Enter
node as the command and the absolute path to dist/index.js as the argument, then name it
codex-offload.
By hand. Command Palette → MCP: Open User Configuration to open your user mcp.json
(%APPDATA%\Code\User\mcp.json on Windows), and add the server:
{
"servers": {
"codex-offload": {
"type": "stdio",
"command": "node",
"args": ["C:/path/to/codex-offload-mcp/dist/index.js"]
}
}
}Use forward slashes on Windows, or escape backslashes as \\ — a raw C:\path is invalid JSON and
the server will silently fail to start.
To scope it to one project instead of your whole profile, use MCP: Open Workspace Folder
Configuration and put the same servers block in .vscode/mcp.json. That file can be committed,
which gives everyone on the repo the same tools.
Verify. Open the Chat view, switch to Agent mode, click Configure Tools, and confirm the
codex_* tools appear and are enabled. MCP: List Servers shows the server's status and its
logs if it failed to start.
Add to Claude Code
claude mcp add codex-offload --scope user -- node /absolute/path/to/dist/index.jsConfirm with /mcp in a session, or claude mcp list from a shell.
After changing the code
A running server keeps serving the old dist/, so rebuild and restart it:
npm run buildVS Code — MCP: List Servers → select the server → Restart. (The experimental
chat.mcp.autoStartsetting can do this for you.)Claude Code — restart the session; MCP servers connect at session start.
Using it
You do not call the tools by name. Ask for what you want and Claude selects them:
You say | Tool |
"Offload to Codex: migrate the auth module to the new API" |
|
"How's that Codex job doing?" |
|
"Get the Codex result" |
|
"Tell Codex it missed the error path" |
|
"What Codex jobs are running?" |
|
"Which models can Codex use?" |
|
"Kill that job" |
|
The pattern worth building a habit around is dispatch-then-continue: "Offload the test migration to Codex, and while that runs, walk me through the router."
Checking its work
codex_result gives you Codex's own report and actualChanges from git. When they disagree,
git wins. For anything that matters, go further than reading the report: run the tests, and break
the thing under test to confirm they actually fail. A suite that passes proves less than a suite
you have watched fail for the right reason.
Collaboration modes
One-shot offload — hand Codex a task, collect the result — is the base case. On top of it are a few named ways to split work between the two models, each with a slash command so you can start one from any repo. Install the commands by copying them into your Claude Code commands directory:
cp .claude/commands/*.md ~/.claude/commands/Command | Pattern | Who does what |
| plan → execute | Claude designs the change and writes a concrete plan; Codex carries it out faithfully in the background. |
| execute → review | Codex does the work; Claude reviews the git diff and sends corrections with |
| split & parallelize | Claude decomposes the task into independent chunks and dispatches several Codex jobs at once, then reassembles. |
| draft → refine | Codex produces a fast first draft cheaply; Claude refines it in-process where context and taste are needed. |
The split is the same one the whole server is built on: the reasoning, the judgement and the
conversation context stay with Claude; the expensive output tokens — typing out a plan, a bulk draft,
a mechanical migration — go to Codex, which bills separately. Three of the four modes need no new
code; they are patterns over codex_start, codex_result and codex_reply, named and given a front
door.
plan→execute is the exception — it has machinery of its own. It adds a tool,
codex_execute_plan, that takes a finished plan rather than an open task. Codex is told the plan came
from another model and to follow it faithfully — and, crucially, to stop and report a blocker
instead of quietly substituting its own design when a step is wrong, so you can revise the plan and
resume with codex_reply. Because the design thinking is already done and lives in the plan,
execution usually needs less reasoning effort than the whole task would; routing still reads the plan
text and effort stays overridable, so pin a lower reasoningEffort when the steps are mechanical.
It fits the same seam as the documentation instruction: the faithful-execution framing is prepended
to what Codex receives, but codex_status and codex_list still show the plan you actually handed
over, not the machinery around it.
Orchestration
Two decisions get made per task, and they are handled very differently.
Whether to delegate is not decided here. There is no scheduler, no queue and no second model
triaging work. The only thing steering that choice is the codex_start tool description, which the
calling model reads at call time and judges against.
That is deliberate. The decision needs the one thing this process cannot see — the conversation.
Whether a task is self-contained, whether there is useful work to do while it runs, whether you are
about to change your mind about the approach: none of that is visible from inside an MCP server, so
the judgment stays with the model that has the context, and this server sticks to running the job
and checking the result. Editing that description in src/index.ts is how you change delegation
behaviour in general.
One dial you can turn without touching code: how much to offload. The whether-decision stays the
model's, but you can bias it. Set the CODEX_MCP_OFFLOAD_LEVEL environment variable on the server to
conservative, balanced (the default) or aggressive, and it appends a matching instruction to
the codex_start description the model reads — aggressive lowers the bar so more work goes to Codex
(useful when you want to conserve Claude's own usage), conservative raises it. It steers judgment
rather than enforcing a rule: the hard exclusions — needs conversation context, exploratory, trivial
triage — still hold at every level. It is read once at startup, so restart the server after changing
it, and codex_models reports the active setting so you can confirm it took. In your MCP config it
goes in the server's env block:
{
"servers": {
"codex-offload": {
"type": "stdio",
"command": "node",
"args": ["C:/path/to/codex-offload-mcp/dist/index.js"],
"env": { "CODEX_MCP_OFFLOAD_LEVEL": "aggressive" }
}
}
}Which model and how hard it thinks are decided here — see below. That part is a genuine heuristic, and the honest framing is that it is the one invented thing in the pipeline: no API reports "this task is mechanical". So it is built to be inspectable rather than trusted. Every job records the tier and the reason it was picked, and any explicit value overrides it.
What should be offloaded
All of these need to hold:
Test | Why |
Self-contained | Codex cannot see the conversation. Anything resting on what was just worked out must be restated in full — and if restating it is most of the work, offloading is a net loss. |
Slow | Minutes, not seconds. Below ~30s the round trip costs more than it saves. |
Real work to do meanwhile | Dispatching and then sitting on |
Verifiable afterwards | Mechanical enough that |
Scoped to one | Codex writes to disk directly; ambiguous scope means unwanted edits. |
Keep it in-process when the task needs conversation context, is fast, blocks the next decision, or is surgical enough that specifying it precisely costs more than just doing it.
Exploratory work is the main trap. Investigations where each measurement changes what you look at next cannot be offloaded — by the time the prompt can be written, the thinking is already done. A useful tell is a wrong hypothesis: if you expect to have one, keep the work in-process.
Misjudging is asymmetric. A bad question wastes a minute; a bad delegation writes files to disk.
That asymmetry is why codex_result checks Codex's self-report against git instead of trusting it —
the design already assumes a delegation can be wrong. Prefer sandbox: "read-only" for anything
analytical.
Choosing model and effort
codex_start picks both from the task text unless you pass them. Here is the whole path a call
takes, from prompt to spawned process:
codex_start(prompt, model?, reasoningEffort?, autoRoute?)
│
├── autoRoute: false ────────────────────► skip everything below
│ (caller's values, else config.toml)
▼
classify(prompt) src/route.ts — keyword match, no model call
│
├── matches HARD patterns? ── yes ──► tier = hard
├── matches MECHANICAL patterns? ─ yes ──► tier = mechanical
└── neither ─────────────────────────────► tier = standard (hard wins ties)
│
▼
pick model match tier's wording against each model's vendor description
│ ("frontier" / "balanced" / "fast, affordable")
│ no match → Codex's own first-ranked model
▼
pick effort mechanical → low standard → medium hard → high
│
▼
clamp to model effort not in this model's supported list?
│ → nearest supported level at or below it
▼
apply overrides explicit model / reasoningEffort replace whatever was chosen
│
▼
validate model accepts this effort?
│ │
│ └── no ──► error returned in ms, nothing spawned
▼
spawn `codex exec --json` detached, recording {tier, rationale, auto} on the jobOnly the classify step is invented. Everything from "pick model" down is driven by the model index
Codex itself maintains, so the lineup and the legal effort levels are facts rather than guesses.
The tiers:
Tier | Signals | Effort | Model |
| renames, moving files, formatting, typos, applying a stated pattern |
| the one described as fast/affordable |
| anything without a strong signal — the default |
| the one described as balanced/everyday |
| concurrency, races, deadlocks, leaks, security, architecture, trade-offs, root-cause work |
| the one described as frontier |
A prompt matching both mechanical and hard is treated as hard: under-thinking a subtle problem
costs more than over-thinking a simple one.
Resolved against a real lineup, the assignment looks like this — the model column is whatever the index currently offers, not a fixed list:
TASK TIER MODEL EFFORT
─────────────────────────────────────────── ──────────── ─────────────── ──────
"Rename getUser to fetchUser across the mechanical gpt-5.6-luna low
repo and update the call sites" │ fast, affordable
└─ matched: "rename"
"Add an endpoint returning the current standard gpt-5.6-terra medium
user's projects" │ balanced/everyday
└─ no strong signal → default
"Fix the race condition in the job hard gpt-5.6-sol high
scheduler that drops events under load" │ frontier
└─ matched: "race condition"
"Rename the lock helper while fixing the hard gpt-5.6-sol high
deadlock it causes" │
└─ matched both; hard winsThe model lineup is discovered, not hardcoded. Models and their accepted effort levels are read
from Codex's own index (~/.codex/models_cache.json), so a model released after this server was
written is picked up automatically. Routing re-reads the index whenever Codex rewrites it (the file's
mtime changes), so a long-running server keeps up with Codex rotating or dropping models instead of
routing to one that no longer exists and failing the job after it spawns. Tiers map onto that lineup
by matching the vendor's own descriptions, which means a renamed model still routes sensibly.
codex_models reports the current lineup, including whether the index was actually readable. (The
model names shown inside the tool descriptions are still the startup snapshot — schema text is
fixed for an MCP session — so a restart refreshes those, but routing itself does not need one.)
This matters more than it sounds. The first version of this feature hardcoded the effort list, and
it was wrong within the hour — it invented none and minimal, which no model advertises, and
omitted ultra, which two models support.
Effort is clamped to what the chosen model accepts, so a model topping out at xhigh never
receives ultra. If you pin a model and an effort it cannot take, the call fails immediately with
the model's real list — rather than failing the job minutes later, which is what Codex does on its
own, since it forwards the value unchecked.
Overriding. An explicit model or reasoningEffort always wins, and partial pinning works —
fix the model, let the effort be routed, or the reverse. autoRoute: false disables inference
entirely and falls back to your ~/.codex/config.toml defaults. Worth knowing that if that file
sets model_reasoning_effort = "high", every un-routed job runs at high until you say otherwise.
Every job records what was chosen and why:
"routing": {
"tier": "mechanical",
"rationale": "classified as mechanical — matched mechanical wording (rename); model gpt-5.6-luna chosen for this tier; effort low",
"auto": true
}The classifier is keyword matching, and deliberately so — a cleverer one would need a model call, which would add latency to a tool whose whole promise is returning immediately. It will misread things. That is why it explains itself and why every part of it can be overridden.
codex_reply takes reasoningEffort too, defaulting to the parent job's. Worth raising when a first
attempt failed for want of thinking, and lowering when the follow-up is a mechanical fixup.
The feedback loop
Work does not come back trusted. Every job produces two independent accounts of itself, and the loop closes by comparing them:
codex_start
│ git baseline captured first: current commit + already-dirty files
▼
codex exec --json (detached — outlives this server)
│ streams events.jsonl while it works
▼
codex_status ── poll as often as you like; never blocks
│
▼
codex_result
│
├───────────────────────────┐
▼ ▼
Codex's report actualChanges
what it believes it did what git says changed,
summary, verification, measured against the
confidence baseline above; files
│ already dirty flagged
│ preexisting
│ │
└─────────────┬─────────────┘
▼
caller compares them ── git wins any disagreement
│
┌──────────┴───────────┐
▼ ▼
agree, tests pass disagree, or partial
│ │
▼ ▼
accept codex_reply
│
▼
new job on the SAME Codex thread
full context retained, effort adjustable
│
└──► back to codex_status, aboveThree properties make this worth the machinery:
The two accounts are produced independently. The report is what Codex says; actualChanges is
computed by diffing the repo against the baseline captured before the job started. Files you had
already modified are flagged preexisting, so they are never mistaken for Codex's work. When the two
disagree, git wins.
The correction path keeps context. codex_reply resumes the original thread by its thread_id,
so "you missed the error path" lands with everything Codex already knows, instead of a cold job that
has to rediscover the codebase.
The loop is closed by you, not by the router. Nothing here learns. A job records its tier and
rationale so you can look back and see whether the classification was reasonable, but a misrouted
task does not adjust anything for next time — you pin model or reasoningEffort on the retry, or
edit the patterns in src/route.ts. If routing keeps getting a particular kind of task wrong, that
is a signal to change the keyword lists, and it is meant to be done by hand.
How it works
What happens when a task is offloaded
Baseline.
codex_startrecords the state ofcwdbefore Codex touches anything: the current commit, plus the set of files already dirty. This is what makes the later verification honest — files you had already modified are flaggedpreexistingrather than blamed on Codex.Routing. Unless you pinned them, model and reasoning effort are chosen from the task text and clamped to what that model accepts. An effort the model cannot take is rejected here, before anything is spawned.
Dispatch. The prompt is written to
prompt.txtandcodex exec --jsonis spawned detached, with the prompt fed over stdin. The call returns a job id immediately; nothing blocks, ever.Streaming. Codex writes JSONL events to
events.jsonlas it works, and its final answer tolast-message.txt. Both are on disk, not in memory.Polling.
codex_statusparses the tail of the event stream into a progress view — commands run, files touched, recent activity. It reports; it never waits.Collection.
codex_resultreturns Codex's structured report and re-diffs the repo against the baseline from step 1 to produceactualChanges. Two independent accounts of the same work.Follow-up.
codex_replyresumes the original Codex thread by itsthread_id, so a correction lands with all of Codex's context intact rather than starting cold.
Under the hood
Each job spawns codex exec --json as a detached process, with:
the prompt piped over stdin — no command-line length or quoting limits
--jsonevents streamed toevents.jsonl(parsed for the progress view)-o last-message.txtcapturing the final answermetadata in
meta.json
Jobs live in ~/.codex-mcp/jobs/<jobId>/ (override with CODEX_MCP_JOBS_DIR). Finished jobs
older than seven days are pruned at startup.
Because jobs are detached, they survive this server restarting. A job started by a previous
process has nobody listening for its exit, so running is only trusted while its pid is alive;
otherwise state is recovered from whether a result file was produced.
The Codex binary is resolved to the real platform executable inside the npm package, so jobs spawn
with an argv array rather than through a shell. Override with CODEX_BIN if needed.
Sandbox
sandbox defaults to workspace-write: Codex may edit files under cwd. That is the point — it
does real work — but it means the working tree changes underneath you. Re-read files after a
job finishes rather than trusting anything cached from before.
read-only— analysis and review, no writesworkspace-write(default) — edits withincwd(plus anyaddDirs)danger-full-access— unrestricted; only when explicitly asked for
Cancelling does not roll back edits already made.
Documentation
Jobs that can write are asked to document themselves. A standing instruction is appended to the
prompt telling Codex to update the project's existing documentation — README.md, AGENTS.md,
CLAUDE.md, files under docs/ — when the change alters behaviour, adds a feature, or makes an
existing statement untrue.
This exists because Codex starts cold. It cannot see the conversation that led to the task, so unless something asks, documentation simply never gets written and the caller has to remember every single time.
The instruction leans hard in one direction: edit what exists, do not create new files. An
unqualified "document your changes" reliably produces a CHANGES.md or NOTES.md in every
repository it touches, which is worse than nothing — it fragments what a reader has to consult and
goes stale immediately. Codex is told not to invent a changelog, to match the surrounding document's
voice and structure, and to prioritise correcting statements the change made untrue over adding new
prose. It is also told to skip documentation when the work does not warrant it, and to say so.
The structured report carries a required documentation field, so a job has to account for what it
wrote or explain why it wrote nothing. actualChanges from git shows which .md files were actually
touched, so the claim is checkable rather than taken on faith — the same split the rest of the
handoff uses.
Setting | Behaviour |
default | On for |
| Suppressed; the prompt is passed through untouched |
| Always off — the job could not write a file even if asked |
A codex_reply follow-up inherits the parent job's setting, so a correction to documented work keeps
the docs in step with it. The instruction is appended only to what Codex receives; codex_status and
codex_list still show the prompt you actually wrote.
Writing good job prompts
Codex starts cold. It cannot see the Claude conversation, so a prompt must carry its own context: what to change, which files, the constraints, and what "done" looks like.
This server cannot be installed
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/jgt87/codex-offload-mcp'
If you have feedback or need assistance with the MCP directory API, please join our Discord server