codex-subagent-mcp
codex-subagent-mcp lets Claude Code delegate coding tasks to a local OpenAI Codex CLI, choosing model and reasoning effort per task and running read-only by default.
Check setup —
codex_doctorverifies the Codex CLI is installed, recent enough, signed in, and reports exact fixes.Discover models —
list_codex_modelsreads the live model catalog and each model's supported reasoning-effort levels.Get advice —
codex_recommendsuggests a model and effort for a described task (quality/balanced/latency/cost priority) without running a model call.Delegate work —
codex_delegateruns a self-contained task with a chosen model, reasoning effort, sandbox, working dir, target files, acceptance criteria, and timeout.Control the sandbox — read-only by default;
workspace-writeordanger-full-accessfor edits, withadd_dirs,use_worktree(edits land in a managed git worktree), andauto_approve.Run in background —
mode=backgroundreturns ajob_idimmediately (up to eight at once);codex_job_status,codex_job_result, andcodex_job_cancelmanage them.Continue sessions —
codex_follow_upreuses a priorthread_idso only the new instruction is sent, at a fraction of the cost.Structure output — pass an
output_schemato get schema-matching JSON back for direct action by Claude.Search the web —
web_searchenables Codex's live search for current external information.Stay bounded — server config caps sandbox, effort, allowed models, and inherited MCP servers/plugins/apps; prompts are sent to OpenAI and spend your own Codex usage.
Delegates coding tasks from Claude to OpenAI's Codex CLI running on the same machine, driving the installed CLI as a subagent. Supports choosing a Codex model (e.g. gpt-6-astra, gpt-5.6-terra) and reasoning effort per task, read-only investigation by default or explicitly allowed file edits, capturing the list of changed files, running long jobs in the background with job status polling, continuing a delegated thread, and confining writes to a managed git worktree. Also exposes a doctor check reporting whether the Codex CLI is installed and signed in.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@codex-subagent-mcphave codex investigate why the auth tests are flaky and report back"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
codex-subagent-mcp
An MCP server that lets Claude Code delegate coding tasks to OpenAI's Codex CLI running on the same machine, with the model and reasoning depth chosen per task. Claude stays the orchestrator. Codex becomes a subagent it can call.
An independent project. Not affiliated with, endorsed by, or supported by OpenAI or Anthropic.
Quick start
Requires Node.js 22+ and the Codex CLI installed,
on PATH, and signed in.
claude mcp add codex-subagent -- npx -y codex-subagent-mcpAsk Claude to run codex_doctor if anything is missing. See docs/INSTALL.md
for platform setup, Claude Desktop and other installation options, and Safety before
enabling writes.
Related MCP server: Hydra
What you get
You pick the model and the reasoning effort per task, from the catalog your Codex CLI reports live.
Read-only by default, with a sandbox ceiling no tool call can exceed.
Every result states what Codex actually applied (model, effort, sandbox, directory), not only what was asked.
Optional
output_schemareturns JSON that matches your schema, for results Claude can act on directly.Cancelling, timing out or shutting down stops Codex and every command it started.
A real read-only delegation asking for the package name, captured 2026-09-30:
codex-subagent-mcp
Commands run (1 total):
- [exit 0] /bin/zsh -lc "sed -n '1,80p' package.json"
model=gpt-5.6-luna | effort=low | sandbox=read-only | working_dir=/Users/you/project | applied=confirmed | duration=8s | tokens=in 35701 (cached 28160, uncached 7541) / out 88 (reasoning 9) | thread_id=01a0f419-9fd2-79e2-8248-91436f04e297Why this exists
A single model doing everything has three recurring problems, and delegation solves each one:
Your context window is finite. Having Claude read forty files to answer one question spends context you need for the actual work. Delegating the investigation returns the answer instead of the forty files.
One model has one set of blind spots. A second opinion is worth most when it comes from a different model family — different training, different failure modes. Asking the same model twice mostly gets you the same answer twice.
Not every task deserves the same reasoning budget. Renaming a variable and diagnosing a race condition are not the same job. Here they are separate dials: the model sets raw capability, the reasoning effort sets how long it deliberates. Cheap work goes to a fast model; a hard problem gets the capable one thinking for as long as it needs.
The server runs on your machine and drives the Codex CLI you already have installed; it holds no
credentials of its own. Prompts reach OpenAI through Codex, exactly as when you run codex
yourself.
See How it compares for a versioned comparison with other Codex MCP servers.
Using it
Write a bounded delegation
A delegation gets expensive when repeated commands keep adding output to the context carried into
later requests. Name the exact question, likely files, stopping condition and evidence the answer
must contain; choose higher effort for ambiguity rather than by habit.
Writing a delegation gives the measured cost model, ranked rules and
weak-versus-strong examples using the real tool parameters. Background goes in context and a
persona or extra rules in system_instructions, both layered on the built-in quality contract; see
the tool reference.
Get a second opinion from a different model family
The value here is not a second run — it is a different set of blind spots.
Ask Codex to review
src/server.tsfor correctness problems, focusing on error paths. Use a high reasoning effort and tell it to report each finding with the line and why it matters.
A delegated review like this found the terminate() defect in this repository's own runner: two
code paths could each arm a timer while only one was ever cleared.
Investigate without spending your context
Forty files go into the delegation; one answer comes back. Codex runs its own searches and reads whatever it needs; your conversation receives the conclusion.
Have Codex trace how a reasoning effort travels from the MCP tool call down to the arguments handed to the Codex CLI, and report just the call chain.
Run long work in the background while you keep going
Kick off a Codex run in the background that writes unit tests for
src/jobs.ts, then keep helping me with the API layer.
You get a job_id immediately. Ask for the status whenever you want, and read the result when it is
done. Up to eight can run at once.
Buy deep reasoning for one hard problem
Raising the reasoning effort for the whole conversation is expensive. Raising it for one delegation is not.
This intermittent test failure has beaten me twice. Ask Codex to work out the root cause at maximum reasoning effort, give it
test/runner.test.tsand the CI log, and tell it not to change anything — I want the diagnosis first.
Keep the thread going
Ask Codex to expand on its second finding.
Follow-ups reuse Codex's context, so they cost a fraction of the original. The server restates the same model, effort and directory on every follow-up because Codex itself does not keep them on resume.
Let it write, when you mean it
Have Codex apply its first two suggestions. Let it edit files, but keep it inside a git worktree so my working tree stays clean.
use_worktree sends the run's edits to ~/.codex/worktrees/; results list the files it touched and
where each landed, up to a thousand distinct files, then report the omitted count. The server does
not clean worktrees up: they may hold unapplied work. The experimental feature is enabled only for
that invocation; your Codex configuration is unchanged. See Safety for the write
boundary.
Whatever the sandbox, a delegation that writes reports what it wrote:
Files changed (2):
- [edit] src/codex/runner.ts
- [add] test/runner.test.tsSafety
This server runs another program on your machine, so it is worth two minutes before you enable writes.
What protects you
Delegations are read-only by default. Writing requires an explicit sandbox: "workspace-write"
or a user-set default. use_worktree sends edits to a managed git worktree instead of your
checkout; the sandbox and add_dirs bound where it can write at all. Unsandboxed runs are
unavailable unless you explicitly opt into that ceiling.
The confinement is the operating system's own sandbox: Seatbelt on macOS, bubblewrap on Linux and
WSL2, and a native sandbox on Windows. The table was measured on macOS;
SECURITY.md records the measurement details.
Linux and Windows have not been measured here, and OpenAI's Windows documentation notes that
sandboxed commands can fail to read some directories, so reads may be stricter there:
|
|
| |
Write inside the working directory | no | yes | yes |
Write outside it (your home) | no | no | yes |
Network access | no | no | yes |
Read outside the working directory | yes | yes | yes |
The sandbox confines the commands Codex runs in its shell. It does not confine MCP or plugin tools:
those run as their own processes with your account's full permissions, network included, and under
read-only a tool that declares itself read-only runs without approval (measured on codex-cli
0.159.2; ADR 16). So a delegation inherits no MCP
server, plugin or app from your Codex setup unless you allow it in the server configuration, and
every result ends with a line saying what the run was allowed.
There is no shell in the server's invocation path: the CLI is spawned with an argv array and the prompt is written to its stdin, never interpolated into a command string. Shell metacharacters in a prompt are inert.
What does not protect you
Reads are not confined. Codex can read anything your user account can, in every mode — your SSH keys, your cloud credentials. That was measured on macOS, and it is the safe assumption on every platform. Sandboxed network access is blocked so it cannot send them anywhere, but its report comes back to you, and that is a channel.
A prompt is untrusted input, and Codex acts on it. Content you did not write — an issue body, a
web page, a log, a file from someone else's repository — can carry instructions. With
workspace-write it can direct Codex to modify your repository; even read-only it can direct Codex
to read something sensitive and put it in the answer. The sandbox bounds where Codex can write. It
does not judge what it should write, or why it was asked. This is prompt injection, and it is the
risk that matters here.
The result is not sanitised. What comes back is text from a model that just read your files. Treat it as data, not as instructions; review what a delegation did rather than assuming it did what you asked.
Reducing the risk
Leave the built-in default alone. Read-only handles investigation, review and diagnosis, which is most delegation.
If you never want writes, cap it:
CODEX_SUBAGENT_MAX_SANDBOX=read-only. No conversation can argue past a ceiling. Register it outside the repository (Claude Code's defaultlocalscope,--scope user, or Claude Desktop's config), not in a project.mcp.jsonthat a write-enabled delegation could edit. See Configuration.When you enable writes, add
use_worktreeso changes land somewhere you can inspect before they touch your branch.Allow MCP servers, plugins or apps in delegations only when a task needs them, and by name. An allowed one runs outside the sandbox, and with
auto_approveits tools need no approval at all.Do not assemble delegation prompts from untrusted content when you intend to act on the answer.
If this threat matters seriously to you, run Codex under an account or container with no access to your secrets. That solves it at the root instead of bounding it.
SECURITY.md has the full threat model, what a deny_read policy could add, and how
to report a vulnerability.
Configuration
Everything is optional, and set through environment variables on the MCP server:
Variable | Values | Default | Effect |
| Slug from | Unset: refused with a suggestion | Model when a call specifies none. |
|
| Model's default | Effort when a call specifies none. |
| Comma-separated slugs from | Unrestricted | Anything else is refused; a single entry acts as the default model. An empty entry ( |
|
|
| Sandbox when a call specifies none; cannot exceed the ceiling. |
|
|
| Calls above it are refused; |
|
| Unrestricted | Higher effort is lowered to a supported level, or refused if none fits. |
|
|
| MCP servers from Codex's config a delegation keeps; this server is always off. |
|
|
| Installed Codex plugins a delegation keeps, with the MCP servers they provide. |
|
|
| Whether a delegation keeps Codex's apps (connectors to external services). |
| Executable path |
| Override CLI resolution; see Installation. |
A model supports a subset of efforts; an unsupported effort is adjusted to the closest supported one with a note. See docs/TOOLS.md#configuration for full semantics and Safety for the sandbox boundary.
Claude decides when to delegate; each run sends its prompt to OpenAI and spends your Codex usage, in
any conversation where the server is available. docs/CONTROL.md covers client
permission prompts, server ceilings, version ranges and your own CLAUDE.md escalation rules. Those
rules stay yours; see ADR 12 and
ADR 14 for the policy split.
Choosing a model
Use list_codex_models for the live catalog from your installed CLI. Model and reasoning effort are
independent: the model sets raw capability; the effort sets how long it deliberates. ultra
additionally delegates subtasks automatically.
The server does not choose for you. A task's model depends on your budget and how costly a wrong answer is. Without a model in the call or configuration, it refuses with a recommendation for you to decide on.
With no default, tool descriptions tell Claude to call codex_recommend first, announce the
suggested model and effort, then delegate with both explicit. This is guidance to a model, not
enforcement. For advice, ask: “Which Codex model should handle migrating this repo's tests to
vitest?” The suggestion respects your model allow-list and effort ceiling; see
the recommendation reference.
To skip that step on later delegations, set a default from list_codex_models once:
claude mcp add codex-subagent -e CODEX_SUBAGENT_DEFAULT_MODEL=gpt-5.6-terra -- npx -y codex-subagent-mcpTools
Tool | What it does |
| Check the Codex CLI installation and report how to fix it. |
| List available models and their reasoning-effort levels. |
| Suggest a model and effort for a described task. |
| Run a task, blocking or in the background, optionally returning JSON that matches a schema. |
| Continue a previous delegation using its |
| Check a background delegation. |
| Read a finished background delegation's output. |
| Stop a running background delegation. |
Full parameter reference: docs/TOOLS.md.
FAQ
Does this cost money?
It uses your existing Codex quota, as running codex yourself does. This server adds nothing and
has no visibility into the cost. Higher efforts consume more, and ultra also delegates subtasks;
codex_recommend helps avoid spending ultra on low work.
Why drive the CLI instead of calling the OpenAI API? Delegated coding is not a single completion — it is an agentic loop with a sandbox, an approval model, session persistence and project instruction files. All of that lives in the Codex client, not in the model endpoint. See ADR 1.
Do I need Claude Code, or does Claude Desktop work? Either. Claude Code gets a one-line install; Claude Desktop needs a manual config entry.
It says Codex is not installed, but codex works in my terminal.
Most likely Windows with a global npm install, which produces a codex.cmd batch shim that cannot
be launched without a command shell. codex_doctor reports this as unsupported-shim and offers
two fixes. On macOS and Linux, check whether a Node version manager moved codex off PATH.
Does it work on Windows and Linux?
CI builds, tests and starts the server on Windows, macOS and Linux on every change, and checks that
the Codex CLI is resolved correctly on each. A real delegation has only been verified on macOS — the
CI runners have no Codex installation or credentials. Reports from Windows and Linux are welcome.
On Windows, install the Codex CLI with the PowerShell installer rather than npm; on Linux, install
bubblewrap for Codex's sandbox. See Installation.
Documentation
docs/INSTALL.md — installation, platform notes and troubleshooting.
docs/CONTROL.md — permissions, ceilings and when Claude delegates.
docs/DELEGATING.md — how to scope a delegation, with measured costs and worked prompts.
docs/TOOLS.md — every tool and parameter.
docs/COMPARISON.md — versioned comparisons with other Codex MCP servers.
docs/adr/ — why the design is what it is, decision by decision.
docs/ROADMAP.md — what is planned, and what is deliberately out of scope.
Issues — what is actually open right now.
docs/VERSIONING.md — what counts as a breaking change here.
CHANGELOG.md — what changed in each release, and the Codex CLI version it was verified against.
CONTRIBUTING.md — setup, and the rules that are not negotiable.
Disclaimer
Not an official product. This is an independent, community project. It is not affiliated with, endorsed by, sponsored by or supported by OpenAI or Anthropic. "Codex", "ChatGPT" and "OpenAI" are trademarks of OpenAI; "Claude" and "Claude Code" are trademarks of Anthropic. They are used here only to describe what this software interoperates with, which is nominative use — no claim is made to any of them. Neither company is responsible for this software, and problems with it should be reported here rather than to them.
No warranty. The software is provided "as is", without warranty of any kind, as stated in LICENSE. You use it at your own risk.
Read Safety before enabling writes; delegations spend your own Codex usage.
License
MIT. See LICENSE.
Available Tools
8 toolscodex_delegateDelegate a task to CodexA
Delegate a task to the local Codex CLI (OpenAI's coding agent), choosing model and reasoning effort. Use it when the user asks for Codex, or when handing work off clearly serves their request: a second opinion from a different model family, or an investigation that would otherwise flood this conversation. When the user has not named a model, call codex_recommend first and present its suggested model and effort to the user in the same message in which you say you are going to delegate, then pass both explicitly here. That recommendation is advice for an already-authorised delegation, not a replacement for the user's own preference. Everything passed in prompt, context and target_files is sent to OpenAI, and every run spends the user's own Codex usage, so do not delegate what you can answer directly, and tell the user when you delegate. Codex runs read-only unless a different default sandbox is configured. Set sandbox to workspace-write to let it edit files. Codex cannot see this conversation, so pass everything it needs in prompt, context, and target_files.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | blocking (default) waits and streams progress; background returns a job_id immediately. | |
| model | No | Catalog slug from list_codex_models. If omitted, the configured default is used; with no default configured the call is refused and the recommended model is returned. | |
| prompt | Yes | The task for Codex. Be specific and self-contained: Codex cannot see this conversation. | |
| context | No | Background Codex needs: prior findings, constraints, relevant excerpts. | |
| sandbox | No | Sandbox policy. Uses the configured default when omitted; without one, Codex runs read-only. | |
| add_dirs | No | Additional absolute directories that should be writable alongside working_dir. | |
| web_search | No | Enable Codex's API-backed live web-search tool for this run. In a read-only sandbox, shell commands have no network access, so this is the route to current external information. When omitted, Codex's own configured web_search mode applies. | |
| working_dir | No | Absolute path Codex uses as its working root. | |
| auto_approve | No | Adds --approve-for-me so Codex auto-approves its own commands. Only applies when sandbox is workspace-write. | |
| target_files | No | Paths Codex should focus on, relative to working_dir. | |
| use_worktree | No | Run in a managed git worktree. Writes outside it remain subject to the sandbox policy and add_dirs. | |
| output_schema | No | A JSON Schema object for this turn's final message, which then comes back as JSON in a delimited block. OpenAI's structured outputs apply: every object needs "additionalProperties": false and every property listed in "required". At most 64 KiB serialised. | |
| timeout_seconds | No | Wall-clock budget. Defaults to 1800s. | |
| reasoning_effort | No | Reasoning depth, independent of model choice. Uses the configured default or the model's default when omitted. Clamped to supported levels within the configured ceiling; refused if none qualify. | |
| acceptance_criteria | No | Concrete conditions that must hold for the task to be considered done. | |
| skip_git_repo_check | No | Allow running outside a git repository. | |
| system_instructions | No | Persona or extra rules inherited from the orchestrator, layered on the built-in quality contract. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false and openWorldHint=true, but the description adds substantial non-obvious context: prompt/context/target_files are sent to OpenAI, each run spends the user's own Codex usage, the user must be told when delegating, and the default sandbox is read-only unless reconfigured. This is exactly the kind of disclosure structured fields cannot carry.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core purpose and usage trigger are front-loaded, and for a 17-parameter tool the length is defensible. It is a dense single block with some redundancy (the 'Codex cannot see this conversation' constraint is echoed by the schema), but each sentence carries operational weight.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a high-complexity tool with 17 parameters, no output schema, and a nested output_schema parameter, the description covers delegation conditions, privacy, cost, sandbox behavior, and the recommendation handoff. Return behavior (job_id for background, structured output block) is left to schema descriptions rather than explained, which is the main remaining gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real meaning beyond the schema: the model parameter is tied to a codex_recommend workflow, sandbox read-only is the default and workspace-write is required for edits, and target_files/context/prompt are flagged as externally transmitted. This lifts it above the schema-only baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (delegate) and resource (a task to the local Codex CLI, described as OpenAI's coding agent), with additional scoping that distinguishes it from siblings by naming codex_recommend as the prerequisite recommendation tool. An agent can tell exactly what this does and how it differs from list_codex_models and codex_recommend.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use triggers ('when the user asks for Codex, or when handing work off clearly serves their request: a second opinion from a different model family, or an investigation that would otherwise flood this conversation') and an explicit when-not ('do not delegate what you can answer directly'). It also names the alternative workflow (codex_recommend first when no model is chosen), leaving nothing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
codex_doctorCheck the Codex CLI installationARead-only
Check whether the local Codex CLI is installed, recent enough, signed in and able to load its configuration, and report the exact steps to fix it if not. Run this when any other tool reports the CLI is unavailable, or before relying on delegation for the first time. It only inspects the installation; it never installs or changes anything.
| Name | Required | Description | Default |
|---|---|---|---|
| refresh | No | Re-probe the CLI instead of reusing the cached diagnosis. | |
| working_dir | No | Absolute directory to run the check in. Codex loads the configuration of the directory it runs in, so pass the one a delegation would use. Defaults to this server's own. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already declare readOnlyHint=true and openWorldHint=false, and the description's 'never installs or changes anything' reinforces that safety profile without contradicting it. It adds useful behavioral scope by explaining the inspection covers signed-in state and configuration loading, and that it emits remediation steps rather than performing fixes. It doesn't detail the exact return format or mention the cached diagnosis behavior described in the refresh parameter, but the schema covers that.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, no fluff, and the most decision-relevant facts come first: what it checks, when to run it, and what it will not do. Every sentence earns its place, and the coverage of purpose, usage, and safety is achieved in extremely compact form.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-required-parameter diagnostic tool, the definition is complete: it states the checks performed, what the output will contain (fix steps), when to invoke it, and that it is read-only. The schema covers the two optional parameters, and the annotations cover the side-effect profile, so nothing needed for correct selection or invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and both parameters already have strong descriptions: refresh explains the cache/re-probe behavior)Skip and working_dir explains the directory-sensitive configuration loading and its default. The tool description contributes contextual motivation by tying the check to delegation, but it does not add meaning beyond what the schema already provides. A baseline 3 is appropriate because the schema carries the parameter documentation burden.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ('check') and a concrete resource (the local Codex CLI installation), then elaborates with four concrete checks: installed, recent enough, signed in, and able to load configuration. It also states the actual outcome ('report the exact steps to fix it'), which differentiates this diagnostic tool from the delegation, recommendation, and job-management siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit trigger conditions: run when another tool reports the CLI is unavailable, or before relying on delegation for the first time. It also states a clear exclusion—it only inspects and never installs or changes anything—so an agent knows not to use it as a fixer and should look elsewhere for remediation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
codex_follow_upContinue a Codex sessionA
Send a follow-up message to a previous delegation using its thread_id. Codex retains the earlier context, so only the new instruction needs to be sent. Like a delegation, it is sent to OpenAI and spends the user's Codex usage.
| Name | Required | Description | Default |
|---|---|---|---|
| model | No | Override the model for this turn. Defaults to the thread's last model from memory or, on a registry miss, Codex's session file; if neither has it, the configured default, else the call is refused. | |
| prompt | Yes | The follow-up instruction. | |
| sandbox | No | Sandbox policy for this turn. Uses the configured default when omitted; read-only when unset. | |
| thread_id | Yes | The thread_id reported by a previous codex_delegate call. | |
| working_dir | No | Absolute directory to resume in. Defaults to the directory the thread last ran in. | |
| auto_approve | No | Not supported on follow-ups: true is refused and nothing runs. | |
| output_schema | No | A JSON Schema object for this turn's final message, which then comes back as JSON in a delimited block. OpenAI's structured outputs apply: every object needs "additionalProperties": false and every property listed in "required". At most 64 KiB serialised. | |
| timeout_seconds | No | ||
| reasoning_effort | No | Override the reasoning effort for this turn. Defaults to the thread's last effort when the model is unchanged, otherwise to the configured or model default. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint=false and openWorldHint=true, so the description carries the rest: the call goes to OpenAI, spends the user's Codex usage, and the thread is stateful (earlier context is reused). That is meaningful disclosure beyond structured fields, though sync/async behavior is left unstated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with who/what/how, followed by the stateful-context point and then the cost warning. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter, stateful, externally-billed call with no output schema, the description covers the essentials: thread linkage, context reuse, and cost. It omits whether the call blocks or returns a job handle (relevant given codex_job_status/codex_job_result siblings), which is the main remaining gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 89%, which already sets a baseline of 3, and the description genuinely adds prompt semantics: Codex retains earlier context, so only the new instruction needs to be sent. That is meaning beyond the schema's 'The follow-up instruction.'
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (send a follow-up message) plus the exact resource and linking key (a previous delegation, via its thread_id). The phrase 'previous delegation' implicitly separates it from codex_delegate, which starts a new one, so an agent can route correctly without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Send a follow-up ... using its thread_id' implies the precondition (you must already hold a thread_id from codex_delegate), and the cost sentence signals weight. However, no alternative is named and no explicit when-not condition is given, so usage is inferred rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
codex_job_cancelCancel a background delegationBDestructive
Terminate a running background delegation.
| Name | Required | Description | Default |
|---|---|---|---|
| job_id | Yes | The job_id returned by codex_delegate. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and readOnlyHint=false, clearly indicating a destructive operation. The description adds 'running' as a state constraint, which is useful. However, it omits key behavioral details like whether the job is killed immediately or gracefully, whether it can be undone, or what happens to any partial results. With annotations covering the safety profile, a 3 is appropriate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence with no wasted words. It is front-loaded with the action, though it is extremely terse, which might leave some ambiguity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (one parameter, no output schema) and the annotations covering destructiveness, the description is minimally adequate but lacks important context about side effects, job state requirements, and the return value or confirmation. It should do more to help an agent understand when and how to use it safely.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the single parameter job_id is fully documented in the schema, including that it comes from codex_delegate. The description adds no additional parameter meaning beyond what the schema provides, making the baseline 3 correct.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Terminate) and resource (running background delegation), which is clearer than the title. However, it does not distinguish itself from siblings like codex_job_status or codex_job_result, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives, or when not to use it. The description gives no context about the conditions under which cancellation is appropriate, such as whether the job must be running or what happens if it has already completed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
codex_job_resultRead a background delegation's resultARead-only
Return the full output of a finished background delegation. Errors if the job is still running.
| Name | Required | Description | Default |
|---|---|---|---|
| job_id | Yes | The job_id returned by codex_delegate. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds a meaningful behavioral constraint: it errors when the job is still running, telling the agent this is a post-completion read with a specific failure mode. It does not discuss output truncation or pagination, so 4.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero waste, and the completion precondition is front-loaded. The error condition follows immediately after the main action, keeping the most important routing information prominent.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter read tool with full schema coverage and annotations covering safety, the description supplies the necessary completion precondition and output scope. No output schema exists, so describing the return as 'full output' is appropriately complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the only parameter 'job_id' is documented as 'returned by codex_delegate'. The description adds no additional syntax, format, or constraint beyond the schema, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Return') and resource ('full output of a finished background delegation'), and the error condition distinguishes it from a status check. It does not explicitly name sibling codex_job_status as the alternative, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Implies use only after completion via 'finished background delegation' and 'Errors if the job is still running'. However, it does not explicitly tell the agent to call codex_job_status first or name alternatives, so usage is inferable but not spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
codex_job_statusCheck a background delegationARead-only
Report the state and recent activity of a background delegation started with mode=background. Call it with no job_id to list every known job.
| Name | Required | Description | Default |
|---|---|---|---|
| job_id | No | Omit to list all jobs. | |
| include_activity | No | Include the recent progress log for the job. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds that it reports state plus recent activity and that omitting job_id enumerates all jobs, but says nothing about pagination, log volume, or failure modes for unknown job ids.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the core purpose and with the no-arg listing behavior second; no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-required-param, read-only status tool with no output schema, the description covers what is queried (state and activity) and the two calling modes. It is nearly complete, with only return-shape hints and error handling left unspecified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and both parameters are already documented in the schema; the description's note about omitting job_id restates that documentation rather than adding format or constraint detail.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (report) plus resource (state and recent activity of a background delegation) and scopes it to delegations started with mode=background, which links it to codex_delegate and separates it from codex_job_result.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives one concrete usage rule — call with no job_id to list every known job — but never says when to prefer this over codex_job_result, codex_follow_up, or codex_job_cancel, so the sibling routing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
codex_recommendRecommend a Codex model and effortARead-only
Given a task description, recommend which Codex model and reasoning effort to delegate it with. Runs no model call; applies a documented matrix reconciled against the installed catalog.
| Name | Required | Description | Default |
|---|---|---|---|
| priority | No | Bias the reasoning effort: quality raises it, latency and cost lower it. Default balanced. | |
| working_dir | No | Absolute directory whose Codex configuration and model catalog should be used. Defaults to this server's own. | |
| task_description | Yes | What the delegated task involves, in one or two sentences. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint and openWorldHint annotations, the description discloses that the tool executes no model call and applies a documented matrix reconciled against the installed catalog. This adds meaningful context about how the recommendation is derived (deterministically, not by running a model) without conflicting with the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The two-sentence description is compact and front-loaded: it immediately states the input ('Given a task description') and the output ('recommend which Codex model and reasoning effort'). The second sentence clarifies a key behavioral trait without waste. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple tool with 3 parameters (1 required) and no output schema, the description is adequately complete. It specifies the recommendation output (model and effort) but does not detail the exact return structure (e.g., object or string). Given the low complexity and clear sibling context (codex_delegate, list_codex_models), this is sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
All three parameters are fully described in the input schema (100% coverage), so the description adds no parameter-specific detail beyond the schema. With full schema coverage, the baseline of 3 is appropriate; the description does not compensate with extra parameter guidance.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'recommend which Codex model and reasoning effort to delegate it with.' It also notes 'Runs no model call,' which clearly differentiates this advisory tool from execution tools like codex_delegate. The purpose is unambiguous and distinct from siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage ('Given a task description... recommend') but does not explicitly say when to use this tool versus alternatives like codex_delegate or list_codex_models. The phrase 'Runs no model call' hints it is not for execution, but there is no explicit 'Use this before delegating' or 'For execution, use codex_delegate.' Guidance is implied, not stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_codex_modelsList Codex modelsARead-only
List the Codex models available on this machine, with the reasoning-effort levels each one supports. Read from the installed Codex CLI, with a warned static fallback if its catalog cannot be read. Call this before codex_delegate when choosing a model explicitly.
| Name | Required | Description | Default |
|---|---|---|---|
| refresh | No | Bypass the cache and re-read the catalog from the CLI. | |
| working_dir | No | Absolute directory to read the catalog in. A project you have trusted in Codex can set its own catalog, so pass the directory a delegation would use. Defaults to this server's own. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark the tool read-only, and the description adds useful behavioral context: it reads from the installed Codex CLI and has a 'warned static fallback' if the catalog cannot be read. This goes beyond the annotation in a meaningful way, though caching behavior is left to the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each earning its place: the core purpose, the fallback behavior, and the usage guidance. It is front-loaded with the essential information and contains no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only listing tool, the description covers the key context: what is listed, where it is read from, the fallback behavior, and when to call it. It does not describe the return format, but no output schema exists and the purpose makes the return shape largely predictable.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema fully documents both parameters. The description does not add parameter-level meaning beyond the schema, which makes the baseline 3 appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific action ('List'), a precise resource ('Codex models available on this machine'), and additional detail ('reasoning-effort levels each one supports'). It also distinguishes itself from a sibling by pointing to codex_delegate, so an agent can tell why this tool exists.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit usage context: 'Call this before codex_delegate when choosing a model explicitly.' It clearly states when the tool should be used, though it does not name alternatives to avoid or describe when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
v0.4.0- Changed
codex_delegate1 field changed- added
Input schema / properties / output_schemaAdded value: +{ + "additionalProperties": {}, + "description": "A JSON Schema object for this turn's final message, which then comes back as JSON in a delimited block. OpenAI's structured outputs apply: every object needs \"additionalProperties\": false and every property listed in \"required\". At most 64 KiB serialised.", + "propertyNames": { + "type": "string" + }, + "type": "object" +}
- Changed
codex_follow_up1 field changed- added
Input schema / properties / output_schemaAdded value: +{ + "additionalProperties": {}, + "description": "A JSON Schema object for this turn's final message, which then comes back as JSON in a delimited block. OpenAI's structured outputs apply: every object needs \"additionalProperties\": false and every property listed in \"required\". At most 64 KiB serialised.", + "propertyNames": { + "type": "string" + }, + "type": "object" +}
5 tool updates
v0.3.0- Changed
codex_delegate2 fields changed- changed
Input schema / properties / sandbox / descriptionPrevious value: -"Sandbox policy. Defaults to read-only: Codex analyses and reports but cannot modify files."New value: +"Sandbox policy. Uses the configured default when omitted; without one, Codex runs read-only." - changed
Input schema / properties / web_search / descriptionPrevious value: -"Enable live web search for this run, through Codex's web_search = \"live\" setting. When omitted, Codex's own configured web_search mode applies."New value: +"Enable Codex's API-backed live web-search tool for this run. In a read-only sandbox, shell commands have no network access, so this is the route to current external information. When omitted, Codex's own configured web_search mode applies."
- Changed
codex_doctor1 field changed- added
Input schema / properties / working_dirAdded value: +{ + "description": "Absolute directory to run the check in. Codex loads the configuration of the directory it runs in, so pass the one a delegation would use. Defaults to this server's own.", + "type": "string" +}
- Changed
codex_follow_up2 fields changed- changed
Input schema / properties / model / descriptionPrevious value: -"Override the model for this turn. Defaults to the model the thread last ran with on this server; for a thread this server has no record of, the configured default, else the call is refused."New value: +"Override the model for this turn. Defaults to the thread's last model from memory or, on a registry miss, Codex's session file; if neither has it, the configured default, else the call is refused." - changed
Input schema / properties / sandbox / descriptionPrevious value: -"Sandbox policy for this turn. Defaults to read-only."New value: +"Sandbox policy for this turn. Uses the configured default when omitted; read-only when unset."
- Changed
codex_recommend1 field changed- added
Input schema / properties / working_dirAdded value: +{ + "description": "Absolute directory whose Codex configuration and model catalog should be used. Defaults to this server's own.", + "type": "string" +}
- Changed
list_codex_models1 field changed- added
Input schema / properties / working_dirAdded value: +{ + "description": "Absolute directory to read the catalog in. A project you have trusted in Codex can set its own catalog, so pass the directory a delegation would use. Defaults to this server's own.", + "type": "string" +}
2 tool updates
v0.2.0- Changed
codex_delegate5 fields changed- changed
Input schema / properties / auto_approve / descriptionPrevious value: -"Adds --approve-for-me so Codex auto-approves its own commands. Only applies when sandbox allows writes."New value: +"Adds --approve-for-me so Codex auto-approves its own commands. Only applies when sandbox is workspace-write." - changed
Input schema / properties / model / descriptionPrevious value: -"Catalog slug from list_codex_models. Omitted means the recommendation matrix picks one."New value: +"Catalog slug from list_codex_models. If omitted, the configured default is used; with no default configured the call is refused and the recommended model is returned." - changed
Input schema / properties / reasoning_effort / descriptionPrevious value: -"Reasoning depth, independent of model choice. Clamped to what the chosen model supports."New value: +"Reasoning depth, independent of model choice. Uses the configured default or the model's default when omitted. Clamped to supported levels within the configured ceiling; refused if none qualify." - changed
Input schema / properties / use_worktree / descriptionPrevious value: -"Run in a managed git worktree so changes never touch the current working tree."New value: +"Run in a managed git worktree. Writes outside it remain subject to the sandbox policy and add_dirs." - changed
Input schema / properties / web_search / descriptionPrevious value: -"Enable Codex's native web search tool."New value: +"Enable live web search for this run, through Codex's web_search = \"live\" setting. When omitted, Codex's own configured web_search mode applies."
- Changed
codex_follow_up4 fields changed- added
Input schema / properties / auto_approve / descriptionAdded value: +"Not supported on follow-ups: true is refused and nothing runs." - changed
Input schema / properties / model / descriptionPrevious value: -"Override the model for this turn."New value: +"Override the model for this turn. Defaults to the model the thread last ran with on this server; for a thread this server has no record of, the configured default, else the call is refused." - changed
Input schema / properties / reasoning_effort / descriptionPrevious value: -"Override the reasoning effort for this turn."New value: +"Override the reasoning effort for this turn. Defaults to the thread's last effort when the model is unchanged, otherwise to the configured or model default." - added
Input schema / properties / working_dir / descriptionAdded value: +"Absolute directory to resume in. Defaults to the directory the thread last ran in."
1 tool update
- Changed
codex_follow_up1 field changed- added
Input schema / properties / thread_id / patternAdded value: +"^[A-Za-z0-9][A-Za-z0-9_-]{0,199}$"
8 tool updates
v0.1.0- First observed
codex_delegate - First observed
codex_doctor - First observed
codex_follow_up - First observed
codex_job_cancel - First observed
codex_job_result - First observed
codex_job_status - First observed
codex_recommend - First observed
list_codex_models
TDQS
Scored across 8 tools
The job-lifecycle tools (cancel, result, status) and the delegation tools (delegate vs follow_up) are clearly distinguished by their descriptions. Only codex_recommend and list_codex_models overlap somewhat in the model-selection space, but one gives advice and the other enumerates the catalog, so confusion is limited.
Seven of eight tools use a predictable codex_<action> or codex_job_<action> snake_case pattern. list_codex_models deviates by inverting to verb-first without the codex_ prefix, a minor but noticeable break in the pattern.
Eight tools is well-scoped for a CLI delegation wrapper: one core delegate tool plus the supporting lifecycle (cancel, result, status), continuation (follow_up), and setup helpers (doctor, recommend, list_codex_models). Each tool earns its place with no redundancy.
The surface covers the full delegation lifecycle: starting, continuing, monitoring, retrieving, and cancelling jobs, plus environment diagnostics and model discovery. Only minor gaps exist, such as no explicit job cleanup/removal or batch operations, which agents can work around.
Maintenance
Related MCP Connectors
No-data MCP handoff for local Claude Code to Codex harness moves. $49 lifetime.
Source-checked CLI guides and model-aware planning for Claude Code, Codex, and Grok Build.
Use your Mac, Windows or Linux computer from ChatGPT, Claude or Codex: files, commands, documents.
A paid remote MCP for OpenAI Codex agent coordination MCP, built to return verdicts, receipts, usage
Related MCP Servers
- AlicenseAqualityDmaintenanceEnables Claude Code to delegate tasks to OpenAI's Codex CLI (GPT-5.4) with structured execution traces, parallel execution, session persistence, and adversarial code review.15MIT
- AlicenseNot gradedqualityBmaintenanceEnables Codex to delegate bounded engineering jobs to Claude Code CLI in isolated Git worktrees with strict security and allowance pacing.MIT
- AlicenseNot gradedqualityAmaintenanceLets Codex delegate coding and repository work to an installed Claude Code CLI with permission-aware inspect/write access, model and effort selection, resumable and cloud-attached sessions, and durable synchronous or asynchronous jobs.MIT
- AlicenseNot gradedqualityBmaintenanceEnables Claude Code to delegate bounded, read-only analysis and isolated patch proposals to the locally installed Codex CLI over MCP, with sanitized live status, revision-aware polling, and reviewable results.1MIT