Skip to main content
Glama

codex-subagent-mcp

CI npm version License: MIT

An MCP server that lets Claude Code delegate coding tasks to OpenAI's Codex CLI running on the same machine, with the model and reasoning depth chosen per task. Claude stays the orchestrator. Codex becomes a subagent it can call.

An independent project. Not affiliated with, endorsed by, or supported by OpenAI or Anthropic.

Quick start

Requires Node.js 22+ and the Codex CLI installed, on PATH, and signed in.

claude mcp add codex-subagent -- npx -y codex-subagent-mcp

Ask Claude to run codex_doctor if anything is missing. See docs/INSTALL.md for platform setup, Claude Desktop and other installation options, and Safety before enabling writes.

Related MCP server: Hydra

What you get

  • You pick the model and the reasoning effort per task, from the catalog your Codex CLI reports live.

  • Read-only by default, with a sandbox ceiling no tool call can exceed.

  • Every result states what Codex actually applied (model, effort, sandbox, directory), not only what was asked.

  • Optional output_schema returns JSON that matches your schema, for results Claude can act on directly.

  • Cancelling, timing out or shutting down stops Codex and every command it started.

A real read-only delegation asking for the package name, captured 2026-09-30:

codex-subagent-mcp
Commands run (1 total):
- [exit 0] /bin/zsh -lc "sed -n '1,80p' package.json"
model=gpt-5.6-luna | effort=low | sandbox=read-only | working_dir=/Users/you/project | applied=confirmed | duration=8s | tokens=in 35701 (cached 28160, uncached 7541) / out 88 (reasoning 9) | thread_id=01a0f419-9fd2-79e2-8248-91436f04e297

Why this exists

A single model doing everything has three recurring problems, and delegation solves each one:

Your context window is finite. Having Claude read forty files to answer one question spends context you need for the actual work. Delegating the investigation returns the answer instead of the forty files.

One model has one set of blind spots. A second opinion is worth most when it comes from a different model family — different training, different failure modes. Asking the same model twice mostly gets you the same answer twice.

Not every task deserves the same reasoning budget. Renaming a variable and diagnosing a race condition are not the same job. Here they are separate dials: the model sets raw capability, the reasoning effort sets how long it deliberates. Cheap work goes to a fast model; a hard problem gets the capable one thinking for as long as it needs.

The server runs on your machine and drives the Codex CLI you already have installed; it holds no credentials of its own. Prompts reach OpenAI through Codex, exactly as when you run codex yourself.

See How it compares for a versioned comparison with other Codex MCP servers.

Using it

Write a bounded delegation

A delegation gets expensive when repeated commands keep adding output to the context carried into later requests. Name the exact question, likely files, stopping condition and evidence the answer must contain; choose higher effort for ambiguity rather than by habit. Writing a delegation gives the measured cost model, ranked rules and weak-versus-strong examples using the real tool parameters. Background goes in context and a persona or extra rules in system_instructions, both layered on the built-in quality contract; see the tool reference.

Get a second opinion from a different model family

The value here is not a second run — it is a different set of blind spots.

Ask Codex to review src/server.ts for correctness problems, focusing on error paths. Use a high reasoning effort and tell it to report each finding with the line and why it matters.

A delegated review like this found the terminate() defect in this repository's own runner: two code paths could each arm a timer while only one was ever cleared.

Investigate without spending your context

Forty files go into the delegation; one answer comes back. Codex runs its own searches and reads whatever it needs; your conversation receives the conclusion.

Have Codex trace how a reasoning effort travels from the MCP tool call down to the arguments handed to the Codex CLI, and report just the call chain.

Run long work in the background while you keep going

Kick off a Codex run in the background that writes unit tests for src/jobs.ts, then keep helping me with the API layer.

You get a job_id immediately. Ask for the status whenever you want, and read the result when it is done. Up to eight can run at once.

Buy deep reasoning for one hard problem

Raising the reasoning effort for the whole conversation is expensive. Raising it for one delegation is not.

This intermittent test failure has beaten me twice. Ask Codex to work out the root cause at maximum reasoning effort, give it test/runner.test.ts and the CI log, and tell it not to change anything — I want the diagnosis first.

Keep the thread going

Ask Codex to expand on its second finding.

Follow-ups reuse Codex's context, so they cost a fraction of the original. The server restates the same model, effort and directory on every follow-up because Codex itself does not keep them on resume.

Let it write, when you mean it

Have Codex apply its first two suggestions. Let it edit files, but keep it inside a git worktree so my working tree stays clean.

use_worktree sends the run's edits to ~/.codex/worktrees/; results list the files it touched and where each landed, up to a thousand distinct files, then report the omitted count. The server does not clean worktrees up: they may hold unapplied work. The experimental feature is enabled only for that invocation; your Codex configuration is unchanged. See Safety for the write boundary.

Whatever the sandbox, a delegation that writes reports what it wrote:

Files changed (2):
- [edit] src/codex/runner.ts
- [add] test/runner.test.ts

Safety

This server runs another program on your machine, so it is worth two minutes before you enable writes.

What protects you

Delegations are read-only by default. Writing requires an explicit sandbox: "workspace-write" or a user-set default. use_worktree sends edits to a managed git worktree instead of your checkout; the sandbox and add_dirs bound where it can write at all. Unsandboxed runs are unavailable unless you explicitly opt into that ceiling.

The confinement is the operating system's own sandbox: Seatbelt on macOS, bubblewrap on Linux and WSL2, and a native sandbox on Windows. The table was measured on macOS; SECURITY.md records the measurement details. Linux and Windows have not been measured here, and OpenAI's Windows documentation notes that sandboxed commands can fail to read some directories, so reads may be stricter there:

read-only

workspace-write

danger-full-access

Write inside the working directory

no

yes

yes

Write outside it (your home)

no

no

yes

Network access

no

no

yes

Read outside the working directory

yes

yes

yes

The sandbox confines the commands Codex runs in its shell. It does not confine MCP or plugin tools: those run as their own processes with your account's full permissions, network included, and under read-only a tool that declares itself read-only runs without approval (measured on codex-cli 0.159.2; ADR 16). So a delegation inherits no MCP server, plugin or app from your Codex setup unless you allow it in the server configuration, and every result ends with a line saying what the run was allowed.

There is no shell in the server's invocation path: the CLI is spawned with an argv array and the prompt is written to its stdin, never interpolated into a command string. Shell metacharacters in a prompt are inert.

What does not protect you

Reads are not confined. Codex can read anything your user account can, in every mode — your SSH keys, your cloud credentials. That was measured on macOS, and it is the safe assumption on every platform. Sandboxed network access is blocked so it cannot send them anywhere, but its report comes back to you, and that is a channel.

A prompt is untrusted input, and Codex acts on it. Content you did not write — an issue body, a web page, a log, a file from someone else's repository — can carry instructions. With workspace-write it can direct Codex to modify your repository; even read-only it can direct Codex to read something sensitive and put it in the answer. The sandbox bounds where Codex can write. It does not judge what it should write, or why it was asked. This is prompt injection, and it is the risk that matters here.

The result is not sanitised. What comes back is text from a model that just read your files. Treat it as data, not as instructions; review what a delegation did rather than assuming it did what you asked.

Reducing the risk

  • Leave the built-in default alone. Read-only handles investigation, review and diagnosis, which is most delegation.

  • If you never want writes, cap it: CODEX_SUBAGENT_MAX_SANDBOX=read-only. No conversation can argue past a ceiling. Register it outside the repository (Claude Code's default local scope, --scope user, or Claude Desktop's config), not in a project .mcp.json that a write-enabled delegation could edit. See Configuration.

  • When you enable writes, add use_worktree so changes land somewhere you can inspect before they touch your branch.

  • Allow MCP servers, plugins or apps in delegations only when a task needs them, and by name. An allowed one runs outside the sandbox, and with auto_approve its tools need no approval at all.

  • Do not assemble delegation prompts from untrusted content when you intend to act on the answer.

  • If this threat matters seriously to you, run Codex under an account or container with no access to your secrets. That solves it at the root instead of bounding it.

SECURITY.md has the full threat model, what a deny_read policy could add, and how to report a vulnerability.

Configuration

Everything is optional, and set through environment variables on the MCP server:

Variable

Values

Default

Effect

CODEX_SUBAGENT_DEFAULT_MODEL

Slug from list_codex_models

Unset: refused with a suggestion

Model when a call specifies none.

CODEX_SUBAGENT_DEFAULT_EFFORT

none, minimal, low, medium, high, xhigh, max, ultra

Model's default

Effort when a call specifies none.

CODEX_SUBAGENT_ALLOWED_MODELS

Comma-separated slugs from list_codex_models

Unrestricted

Anything else is refused; a single entry acts as the default model. An empty entry (a,,b, a trailing comma) is a configuration error.

CODEX_SUBAGENT_DEFAULT_SANDBOX

read-only, workspace-write, danger-full-access

read-only

Sandbox when a call specifies none; cannot exceed the ceiling.

CODEX_SUBAGENT_MAX_SANDBOX

read-only, workspace-write, danger-full-access

workspace-write

Calls above it are refused; danger-full-access needs explicit opt-in.

CODEX_SUBAGENT_MAX_EFFORT

none, minimal, low, medium, high, xhigh, max, ultra

Unrestricted

Higher effort is lowered to a supported level, or refused if none fits.

CODEX_SUBAGENT_MCP_SERVERS

none, all, or comma-separated server names

none

MCP servers from Codex's config a delegation keeps; this server is always off.

CODEX_SUBAGENT_PLUGINS

none, all, or comma-separated plugin ids (name@marketplace)

none

Installed Codex plugins a delegation keeps, with the MCP servers they provide.

CODEX_SUBAGENT_APPS

on, off

off

Whether a delegation keeps Codex's apps (connectors to external services).

CODEX_BIN

Executable path

codex on PATH

Override CLI resolution; see Installation.

A model supports a subset of efforts; an unsupported effort is adjusted to the closest supported one with a note. See docs/TOOLS.md#configuration for full semantics and Safety for the sandbox boundary.

Claude decides when to delegate; each run sends its prompt to OpenAI and spends your Codex usage, in any conversation where the server is available. docs/CONTROL.md covers client permission prompts, server ceilings, version ranges and your own CLAUDE.md escalation rules. Those rules stay yours; see ADR 12 and ADR 14 for the policy split.

Choosing a model

Use list_codex_models for the live catalog from your installed CLI. Model and reasoning effort are independent: the model sets raw capability; the effort sets how long it deliberates. ultra additionally delegates subtasks automatically.

The server does not choose for you. A task's model depends on your budget and how costly a wrong answer is. Without a model in the call or configuration, it refuses with a recommendation for you to decide on.

With no default, tool descriptions tell Claude to call codex_recommend first, announce the suggested model and effort, then delegate with both explicit. This is guidance to a model, not enforcement. For advice, ask: “Which Codex model should handle migrating this repo's tests to vitest?” The suggestion respects your model allow-list and effort ceiling; see the recommendation reference.

To skip that step on later delegations, set a default from list_codex_models once:

claude mcp add codex-subagent -e CODEX_SUBAGENT_DEFAULT_MODEL=gpt-5.6-terra -- npx -y codex-subagent-mcp

Tools

Tool

What it does

codex_doctor

Check the Codex CLI installation and report how to fix it.

list_codex_models

List available models and their reasoning-effort levels.

codex_recommend

Suggest a model and effort for a described task.

codex_delegate

Run a task, blocking or in the background, optionally returning JSON that matches a schema.

codex_follow_up

Continue a previous delegation using its thread_id, with or without a schema.

codex_job_status

Check a background delegation.

codex_job_result

Read a finished background delegation's output.

codex_job_cancel

Stop a running background delegation.

Full parameter reference: docs/TOOLS.md.

FAQ

Does this cost money? It uses your existing Codex quota, as running codex yourself does. This server adds nothing and has no visibility into the cost. Higher efforts consume more, and ultra also delegates subtasks; codex_recommend helps avoid spending ultra on low work.

Why drive the CLI instead of calling the OpenAI API? Delegated coding is not a single completion — it is an agentic loop with a sandbox, an approval model, session persistence and project instruction files. All of that lives in the Codex client, not in the model endpoint. See ADR 1.

Do I need Claude Code, or does Claude Desktop work? Either. Claude Code gets a one-line install; Claude Desktop needs a manual config entry.

It says Codex is not installed, but codex works in my terminal. Most likely Windows with a global npm install, which produces a codex.cmd batch shim that cannot be launched without a command shell. codex_doctor reports this as unsupported-shim and offers two fixes. On macOS and Linux, check whether a Node version manager moved codex off PATH.

Does it work on Windows and Linux? CI builds, tests and starts the server on Windows, macOS and Linux on every change, and checks that the Codex CLI is resolved correctly on each. A real delegation has only been verified on macOS — the CI runners have no Codex installation or credentials. Reports from Windows and Linux are welcome. On Windows, install the Codex CLI with the PowerShell installer rather than npm; on Linux, install bubblewrap for Codex's sandbox. See Installation.

Documentation

Disclaimer

Not an official product. This is an independent, community project. It is not affiliated with, endorsed by, sponsored by or supported by OpenAI or Anthropic. "Codex", "ChatGPT" and "OpenAI" are trademarks of OpenAI; "Claude" and "Claude Code" are trademarks of Anthropic. They are used here only to describe what this software interoperates with, which is nominative use — no claim is made to any of them. Neither company is responsible for this software, and problems with it should be reported here rather than to them.

No warranty. The software is provided "as is", without warranty of any kind, as stated in LICENSE. You use it at your own risk.

Read Safety before enabling writes; delegations spend your own Codex usage.

License

MIT. See LICENSE.

Available Tools

8 tools
codex_delegateDelegate a task to CodexA

Delegate a task to the local Codex CLI (OpenAI's coding agent), choosing model and reasoning effort. Use it when the user asks for Codex, or when handing work off clearly serves their request: a second opinion from a different model family, or an investigation that would otherwise flood this conversation. When the user has not named a model, call codex_recommend first and present its suggested model and effort to the user in the same message in which you say you are going to delegate, then pass both explicitly here. That recommendation is advice for an already-authorised delegation, not a replacement for the user's own preference. Everything passed in prompt, context and target_files is sent to OpenAI, and every run spends the user's own Codex usage, so do not delegate what you can answer directly, and tell the user when you delegate. Codex runs read-only unless a different default sandbox is configured. Set sandbox to workspace-write to let it edit files. Codex cannot see this conversation, so pass everything it needs in prompt, context, and target_files.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNoblocking (default) waits and streams progress; background returns a job_id immediately.
modelNoCatalog slug from list_codex_models. If omitted, the configured default is used; with no default configured the call is refused and the recommended model is returned.
promptYesThe task for Codex. Be specific and self-contained: Codex cannot see this conversation.
contextNoBackground Codex needs: prior findings, constraints, relevant excerpts.
sandboxNoSandbox policy. Uses the configured default when omitted; without one, Codex runs read-only.
add_dirsNoAdditional absolute directories that should be writable alongside working_dir.
web_searchNoEnable Codex's API-backed live web-search tool for this run. In a read-only sandbox, shell commands have no network access, so this is the route to current external information. When omitted, Codex's own configured web_search mode applies.
working_dirNoAbsolute path Codex uses as its working root.
auto_approveNoAdds --approve-for-me so Codex auto-approves its own commands. Only applies when sandbox is workspace-write.
target_filesNoPaths Codex should focus on, relative to working_dir.
use_worktreeNoRun in a managed git worktree. Writes outside it remain subject to the sandbox policy and add_dirs.
output_schemaNoA JSON Schema object for this turn's final message, which then comes back as JSON in a delimited block. OpenAI's structured outputs apply: every object needs "additionalProperties": false and every property listed in "required". At most 64 KiB serialised.
timeout_secondsNoWall-clock budget. Defaults to 1800s.
reasoning_effortNoReasoning depth, independent of model choice. Uses the configured default or the model's default when omitted. Clamped to supported levels within the configured ceiling; refused if none qualify.
acceptance_criteriaNoConcrete conditions that must hold for the task to be considered done.
skip_git_repo_checkNoAllow running outside a git repository.
system_instructionsNoPersona or extra rules inherited from the orchestrator, layered on the built-in quality contract.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false and openWorldHint=true, but the description adds substantial non-obvious context: prompt/context/target_files are sent to OpenAI, each run spends the user's own Codex usage, the user must be told when delegating, and the default sandbox is read-only unless reconfigured. This is exactly the kind of disclosure structured fields cannot carry.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core purpose and usage trigger are front-loaded, and for a 17-parameter tool the length is defensible. It is a dense single block with some redundancy (the 'Codex cannot see this conversation' constraint is echoed by the schema), but each sentence carries operational weight.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a high-complexity tool with 17 parameters, no output schema, and a nested output_schema parameter, the description covers delegation conditions, privacy, cost, sandbox behavior, and the recommendation handoff. Return behavior (job_id for background, structured output block) is left to schema descriptions rather than explained, which is the main remaining gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real meaning beyond the schema: the model parameter is tied to a codex_recommend workflow, sandbox read-only is the default and workspace-write is required for edits, and target_files/context/prompt are flagged as externally transmitted. This lifts it above the schema-only baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (delegate) and resource (a task to the local Codex CLI, described as OpenAI's coding agent), with additional scoping that distinguishes it from siblings by naming codex_recommend as the prerequisite recommendation tool. An agent can tell exactly what this does and how it differs from list_codex_models and codex_recommend.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use triggers ('when the user asks for Codex, or when handing work off clearly serves their request: a second opinion from a different model family, or an investigation that would otherwise flood this conversation') and an explicit when-not ('do not delegate what you can answer directly'). It also names the alternative workflow (codex_recommend first when no model is chosen), leaving nothing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

codex_doctorCheck the Codex CLI installationA
Read-only

Check whether the local Codex CLI is installed, recent enough, signed in and able to load its configuration, and report the exact steps to fix it if not. Run this when any other tool reports the CLI is unavailable, or before relying on delegation for the first time. It only inspects the installation; it never installs or changes anything.

ParametersJSON Schema
NameRequiredDescriptionDefault
refreshNoRe-probe the CLI instead of reusing the cached diagnosis.
working_dirNoAbsolute directory to run the check in. Codex loads the configuration of the directory it runs in, so pass the one a delegation would use. Defaults to this server's own.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already declare readOnlyHint=true and openWorldHint=false, and the description's 'never installs or changes anything' reinforces that safety profile without contradicting it. It adds useful behavioral scope by explaining the inspection covers signed-in state and configuration loading, and that it emits remediation steps rather than performing fixes. It doesn't detail the exact return format or mention the cached diagnosis behavior described in the refresh parameter, but the schema covers that.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, no fluff, and the most decision-relevant facts come first: what it checks, when to run it, and what it will not do. Every sentence earns its place, and the coverage of purpose, usage, and safety is achieved in extremely compact form.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-required-parameter diagnostic tool, the definition is complete: it states the checks performed, what the output will contain (fix steps), when to invoke it, and that it is read-only. The schema covers the two optional parameters, and the annotations cover the side-effect profile, so nothing needed for correct selection or invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and both parameters already have strong descriptions: refresh explains the cache/re-probe behavior)Skip and working_dir explains the directory-sensitive configuration loading and its default. The tool description contributes contextual motivation by tying the check to delegation, but it does not add meaning beyond what the schema already provides. A baseline 3 is appropriate because the schema carries the parameter documentation burden.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb ('check') and a concrete resource (the local Codex CLI installation), then elaborates with four concrete checks: installed, recent enough, signed in, and able to load configuration. It also states the actual outcome ('report the exact steps to fix it'), which differentiates this diagnostic tool from the delegation, recommendation, and job-management siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit trigger conditions: run when another tool reports the CLI is unavailable, or before relying on delegation for the first time. It also states a clear exclusion—it only inspects and never installs or changes anything—so an agent knows not to use it as a fixer and should look elsewhere for remediation.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

codex_follow_upContinue a Codex sessionA

Send a follow-up message to a previous delegation using its thread_id. Codex retains the earlier context, so only the new instruction needs to be sent. Like a delegation, it is sent to OpenAI and spends the user's Codex usage.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelNoOverride the model for this turn. Defaults to the thread's last model from memory or, on a registry miss, Codex's session file; if neither has it, the configured default, else the call is refused.
promptYesThe follow-up instruction.
sandboxNoSandbox policy for this turn. Uses the configured default when omitted; read-only when unset.
thread_idYesThe thread_id reported by a previous codex_delegate call.
working_dirNoAbsolute directory to resume in. Defaults to the directory the thread last ran in.
auto_approveNoNot supported on follow-ups: true is refused and nothing runs.
output_schemaNoA JSON Schema object for this turn's final message, which then comes back as JSON in a delimited block. OpenAI's structured outputs apply: every object needs "additionalProperties": false and every property listed in "required". At most 64 KiB serialised.
timeout_secondsNo
reasoning_effortNoOverride the reasoning effort for this turn. Defaults to the thread's last effort when the model is unchanged, otherwise to the configured or model default.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=false and openWorldHint=true, so the description carries the rest: the call goes to OpenAI, spends the user's Codex usage, and the thread is stateful (earlier context is reused). That is meaningful disclosure beyond structured fields, though sync/async behavior is left unstated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with who/what/how, followed by the stateful-context point and then the cost warning. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 9-parameter, stateful, externally-billed call with no output schema, the description covers the essentials: thread linkage, context reuse, and cost. It omits whether the call blocks or returns a job handle (relevant given codex_job_status/codex_job_result siblings), which is the main remaining gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 89%, which already sets a baseline of 3, and the description genuinely adds prompt semantics: Codex retains earlier context, so only the new instruction needs to be sent. That is meaning beyond the schema's 'The follow-up instruction.'

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (send a follow-up message) plus the exact resource and linking key (a previous delegation, via its thread_id). The phrase 'previous delegation' implicitly separates it from codex_delegate, which starts a new one, so an agent can route correctly without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Send a follow-up ... using its thread_id' implies the precondition (you must already hold a thread_id from codex_delegate), and the cost sentence signals weight. However, no alternative is named and no explicit when-not condition is given, so usage is inferred rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

codex_job_cancelCancel a background delegationB
Destructive

Terminate a running background delegation.

ParametersJSON Schema
NameRequiredDescriptionDefault
job_idYesThe job_id returned by codex_delegate.

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and readOnlyHint=false, clearly indicating a destructive operation. The description adds 'running' as a state constraint, which is useful. However, it omits key behavioral details like whether the job is killed immediately or gracefully, whether it can be undone, or what happens to any partial results. With annotations covering the safety profile, a 3 is appropriate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence with no wasted words. It is front-loaded with the action, though it is extremely terse, which might leave some ambiguity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (one parameter, no output schema) and the annotations covering destructiveness, the description is minimally adequate but lacks important context about side effects, job state requirements, and the return value or confirmation. It should do more to help an agent understand when and how to use it safely.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the single parameter job_id is fully documented in the schema, including that it comes from codex_delegate. The description adds no additional parameter meaning beyond what the schema provides, making the baseline 3 correct.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Terminate) and resource (running background delegation), which is clearer than the title. However, it does not distinguish itself from siblings like codex_job_status or codex_job_result, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives, or when not to use it. The description gives no context about the conditions under which cancellation is appropriate, such as whether the job must be running or what happens if it has already completed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

codex_job_resultRead a background delegation's resultA
Read-only

Return the full output of a finished background delegation. Errors if the job is still running.

ParametersJSON Schema
NameRequiredDescriptionDefault
job_idYesThe job_id returned by codex_delegate.

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds a meaningful behavioral constraint: it errors when the job is still running, telling the agent this is a post-completion read with a specific failure mode. It does not discuss output truncation or pagination, so 4.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, zero waste, and the completion precondition is front-loaded. The error condition follows immediately after the main action, keeping the most important routing information prominent.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter read tool with full schema coverage and annotations covering safety, the description supplies the necessary completion precondition and output scope. No output schema exists, so describing the return as 'full output' is appropriately complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the only parameter 'job_id' is documented as 'returned by codex_delegate'. The description adds no additional syntax, format, or constraint beyond the schema, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Return') and resource ('full output of a finished background delegation'), and the error condition distinguishes it from a status check. It does not explicitly name sibling codex_job_status as the alternative, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Implies use only after completion via 'finished background delegation' and 'Errors if the job is still running'. However, it does not explicitly tell the agent to call codex_job_status first or name alternatives, so usage is inferable but not spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

codex_job_statusCheck a background delegationA
Read-only

Report the state and recent activity of a background delegation started with mode=background. Call it with no job_id to list every known job.

ParametersJSON Schema
NameRequiredDescriptionDefault
job_idNoOmit to list all jobs.
include_activityNoInclude the recent progress log for the job.

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds that it reports state plus recent activity and that omitting job_id enumerates all jobs, but says nothing about pagination, log volume, or failure modes for unknown job ids.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the core purpose and with the no-arg listing behavior second; no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-required-param, read-only status tool with no output schema, the description covers what is queried (state and activity) and the two calling modes. It is nearly complete, with only return-shape hints and error handling left unspecified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and both parameters are already documented in the schema; the description's note about omitting job_id restates that documentation rather than adding format or constraint detail.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (report) plus resource (state and recent activity of a background delegation) and scopes it to delegations started with mode=background, which links it to codex_delegate and separates it from codex_job_result.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives one concrete usage rule — call with no job_id to list every known job — but never says when to prefer this over codex_job_result, codex_follow_up, or codex_job_cancel, so the sibling routing is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

codex_recommendRecommend a Codex model and effortA
Read-only

Given a task description, recommend which Codex model and reasoning effort to delegate it with. Runs no model call; applies a documented matrix reconciled against the installed catalog.

ParametersJSON Schema
NameRequiredDescriptionDefault
priorityNoBias the reasoning effort: quality raises it, latency and cost lower it. Default balanced.
working_dirNoAbsolute directory whose Codex configuration and model catalog should be used. Defaults to this server's own.
task_descriptionYesWhat the delegated task involves, in one or two sentences.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnlyHint and openWorldHint annotations, the description discloses that the tool executes no model call and applies a documented matrix reconciled against the installed catalog. This adds meaningful context about how the recommendation is derived (deterministically, not by running a model) without conflicting with the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The two-sentence description is compact and front-loaded: it immediately states the input ('Given a task description') and the output ('recommend which Codex model and reasoning effort'). The second sentence clarifies a key behavioral trait without waste. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple tool with 3 parameters (1 required) and no output schema, the description is adequately complete. It specifies the recommendation output (model and effort) but does not detail the exact return structure (e.g., object or string). Given the low complexity and clear sibling context (codex_delegate, list_codex_models), this is sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

All three parameters are fully described in the input schema (100% coverage), so the description adds no parameter-specific detail beyond the schema. With full schema coverage, the baseline of 3 is appropriate; the description does not compensate with extra parameter guidance.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'recommend which Codex model and reasoning effort to delegate it with.' It also notes 'Runs no model call,' which clearly differentiates this advisory tool from execution tools like codex_delegate. The purpose is unambiguous and distinct from siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage ('Given a task description... recommend') but does not explicitly say when to use this tool versus alternatives like codex_delegate or list_codex_models. The phrase 'Runs no model call' hints it is not for execution, but there is no explicit 'Use this before delegating' or 'For execution, use codex_delegate.' Guidance is implied, not stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_codex_modelsList Codex modelsA
Read-only

List the Codex models available on this machine, with the reasoning-effort levels each one supports. Read from the installed Codex CLI, with a warned static fallback if its catalog cannot be read. Call this before codex_delegate when choosing a model explicitly.

ParametersJSON Schema
NameRequiredDescriptionDefault
refreshNoBypass the cache and re-read the catalog from the CLI.
working_dirNoAbsolute directory to read the catalog in. A project you have trusted in Codex can set its own catalog, so pass the directory a delegation would use. Defaults to this server's own.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark the tool read-only, and the description adds useful behavioral context: it reads from the installed Codex CLI and has a 'warned static fallback' if the catalog cannot be read. This goes beyond the annotation in a meaningful way, though caching behavior is left to the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, each earning its place: the core purpose, the fallback behavior, and the usage guidance. It is front-loaded with the essential information and contains no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only listing tool, the description covers the key context: what is listed, where it is read from, the fallback behavior, and when to call it. It does not describe the return format, but no output schema exists and the purpose makes the return shape largely predictable.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema fully documents both parameters. The description does not add parameter-level meaning beyond the schema, which makes the baseline 3 appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific action ('List'), a precise resource ('Codex models available on this machine'), and additional detail ('reasoning-effort levels each one supports'). It also distinguishes itself from a sibling by pointing to codex_delegate, so an agent can tell why this tool exists.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit usage context: 'Call this before codex_delegate when choosing a model explicitly.' It clearly states when the tool should be used, though it does not name alternatives to avoid or describe when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 2 tool updatesv0.4.0
    • Changedcodex_delegate1 field changed
      • addedInput schema / properties / output_schema
        Added value: +{
        +  "additionalProperties": {},
        +  "description": "A JSON Schema object for this turn's final message, which then comes back as JSON in a delimited block. OpenAI's structured outputs apply: every object needs \"additionalProperties\": false and every property listed in \"required\". At most 64 KiB serialised.",
        +  "propertyNames": {
        +    "type": "string"
        +  },
        +  "type": "object"
        +}
    • Changedcodex_follow_up1 field changed
      • addedInput schema / properties / output_schema
        Added value: +{
        +  "additionalProperties": {},
        +  "description": "A JSON Schema object for this turn's final message, which then comes back as JSON in a delimited block. OpenAI's structured outputs apply: every object needs \"additionalProperties\": false and every property listed in \"required\". At most 64 KiB serialised.",
        +  "propertyNames": {
        +    "type": "string"
        +  },
        +  "type": "object"
        +}
  2. 5 tool updatesv0.3.0
    • Changedcodex_delegate2 fields changed
      • changedInput schema / properties / sandbox / description
        Previous value: -"Sandbox policy. Defaults to read-only: Codex analyses and reports but cannot modify files."New value: +"Sandbox policy. Uses the configured default when omitted; without one, Codex runs read-only."
      • changedInput schema / properties / web_search / description
        Previous value: -"Enable live web search for this run, through Codex's web_search = \"live\" setting. When omitted, Codex's own configured web_search mode applies."New value: +"Enable Codex's API-backed live web-search tool for this run. In a read-only sandbox, shell commands have no network access, so this is the route to current external information. When omitted, Codex's own configured web_search mode applies."
    • Changedcodex_doctor1 field changed
      • addedInput schema / properties / working_dir
        Added value: +{
        +  "description": "Absolute directory to run the check in. Codex loads the configuration of the directory it runs in, so pass the one a delegation would use. Defaults to this server's own.",
        +  "type": "string"
        +}
    • Changedcodex_follow_up2 fields changed
      • changedInput schema / properties / model / description
        Previous value: -"Override the model for this turn. Defaults to the model the thread last ran with on this server; for a thread this server has no record of, the configured default, else the call is refused."New value: +"Override the model for this turn. Defaults to the thread's last model from memory or, on a registry miss, Codex's session file; if neither has it, the configured default, else the call is refused."
      • changedInput schema / properties / sandbox / description
        Previous value: -"Sandbox policy for this turn. Defaults to read-only."New value: +"Sandbox policy for this turn. Uses the configured default when omitted; read-only when unset."
    • Changedcodex_recommend1 field changed
      • addedInput schema / properties / working_dir
        Added value: +{
        +  "description": "Absolute directory whose Codex configuration and model catalog should be used. Defaults to this server's own.",
        +  "type": "string"
        +}
    • Changedlist_codex_models1 field changed
      • addedInput schema / properties / working_dir
        Added value: +{
        +  "description": "Absolute directory to read the catalog in. A project you have trusted in Codex can set its own catalog, so pass the directory a delegation would use. Defaults to this server's own.",
        +  "type": "string"
        +}
  3. 2 tool updatesv0.2.0
    • Changedcodex_delegate5 fields changed
      • changedInput schema / properties / auto_approve / description
        Previous value: -"Adds --approve-for-me so Codex auto-approves its own commands. Only applies when sandbox allows writes."New value: +"Adds --approve-for-me so Codex auto-approves its own commands. Only applies when sandbox is workspace-write."
      • changedInput schema / properties / model / description
        Previous value: -"Catalog slug from list_codex_models. Omitted means the recommendation matrix picks one."New value: +"Catalog slug from list_codex_models. If omitted, the configured default is used; with no default configured the call is refused and the recommended model is returned."
      • changedInput schema / properties / reasoning_effort / description
        Previous value: -"Reasoning depth, independent of model choice. Clamped to what the chosen model supports."New value: +"Reasoning depth, independent of model choice. Uses the configured default or the model's default when omitted. Clamped to supported levels within the configured ceiling; refused if none qualify."
      • changedInput schema / properties / use_worktree / description
        Previous value: -"Run in a managed git worktree so changes never touch the current working tree."New value: +"Run in a managed git worktree. Writes outside it remain subject to the sandbox policy and add_dirs."
      • changedInput schema / properties / web_search / description
        Previous value: -"Enable Codex's native web search tool."New value: +"Enable live web search for this run, through Codex's web_search = \"live\" setting. When omitted, Codex's own configured web_search mode applies."
    • Changedcodex_follow_up4 fields changed
      • addedInput schema / properties / auto_approve / description
        Added value: +"Not supported on follow-ups: true is refused and nothing runs."
      • changedInput schema / properties / model / description
        Previous value: -"Override the model for this turn."New value: +"Override the model for this turn. Defaults to the model the thread last ran with on this server; for a thread this server has no record of, the configured default, else the call is refused."
      • changedInput schema / properties / reasoning_effort / description
        Previous value: -"Override the reasoning effort for this turn."New value: +"Override the reasoning effort for this turn. Defaults to the thread's last effort when the model is unchanged, otherwise to the configured or model default."
      • addedInput schema / properties / working_dir / description
        Added value: +"Absolute directory to resume in. Defaults to the directory the thread last ran in."
  4. 1 tool update
    • Changedcodex_follow_up1 field changed
      • addedInput schema / properties / thread_id / pattern
        Added value: +"^[A-Za-z0-9][A-Za-z0-9_-]{0,199}$"
  5. 8 tool updatesv0.1.0
    • First observedcodex_delegate
    • First observedcodex_doctor
    • First observedcodex_follow_up
    • First observedcodex_job_cancel
    • First observedcodex_job_result
    • First observedcodex_job_status
    • First observedcodex_recommend
    • First observedlist_codex_models

TDQS

A3.9/5.0

Scored across 8 tools

Disambiguation4/5

The job-lifecycle tools (cancel, result, status) and the delegation tools (delegate vs follow_up) are clearly distinguished by their descriptions. Only codex_recommend and list_codex_models overlap somewhat in the model-selection space, but one gives advice and the other enumerates the catalog, so confusion is limited.

Naming Consistency4/5

Seven of eight tools use a predictable codex_<action> or codex_job_<action> snake_case pattern. list_codex_models deviates by inverting to verb-first without the codex_ prefix, a minor but noticeable break in the pattern.

Tool Count5/5

Eight tools is well-scoped for a CLI delegation wrapper: one core delegate tool plus the supporting lifecycle (cancel, result, status), continuation (follow_up), and setup helpers (doctor, recommend, list_codex_models). Each tool earns its place with no redundancy.

Completeness4/5

The surface covers the full delegation lifecycle: starting, continuing, monitoring, retrieving, and cancelling jobs, plus environment diagnostics and model discovery. Only minor gaps exist, such as no explicit job cleanup/removal or batch operations, which agents can work around.

Maintenance

ActivityMaintained
ResponsivenessResponsive

Related MCP Connectors

Related MCP Servers