Skip to main content
Glama

codex-subagent-mcp

CI License: MIT

An MCP server that lets Claude Code delegate coding tasks to OpenAI's Codex CLI running on the same machine — multi-model orchestration, locally, with the model and reasoning depth chosen per task.

Claude stays the orchestrator. Codex becomes a subagent it can call.

This runs another agent on your machine. Codex can read local files and run commands with the permissions you grant it. Read the security model before enabling writes or unsandboxed runs.

An independent project. Not affiliated with, endorsed by, or supported by OpenAI or Anthropic.

Why this exists

A single model doing everything has three recurring problems, and delegation solves each one:

Your context window is finite. Having Claude read forty files to answer one question spends context you need for the actual work. Delegating the investigation returns the answer instead of the forty files.

One model has one set of blind spots. A second opinion is worth most when it comes from a different model family — different training, different failure modes. Asking the same model twice mostly gets you the same answer twice.

Not every task deserves the same reasoning budget. Renaming a variable and diagnosing a race condition are not the same job. Here they are separate dials: the model sets raw capability, the reasoning effort sets how long it deliberates. Cheap work goes to a fast model; a hard problem gets the capable one thinking for as long as it needs.

Everything stays on your machine. The server drives the Codex CLI you already have installed and holds no credentials of its own.

Related MCP server: codex-mcp

Requirements

  • Node.js 22 or newer.

  • The Codex CLI, installed, on PATH, and signed in.

You do not have to check this by hand. Run the codex_doctor tool — or just ask Claude to — and it reports what is missing and the exact commands for your platform. Every tool that reaches the CLI runs the same check first, so you never get a bare spawn ENOENT. Nothing is ever installed on your behalf.

If you do not have the Codex CLI yet, install it without npm:

# macOS — recommended
brew install --cask codex
# macOS / Linux — standalone installer
curl -fsSL https://chatgpt.com/codex/install.sh | sh
# Windows (PowerShell)
powershell -ExecutionPolicy ByPass -c "irm https://chatgpt.com/codex/install.ps1 | iex"

Then run codex once to sign in, and confirm with codex login status. On Windows, open a new terminal first so the updated PATH is picked up.

Codex's sandbox depends on the platform, so two notes from OpenAI's documentation:

  • Linux and WSL2 — Codex sandboxes commands with bubblewrap. Install it with your package manager before the first delegation. Without it Codex falls back to a bundled helper that needs unprivileged user namespaces, which some distributions restrict. See sandboxing.

  • Windows — Codex runs natively, without WSL, and uses its own Windows sandbox. Windows 11 is recommended; Windows 10 version 1809 or newer is the practical minimum. See the Windows sandbox documentation.

The npm package is not the program: bin/codex.js is a Node wrapper that spawns the real Rust binary. That has two consequences.

Every invocation pays a Node startup. Measured on macOS: about 80 ms through the wrapper against about 20 ms calling the binary directly. This server spawns the CLI once per tool call, so the cost recurs — though it is still noise next to a delegation that runs for seconds.

A global npm install lives inside the active Node version. Under a version manager such as nvm it lands in ~/.nvm/versions/node/<version>/lib/node_modules, so switching Node versions takes codex off PATH until you reinstall it. This is the bigger problem in practice.

On Windows there is a third, harder consequence: a global npm install produces a codex.cmd batch shim, which cannot be launched without a command shell — and this server never uses one. It detects that case and says so, but the installer avoids it entirely. See ADR 11.

Switching is two commands, and your sign-in survives because credentials live in Codex's home directory — ~/.codex, or %USERPROFILE%\.codex on native Windows — not in the npm package:

# macOS
npm uninstall -g @openai/codex && brew install --cask codex
# Linux
npm uninstall -g @openai/codex && curl -fsSL https://chatgpt.com/codex/install.sh | sh
# Windows (PowerShell) — two lines, because Windows PowerShell rejects `&&`
npm uninstall -g @openai/codex
powershell -ExecutionPolicy ByPass -c "irm https://chatgpt.com/codex/install.ps1 | iex"

Install

claude mcp add codex-subagent -- npx -y codex-subagent-mcp

That works in both the Claude Code CLI and the desktop app; they share the same configuration.

Global install, if you prefer not to go through npx:

npm install -g codex-subagent-mcp

Then point Claude Code at the codex-subagent binary.

Claude Desktop has no equivalent command. Add an entry to mcpServers in claude_desktop_config.json and restart the app:

{
  "mcpServers": {
    "codex-subagent": {
      "command": "npx",
      "args": ["-y", "codex-subagent-mcp"]
    }
  }
}

Settings → Developer → Edit Config opens the file and creates it if it does not exist. Its documented locations are ~/Library/Application Support/Claude/claude_desktop_config.json on macOS and %APPDATA%\Claude\claude_desktop_config.json on Windows. The Linux desktop app is in beta, and its documentation does not say where the file lives. If the server does not appear after a restart, the MCP logs are in ~/Library/Logs/Claude on macOS and %APPDATA%\Claude\logs on Windows. See Connect to local MCP servers.

From a clone, for development:

git clone https://github.com/parisbs/codex-subagent-mcp.git
cd codex-subagent-mcp && npm ci && npm run build

The repository ships a .mcp.json, so running Claude Code from the project root picks the server up.

Set CODEX_BIN if your Codex executable is not called codex or is not on PATH. On Windows, point it at the real codex.exe: a .cmd or .bat shim is refused rather than run through a shell.

First steps

Ask Claude to check the installation:

Check that the Codex subagent is set up correctly.

You should see status: ok, a version, and signed in: yes. Then see what you can delegate to:

What Codex models are available, and what are they each good for?

Then try a real one. With the built-in default this is read-only, so Codex investigates and reports without touching anything:

Have Codex look at this repository and explain how the build is wired together.

If you have not set a default model, the tool descriptions tell Claude to call codex_recommend first, to give you the suggested model and effort in the same message in which it says it is going to delegate, and then to call codex_delegate with both values explicit. Those descriptions are guidance to a model rather than a rule it cannot break, and a delegation that reaches this server with no model at all is refused rather than guessed at. The recommendation is advice, not a decision on your behalf. To skip that step on later delegations, set CODEX_SUBAGENT_DEFAULT_MODEL once as shown in Choosing a model.

Using it

Delegations run read-only by default: Codex investigates and reports, but cannot modify files. Letting it write is a deliberate call argument or a default you set in the server environment.

Write a bounded delegation

A delegation gets expensive when repeated commands keep adding output to the context carried into later requests. Name the exact question, likely files, stopping condition and evidence the answer must contain; choose higher effort for ambiguity rather than by habit. Writing a delegation gives the measured cost model, ranked rules and weak-versus-strong examples using the real tool parameters.

The examples below are the four situations where delegating beats doing it in the main conversation. Each one has been run against the real Codex CLI — writing them is how two defects in this server were found and fixed.

Get a second opinion from a different model family

The value here is not a second run — it is a different set of blind spots.

Ask Codex to review src/server.ts for correctness problems, focusing on error paths. Use a high reasoning effort and tell it to report each finding with the line and why it matters.

Claude picks the model, passes your file as the focus, and returns the findings. This is how the terminate() defect in this repository's own runner was found: a delegated review spotted that two code paths could each arm a timer while only one was ever cleared.

Investigate without spending your context

Forty files go into the delegation; one answer comes back.

Have Codex trace how a reasoning effort travels from the MCP tool call down to the arguments handed to the Codex CLI, and report just the call chain.

Codex runs its own searches and reads whatever it needs. Your conversation receives the conclusion, not the search results.

Run long work in the background while you keep going

Kick off a Codex run in the background that writes unit tests for src/jobs.ts, then keep helping me with the API layer.

You get a job_id immediately. Ask for the status whenever you want, and read the result when it is done. Up to eight can run at once.

Buy deep reasoning for one hard problem

Raising the reasoning effort for the whole conversation is expensive. Raising it for one delegation is not.

This intermittent test failure has beaten me twice. Ask Codex to work out the root cause at maximum reasoning effort, give it test/runner.test.ts and the CI log, and tell it not to change anything — I want the diagnosis first.

Keep the thread going

Follow-ups reuse the context Codex already has, so they cost a fraction of the original:

Ask Codex to expand on its second finding.

The follow-up runs on the same model, effort and directory as the original. Codex itself does not keep those when a session resumes, so the server restates them for every thread it started.

Let it write, when you mean it

Have Codex apply its first two suggestions. Let it edit files, but keep it inside a git worktree so my working tree stays clean.

That last clause matters: use_worktree sends the run's own edits to a managed git worktree under ~/.codex/worktrees/ instead of your checkout, and the result lists the files it touched with the path where each landed — up to a thousand distinct files, after which it says how many it left out. It is not a second sandbox: what a run may write outside that worktree is still decided by the sandbox and by add_dirs. Worktrees rely on an experimental Codex feature, which the server turns on for that invocation only — it never changes your Codex configuration.

Whatever the sandbox, a delegation that writes reports what it wrote:

Files changed (2):
- [edit] src/codex/runner.ts
- [add] test/runner.test.ts

Before you start enabling writes as a habit, read the next section. It is short.

Safety

This server runs another program on your machine, so it is worth two minutes before you enable writes.

What protects you

Delegations are read-only by default. Writing requires an explicit sandbox: "workspace-write" or a user-set default, and use_worktree sends the run's edits to a managed git worktree instead of your checkout — while the sandbox and add_dirs, not the worktree, are what bound where it can write at all. Unsandboxed runs are unavailable unless you explicitly opt into that ceiling.

The confinement is not a promise from the model — it is the operating system's own sandbox: Seatbelt on macOS, bubblewrap on Linux and WSL2, and a native sandbox on Windows. The table was measured on macOS against Codex CLI 0.154.0. Linux and Windows have not been measured here, and on Windows OpenAI's documentation notes that sandboxed commands can fail to read some directories, so reads may be stricter there:

read-only

workspace-write

danger-full-access

Write inside the working directory

no

yes

yes

Write outside it (your home)

no

no

yes

Network access

no

no

yes

Read outside the working directory

yes

yes

yes

There is also no shell anywhere in the path: the CLI is spawned with an argv array and the prompt is written to its stdin, never interpolated into a command string. Shell metacharacters in a prompt are inert.

What does not protect you

Reads are not confined. That last row is not a typo. Codex can read anything your user account can, in every mode — your SSH keys, your cloud credentials. That was measured on macOS, and it is the safe assumption on every platform. Network access is blocked so it cannot send them anywhere, but its report comes back to you, and that is a channel.

A prompt is untrusted input, and Codex acts on it. This is prompt injection, and it is the risk that matters here. If you build a delegation from content you did not write — an issue body, a web page, a log, a file from someone else's repository — that content can carry instructions. With workspace-write it can direct Codex to modify your repository; even read-only it can direct Codex to read something sensitive and put it in the answer. The sandbox bounds where Codex can write. It does not judge what it should write, or why it was asked.

The result is not sanitised. What comes back is text from a model that just read your files. Treat it as data, not as instructions.

Reducing the risk

  • Leave the built-in default alone. Read-only handles investigation, review and diagnosis, which is most delegation.

  • If you never want writes from this server, cap it: CODEX_SUBAGENT_MAX_SANDBOX=read-only. A ceiling cannot be argued past by anything in the conversation, which is what makes it different from a default. Register it outside the repository (Claude Code's default local scope, --scope user, or Claude Desktop's config), not in a project .mcp.json that a write-enabled delegation could edit. See Configuration.

  • When you do enable writes, add use_worktree so changes land somewhere you can inspect before they touch your branch.

  • Do not assemble delegation prompts from untrusted content when you intend to act on the answer.

  • If this threat matters seriously to you, run Codex under an account or container with no access to your secrets. That solves it at the root instead of bounding it.

SECURITY.md has the full threat model, what a deny_read policy could add, and how to report a vulnerability.

Staying in control

Claude decides when to delegate, and every delegation sends its prompt to OpenAI and spends your Codex usage — in any conversation where the server is available, not only programming ones. The defaults are safe, and you can tighten them in layers:

  • Your client's permission prompt. Let the inspection tools run freely, and keep confirming codex_delegate and codex_follow_up, the two that spend usage.

  • Ceilings on the server, such as CODEX_SUBAGENT_MAX_SANDBOX and CODEX_SUBAGENT_MAX_EFFORT, which no argument can get past.

  • A version range such as codex-subagent-mcp@^0.3.0, so new behaviour arrives when you choose.

  • Your own rules in CLAUDE.md, for when Claude should delegate at all.

docs/CONTROL.md shows how to set each one, what the server already does on its own, and what no setting can guarantee.

Choosing a model

Read live from your installed CLI, so this list tracks whatever you have. As of Codex CLI 0.154.0:

Slug

Positioning

Reasoning efforts

Default

gpt-6-astra

Most capable, for complex demanding work

low … ultra

low

gpt-5.6-sol

Reliable agentic workhorse

low … ultra

low

gpt-5.6-terra

Balanced everyday coding

low … ultra

medium

gpt-5.6-luna

Fast and affordable

low … max

medium

gpt-5.5

Previous generation

low … xhigh

medium

Model and reasoning effort are independent. The model sets raw capability; the effort — low, medium, high, xhigh, max, ultra — sets how long it deliberates before acting. ultra additionally delegates subtasks automatically.

The server does not choose for you. Which model a task deserves depends on your budget and on how costly a wrong answer is, and a regular expression over a prompt cannot know either. Ask for a delegation without naming a model and it refuses — but the refusal carries the recommendation it would have made, so you decide in one more exchange instead of paying for a guess.

If you would rather not be asked, set a default once and it stops asking:

claude mcp add codex-subagent -e CODEX_SUBAGENT_DEFAULT_MODEL=gpt-5.6-terra -- npx -y codex-subagent-mcp

For advice rather than a decision, ask:

Which Codex model should handle migrating this repo's tests to vitest?

That routes mechanical edits to the fast model at low, everyday work to the balanced one at medium, multi-file migrations to the agentic workhorse at high, and hard reasoning problems to the most capable model at xhigh or ultra. It is a suggestion you can ignore, and it stays within the model allow-list and effort ceiling you configure. An effort the chosen model does not support is adjusted to the closest level it does, with a note saying so.

Configuration

Everything is optional, and set through environment variables on the MCP server:

Variable

Effect

CODEX_SUBAGENT_DEFAULT_MODEL

Stops the server asking which model to use.

CODEX_SUBAGENT_DEFAULT_EFFORT

Reasoning effort when a call specifies none.

CODEX_SUBAGENT_ALLOWED_MODELS

Comma-separated allow-list. Anything else is refused.

CODEX_SUBAGENT_DEFAULT_SANDBOX

Sandbox when a call specifies none. Defaults to read-only and cannot exceed the ceiling.

CODEX_SUBAGENT_MAX_SANDBOX

Ceiling on what a delegation may do. Defaults to workspace-write; it must be set to danger-full-access explicitly before unsandboxed calls are allowed.

CODEX_SUBAGENT_MAX_EFFORT

Ceiling on reasoning effort. Useful for keeping ultra off the table. A call above it is lowered to a level the model supports, or refused if the model has none that low.

CODEX_BIN

Path to the Codex executable, if it is not codex on PATH. On Windows it must be codex.exe, not a .cmd shim.

The sandbox settings express policy you choose outside the repository: a default saves repeated arguments, while the ceiling is the boundary no call can cross. See Safety for why the built-in ceiling stops at workspace-write.

Your own escalation rules belong in your CLAUDE.md, in plain language, where Claude applies them with actual understanding and they stay yours. See ADR 12 for why they are not built into this server, and ADR 14 for the sandbox policy split.

Tools

Tool

What it does

codex_doctor

Check the Codex CLI installation and report how to fix it.

list_codex_models

List available models and their reasoning-effort levels.

codex_recommend

Suggest a model and effort for a described task.

codex_delegate

Run a task, blocking or in the background.

codex_follow_up

Continue a previous delegation using its thread_id.

codex_job_status

Check a background delegation.

codex_job_result

Read a finished background delegation's output.

codex_job_cancel

Stop a running background delegation.

Full parameter reference: docs/TOOLS.md.

FAQ

Does this cost money? It uses your existing Codex quota, the same as running codex yourself. This server adds nothing. Higher reasoning efforts consume more; codex_recommend exists partly so you do not spend ultra on work that low would have handled.

Can it modify my files? Not by default. Delegations run read-only unless you explicitly ask for write access, and use_worktree keeps even those changes out of your working tree.

Why drive the CLI instead of calling the OpenAI API? Delegated coding is not a single completion — it is an agentic loop with a sandbox, an approval model, session persistence and project instruction files. All of that lives in the Codex client, not in the model endpoint. See ADR 1.

Do I need Claude Code, or does Claude Desktop work? Either. Claude Code gets a one-line install; Claude Desktop needs a manual config entry.

It says Codex is not installed, but codex works in my terminal. Most likely Windows with a global npm install, which produces a codex.cmd batch shim that cannot be launched without a command shell. codex_doctor reports this as unsupported-shim and offers two fixes. On macOS and Linux, check whether a Node version manager moved codex off PATH.

Does it work on Windows and Linux? CI builds, tests and starts the server on Windows, macOS and Linux on every change, and checks that the Codex CLI is resolved correctly on each. A real delegation has only been verified on macOS — the CI runners have no Codex installation or credentials. Reports from Windows and Linux are welcome. On Windows, install the Codex CLI with the PowerShell installer rather than npm; on Linux, install bubblewrap for Codex's sandbox. See Requirements.

Can Codex read files outside the directory I point it at? Yes, in every sandbox mode — the sandbox restricts writes and network access, not reads. See Safety for what that means in practice and what to do about it.

Where do worktree changes end up? Under ~/.codex/worktrees/, and the delegation result gives you the full path of each file it touched, up to a thousand distinct files. The server does not clean those worktrees up: they may hold work you have not applied yet.

Documentation

  • docs/DELEGATING.md — how to scope a delegation, with measured costs and worked prompts.

  • docs/TOOLS.md — every tool and parameter.

  • docs/adr/ — why the design is what it is, decision by decision.

  • docs/ROADMAP.md — what is planned, and what is deliberately out of scope.

  • Issues — what is actually open right now.

  • docs/VERSIONING.md — what counts as a breaking change here.

  • CHANGELOG.md — what changed in each release, and the Codex CLI version it was verified against.

  • CONTRIBUTING.md — setup, and the rules that are not negotiable.

Disclaimer

Not an official product. This is an independent, community project. It is not affiliated with, endorsed by, sponsored by or supported by OpenAI or Anthropic. "Codex", "ChatGPT" and "OpenAI" are trademarks of OpenAI; "Claude" and "Claude Code" are trademarks of Anthropic. They are used here only to describe what this software interoperates with, which is nominative use — no claim is made to any of them. Neither company is responsible for this software, and problems with it should be reported here rather than to them.

No warranty. The software is provided "as is", without warranty of any kind, as stated in LICENSE. You use it at your own risk.

It runs an agent on your machine. This server spawns the Codex CLI as a child process. Depending on the sandbox you allow, that process can read your files, run shell commands and modify your working tree. Read Safety before enabling writes, and review what a delegation did rather than assuming it did what you asked.

It spends your quota. Delegations consume your own OpenAI Codex usage, at whatever rate your account is billed. Higher reasoning efforts consume more, and ultra delegates subtasks of its own. This project has no visibility into that cost and does not cap it beyond the limits you configure yourself.

License

MIT. See LICENSE.

Available Tools

8 tools
codex_delegateDelegate a task to CodexA

Delegate a task to the local Codex CLI (OpenAI's coding agent), choosing model and reasoning effort. Use it when the user asks for Codex, or when handing work off clearly serves their request: a second opinion from a different model family, or an investigation that would otherwise flood this conversation. When the user has not named a model, call codex_recommend first and present its suggested model and effort to the user in the same message in which you say you are going to delegate, then pass both explicitly here. That recommendation is advice for an already-authorised delegation, not a replacement for the user's own preference. Everything passed in prompt, context and target_files is sent to OpenAI, and every run spends the user's own Codex usage, so do not delegate what you can answer directly, and tell the user when you delegate. Codex runs read-only unless a different default sandbox is configured. Set sandbox to workspace-write to let it edit files. Codex cannot see this conversation, so pass everything it needs in prompt, context, and target_files.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNoblocking (default) waits and streams progress; background returns a job_id immediately.
modelNoCatalog slug from list_codex_models. If omitted, the configured default is used; with no default configured the call is refused and the recommended model is returned.
promptYesThe task for Codex. Be specific and self-contained: Codex cannot see this conversation.
contextNoBackground Codex needs: prior findings, constraints, relevant excerpts.
sandboxNoSandbox policy. Uses the configured default when omitted; without one, Codex runs read-only.
add_dirsNoAdditional absolute directories that should be writable alongside working_dir.
web_searchNoEnable Codex's API-backed live web-search tool for this run. In a read-only sandbox, shell commands have no network access, so this is the route to current external information. When omitted, Codex's own configured web_search mode applies.
working_dirNoAbsolute path Codex uses as its working root.
auto_approveNoAdds --approve-for-me so Codex auto-approves its own commands. Only applies when sandbox is workspace-write.
target_filesNoPaths Codex should focus on, relative to working_dir.
use_worktreeNoRun in a managed git worktree. Writes outside it remain subject to the sandbox policy and add_dirs.
timeout_secondsNoWall-clock budget. Defaults to 1800s.
reasoning_effortNoReasoning depth, independent of model choice. Uses the configured default or the model's default when omitted. Clamped to supported levels within the configured ceiling; refused if none qualify.
acceptance_criteriaNoConcrete conditions that must hold for the task to be considered done.
skip_git_repo_checkNoAllow running outside a git repository.
system_instructionsNoPersona or extra rules inherited from the orchestrator, layered on the built-in quality contract.

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description goes well beyond the annotations by disclosing that all prompt/context/target_files content is sent to OpenAI, that runs spend the user's Codex usage, that Codex cannot see the conversation, and that the default sandbox is read-only unless explicitly configured otherwise. These are material behavioral facts an agent needs before invoking this side-effecting tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Although long, every sentence earns its place: purpose, usage conditions, model-selection flow, cost/privacy warning, sandbox behavior, and a crucial blind-spot reminder are all packed into a dense but structured paragraph. The critical caveats are front-loaded after the opening purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's high complexity (16 params, no output schema, external sendouts, cost implications), the description is exceptionally complete. It covers authorization, prerequisite calls, parameter semantics for the foundational fields, sandbox behavior, and the user-facing transparency obligation. Nothing needed to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds meaningful cross-parameter semantics: it explains that prompt/context/target_files must carry everything Codex needs, that sandbox=workspace-write is required for file edits, and that model selection should follow codex_recommend. This enriches the schema descriptions rather than just repeating them.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action and resource: 'Delegate a task to the local Codex CLI', and clarifies the scope of the tool by naming the choices it makes (model, reasoning effort). It is clearly distinct from siblings like codex_recommend, list_codex_models, and codex_job_status.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit when-to-use conditions: when the user asks for Codex, for a second opinion from a different model family, or when an investigation would flood the conversation. It also gives a when-not-to-use rule ('do not delegate what you can answer directly') and a precise prerequisite protocol involving calling codex_recommend first when no model is named.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

codex_doctorCheck the Codex CLI installationA
Read-only

Check whether the local Codex CLI is installed, recent enough, signed in and able to load its configuration, and report the exact steps to fix it if not. Run this when any other tool reports the CLI is unavailable, or before relying on delegation for the first time. It only inspects the installation; it never installs or changes anything.

ParametersJSON Schema
NameRequiredDescriptionDefault
refreshNoRe-probe the CLI instead of reusing the cached diagnosis.
working_dirNoAbsolute directory to run the check in. Codex loads the configuration of the directory it runs in, so pass the one a delegation would use. Defaults to this server's own.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already declare readOnlyHint=true and openWorldHint=false, and the description's 'never installs or changes anything' reinforces that safety profile without contradicting it. It adds useful behavioral scope by explaining the inspection covers signed-in state and configuration loading, and that it emits remediation steps rather than performing fixes. It doesn't detail the exact return format or mention the cached diagnosis behavior described in the refresh parameter, but the schema covers that.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, no fluff, and the most decision-relevant facts come first: what it checks, when to run it, and what it will not do. Every sentence earns its place, and the coverage of purpose, usage, and safety is achieved in extremely compact form.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-required-parameter diagnostic tool, the definition is complete: it states the checks performed, what the output will contain (fix steps), when to invoke it, and that it is read-only. The schema covers the two optional parameters, and the annotations cover the side-effect profile, so nothing needed for correct selection or invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and both parameters already have strong descriptions: refresh explains the cache/re-probe behavior)Skip and working_dir explains the directory-sensitive configuration loading and its default. The tool description contributes contextual motivation by tying the check to delegation, but it does not add meaning beyond what the schema already provides. A baseline 3 is appropriate because the schema carries the parameter documentation burden.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb ('check') and a concrete resource (the local Codex CLI installation), then elaborates with four concrete checks: installed, recent enough, signed in, and able to load configuration. It also states the actual outcome ('report the exact steps to fix it'), which differentiates this diagnostic tool from the delegation, recommendation, and job-management siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit trigger conditions: run when another tool reports the CLI is unavailable, or before relying on delegation for the first time. It also states a clear exclusion—it only inspects and never installs or changes anything—so an agent knows not to use it as a fixer and should look elsewhere for remediation.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

codex_follow_upContinue a Codex sessionA

Send a follow-up message to a previous delegation using its thread_id. Codex retains the earlier context, so only the new instruction needs to be sent. Like a delegation, it is sent to OpenAI and spends the user's Codex usage.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelNoOverride the model for this turn. Defaults to the thread's last model from memory or, on a registry miss, Codex's session file; if neither has it, the configured default, else the call is refused.
promptYesThe follow-up instruction.
sandboxNoSandbox policy for this turn. Uses the configured default when omitted; read-only when unset.
thread_idYesThe thread_id reported by a previous codex_delegate call.
working_dirNoAbsolute directory to resume in. Defaults to the directory the thread last ran in.
auto_approveNoNot supported on follow-ups: true is refused and nothing runs.
timeout_secondsNo
reasoning_effortNoOverride the reasoning effort for this turn. Defaults to the thread's last effort when the model is unchanged, otherwise to the configured or model default.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations (readOnlyHint=false, openWorldHint=true), the description accurately discloses that the call is sent to OpenAI and consumes the user's Codex usage, and that context from the prior thread is used. It does not exhaustively describe side effects such as sandbox execution, but the schema covers those controls.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, each carrying distinct information: purpose, context-retention behavior, and cost/external call. No filler or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The core call semantics are clear, but with no output schema the description does not explain what a successful call returns or how to retrieve the follow-up's outcome (e.g., via job status/result siblings). It also leaves parameter-level behavioral caveats like auto_approve being refused to the schema, which is acceptable but means the description is not fully self-contained.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is high (88%), so the baseline is 3. The description adds useful semantic context by linking thread_id to a previous delegation and explaining that prompt only needs the new instruction because context is retained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and object: 'Send a follow-up message to a previous delegation using its thread_id.' It clearly identifies the continuation use case and distinguishes it from a fresh delegation by noting that prior context is retained.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear context for use: only after a previous delegation, identified by thread_id, and explains that only the new instruction is needed. It does not explicitly name codex_delegate or state when not to use alternatives, but the distinction is strongly implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

codex_job_cancelCancel a background delegationB
Destructive

Terminate a running background delegation.

ParametersJSON Schema
NameRequiredDescriptionDefault
job_idYesThe job_id returned by codex_delegate.

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and readOnlyHint=false, clearly indicating a destructive operation. The description adds 'running' as a state constraint, which is useful. However, it omits key behavioral details like whether the job is killed immediately or gracefully, whether it can be undone, or what happens to any partial results. With annotations covering the safety profile, a 3 is appropriate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence with no wasted words. It is front-loaded with the action, though it is extremely terse, which might leave some ambiguity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (one parameter, no output schema) and the annotations covering destructiveness, the description is minimally adequate but lacks important context about side effects, job state requirements, and the return value or confirmation. It should do more to help an agent understand when and how to use it safely.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the single parameter job_id is fully documented in the schema, including that it comes from codex_delegate. The description adds no additional parameter meaning beyond what the schema provides, making the baseline 3 correct.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Terminate) and resource (running background delegation), which is clearer than the title. However, it does not distinguish itself from siblings like codex_job_status or codex_job_result, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives, or when not to use it. The description gives no context about the conditions under which cancellation is appropriate, such as whether the job must be running or what happens if it has already completed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

codex_job_resultRead a background delegation's resultA
Read-only

Return the full output of a finished background delegation. Errors if the job is still running.

ParametersJSON Schema
NameRequiredDescriptionDefault
job_idYesThe job_id returned by codex_delegate.

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds a meaningful behavioral constraint: it errors when the job is still running, telling the agent this is a post-completion read with a specific failure mode. It does not discuss output truncation or pagination, so 4.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, zero waste, and the completion precondition is front-loaded. The error condition follows immediately after the main action, keeping the most important routing information prominent.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter read tool with full schema coverage and annotations covering safety, the description supplies the necessary completion precondition and output scope. No output schema exists, so describing the return as 'full output' is appropriately complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the only parameter 'job_id' is documented as 'returned by codex_delegate'. The description adds no additional syntax, format, or constraint beyond the schema, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Return') and resource ('full output of a finished background delegation'), and the error condition distinguishes it from a status check. It does not explicitly name sibling codex_job_status as the alternative, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Implies use only after completion via 'finished background delegation' and 'Errors if the job is still running'. However, it does not explicitly tell the agent to call codex_job_status first or name alternatives, so usage is inferable but not spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

codex_job_statusCheck a background delegationA
Read-only

Report the state and recent activity of a background delegation started with mode=background. Call it with no job_id to list every known job.

ParametersJSON Schema
NameRequiredDescriptionDefault
job_idNoOmit to list all jobs.
include_activityNoInclude the recent progress log for the job.

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds that it reports state plus recent activity and that omitting job_id enumerates all jobs, but says nothing about pagination, log volume, or failure modes for unknown job ids.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the core purpose and with the no-arg listing behavior second; no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-required-param, read-only status tool with no output schema, the description covers what is queried (state and activity) and the two calling modes. It is nearly complete, with only return-shape hints and error handling left unspecified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and both parameters are already documented in the schema; the description's note about omitting job_id restates that documentation rather than adding format or constraint detail.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (report) plus resource (state and recent activity of a background delegation) and scopes it to delegations started with mode=background, which links it to codex_delegate and separates it from codex_job_result.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives one concrete usage rule — call with no job_id to list every known job — but never says when to prefer this over codex_job_result, codex_follow_up, or codex_job_cancel, so the sibling routing is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

codex_recommendRecommend a Codex model and effortA
Read-only

Given a task description, recommend which Codex model and reasoning effort to delegate it with. Runs no model call; applies a documented matrix reconciled against the installed catalog.

ParametersJSON Schema
NameRequiredDescriptionDefault
priorityNoBias the reasoning effort: quality raises it, latency and cost lower it. Default balanced.
working_dirNoAbsolute directory whose Codex configuration and model catalog should be used. Defaults to this server's own.
task_descriptionYesWhat the delegated task involves, in one or two sentences.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnlyHint and openWorldHint annotations, the description discloses that the tool executes no model call and applies a documented matrix reconciled against the installed catalog. This adds meaningful context about how the recommendation is derived (deterministically, not by running a model) without conflicting with the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The two-sentence description is compact and front-loaded: it immediately states the input ('Given a task description') and the output ('recommend which Codex model and reasoning effort'). The second sentence clarifies a key behavioral trait without waste. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple tool with 3 parameters (1 required) and no output schema, the description is adequately complete. It specifies the recommendation output (model and effort) but does not detail the exact return structure (e.g., object or string). Given the low complexity and clear sibling context (codex_delegate, list_codex_models), this is sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

All three parameters are fully described in the input schema (100% coverage), so the description adds no parameter-specific detail beyond the schema. With full schema coverage, the baseline of 3 is appropriate; the description does not compensate with extra parameter guidance.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'recommend which Codex model and reasoning effort to delegate it with.' It also notes 'Runs no model call,' which clearly differentiates this advisory tool from execution tools like codex_delegate. The purpose is unambiguous and distinct from siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage ('Given a task description... recommend') but does not explicitly say when to use this tool versus alternatives like codex_delegate or list_codex_models. The phrase 'Runs no model call' hints it is not for execution, but there is no explicit 'Use this before delegating' or 'For execution, use codex_delegate.' Guidance is implied, not stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_codex_modelsList Codex modelsA
Read-only

List the Codex models available on this machine, with the reasoning-effort levels each one supports. Read from the installed Codex CLI, with a warned static fallback if its catalog cannot be read. Call this before codex_delegate when choosing a model explicitly.

ParametersJSON Schema
NameRequiredDescriptionDefault
refreshNoBypass the cache and re-read the catalog from the CLI.
working_dirNoAbsolute directory to read the catalog in. A project you have trusted in Codex can set its own catalog, so pass the directory a delegation would use. Defaults to this server's own.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark the tool read-only, and the description adds useful behavioral context: it reads from the installed Codex CLI and has a 'warned static fallback' if the catalog cannot be read. This goes beyond the annotation in a meaningful way, though caching behavior is left to the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, each earning its place: the core purpose, the fallback behavior, and the usage guidance. It is front-loaded with the essential information and contains no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only listing tool, the description covers the key context: what is listed, where it is read from, the fallback behavior, and when to call it. It does not describe the return format, but no output schema exists and the purpose makes the return shape largely predictable.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema fully documents both parameters. The description does not add parameter-level meaning beyond the schema, which makes the baseline 3 appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific action ('List'), a precise resource ('Codex models available on this machine'), and additional detail ('reasoning-effort levels each one supports'). It also distinguishes itself from a sibling by pointing to codex_delegate, so an agent can tell why this tool exists.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit usage context: 'Call this before codex_delegate when choosing a model explicitly.' It clearly states when the tool should be used, though it does not name alternatives to avoid or describe when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 5 tool updatesv0.3.0
    • Changedcodex_delegate2 fields changed
      • changedInput schema / properties / sandbox / description
        Previous value: -"Sandbox policy. Defaults to read-only: Codex analyses and reports but cannot modify files."New value: +"Sandbox policy. Uses the configured default when omitted; without one, Codex runs read-only."
      • changedInput schema / properties / web_search / description
        Previous value: -"Enable live web search for this run, through Codex's web_search = \"live\" setting. When omitted, Codex's own configured web_search mode applies."New value: +"Enable Codex's API-backed live web-search tool for this run. In a read-only sandbox, shell commands have no network access, so this is the route to current external information. When omitted, Codex's own configured web_search mode applies."
    • Changedcodex_doctor1 field changed
      • addedInput schema / properties / working_dir
        Added value: +{
        +  "description": "Absolute directory to run the check in. Codex loads the configuration of the directory it runs in, so pass the one a delegation would use. Defaults to this server's own.",
        +  "type": "string"
        +}
    • Changedcodex_follow_up2 fields changed
      • changedInput schema / properties / model / description
        Previous value: -"Override the model for this turn. Defaults to the model the thread last ran with on this server; for a thread this server has no record of, the configured default, else the call is refused."New value: +"Override the model for this turn. Defaults to the thread's last model from memory or, on a registry miss, Codex's session file; if neither has it, the configured default, else the call is refused."
      • changedInput schema / properties / sandbox / description
        Previous value: -"Sandbox policy for this turn. Defaults to read-only."New value: +"Sandbox policy for this turn. Uses the configured default when omitted; read-only when unset."
    • Changedcodex_recommend1 field changed
      • addedInput schema / properties / working_dir
        Added value: +{
        +  "description": "Absolute directory whose Codex configuration and model catalog should be used. Defaults to this server's own.",
        +  "type": "string"
        +}
    • Changedlist_codex_models1 field changed
      • addedInput schema / properties / working_dir
        Added value: +{
        +  "description": "Absolute directory to read the catalog in. A project you have trusted in Codex can set its own catalog, so pass the directory a delegation would use. Defaults to this server's own.",
        +  "type": "string"
        +}
  2. 2 tool updatesv0.2.0
    • Changedcodex_delegate5 fields changed
      • changedInput schema / properties / auto_approve / description
        Previous value: -"Adds --approve-for-me so Codex auto-approves its own commands. Only applies when sandbox allows writes."New value: +"Adds --approve-for-me so Codex auto-approves its own commands. Only applies when sandbox is workspace-write."
      • changedInput schema / properties / model / description
        Previous value: -"Catalog slug from list_codex_models. Omitted means the recommendation matrix picks one."New value: +"Catalog slug from list_codex_models. If omitted, the configured default is used; with no default configured the call is refused and the recommended model is returned."
      • changedInput schema / properties / reasoning_effort / description
        Previous value: -"Reasoning depth, independent of model choice. Clamped to what the chosen model supports."New value: +"Reasoning depth, independent of model choice. Uses the configured default or the model's default when omitted. Clamped to supported levels within the configured ceiling; refused if none qualify."
      • changedInput schema / properties / use_worktree / description
        Previous value: -"Run in a managed git worktree so changes never touch the current working tree."New value: +"Run in a managed git worktree. Writes outside it remain subject to the sandbox policy and add_dirs."
      • changedInput schema / properties / web_search / description
        Previous value: -"Enable Codex's native web search tool."New value: +"Enable live web search for this run, through Codex's web_search = \"live\" setting. When omitted, Codex's own configured web_search mode applies."
    • Changedcodex_follow_up4 fields changed
      • addedInput schema / properties / auto_approve / description
        Added value: +"Not supported on follow-ups: true is refused and nothing runs."
      • changedInput schema / properties / model / description
        Previous value: -"Override the model for this turn."New value: +"Override the model for this turn. Defaults to the model the thread last ran with on this server; for a thread this server has no record of, the configured default, else the call is refused."
      • changedInput schema / properties / reasoning_effort / description
        Previous value: -"Override the reasoning effort for this turn."New value: +"Override the reasoning effort for this turn. Defaults to the thread's last effort when the model is unchanged, otherwise to the configured or model default."
      • addedInput schema / properties / working_dir / description
        Added value: +"Absolute directory to resume in. Defaults to the directory the thread last ran in."
  3. 1 tool update
    • Changedcodex_follow_up1 field changed
      • addedInput schema / properties / thread_id / pattern
        Added value: +"^[A-Za-z0-9][A-Za-z0-9_-]{0,199}$"
  4. 8 tool updatesv0.1.0
    • First observedcodex_delegate
    • First observedcodex_doctor
    • First observedcodex_follow_up
    • First observedcodex_job_cancel
    • First observedcodex_job_result
    • First observedcodex_job_status
    • First observedcodex_recommend
    • First observedlist_codex_models

TDQS

A3.9/5.0

Scored across 8 tools

Disambiguation5/5

Each tool targets a clearly distinct concern: installation health, model listing, model recommendation, delegation, follow-up, and background job management. Even related tools like codex_recommend and codex_delegate are cleanly separated by role.

Naming Consistency3/5

The set mostly uses a codex_ prefix, but the verb placement varies: codex_delegate and codex_follow_up put the verb first, codex_job_cancel/result/status put the noun first, and list_codex_models breaks the prefix pattern entirely. The naming is readable but not uniformly consistent.

Tool Count5/5

Eight tools is well-scoped for a Codex delegation server. Each tool covers a necessary part of the workflow—diagnostics, model selection, delegation, follow-up, and background job control—without redundancy.

Completeness4/5

The core delegation lifecycle is covered: check readiness, choose a model, delegate, follow up, and manage background jobs. Minor gaps exist, such as no way to list prior threads or retrieve historical delegation summaries, but agents can generally accomplish the main workflows without dead ends.

Maintenance

ActivityMaintained
ResponsivenessResponsive

Related MCP Connectors

Related MCP Servers