Skip to main content
Glama

claude-codex-mcp

Let Claude hand work to OpenAI Codex — code tasks, second opinions and image assets — on the ChatGPT/Codex subscription you already pay for.

CI License: MIT Node >= 20 Dependencies: 0 Windows | macOS | Linux

English · Bahasa Indonesia

claude-codex-mcp is a small MCP server that drives the official Codex CLI (codex exec) on your machine. Claude gets five tools: delegate a task, generate images with Codex's built-in image_gen, poll/cancel jobs, and list models. No OpenAI API key, no token scraping, no third-party service — just the CLI you are already logged into.

It works in Claude Desktop (Chat and Cowork) and Claude Code, and it was built Windows-first.

You:    Ask Codex to review the auth module read-only, then fix what you both agree on.
Claude: → codex_task { cwd: "D:\\work\\shop", sandbox: "read-only", model: "gpt-6-astra", reasoning_effort: "xhigh" }
        ← 3 findings, session_id 01a1…   (Claude verifies them, applies the fix, runs the tests)

You:    Make 4 transparent gold-coin icons for the HUD.
Claude: → codex_image { prompt: "...", count: 4, transparent: true, out_dir: "D:\\work\\game\\assets\\icons" }
        ← coin.png, coin-2.png, coin-3.png, coin-4.png + previews

Features

  • Delegation — codex_task runs a full Codex agent turn in a folder you choose, returns its final message and a session_id to continue the same Codex thread.

  • Image assets — codex_image uses Codex's built-in GPT Image tool, copies results into your project (never overwriting), supports transparent backgrounds and reference images, and returns previews so Claude can check them.

  • Model choice — codex_models reads the live catalog for your account; pick any model and reasoning_effort (low … ultra) per call.

  • Never blocks the client — calls wait up to wait_seconds (default 50 s) and then hand back a job_id; progress notifications keep long calls alive. Image jobs are serialized so outputs are attributed to the right job.

  • Safe defaults — workspace-write sandbox (or read-only), network off unless asked, no danger-full-access. Prompts go over stdin and nothing ever passes through a shell.

  • Zero runtime dependencies — plain Node ≥ 20. Easy to audit: about 1,000 lines in src/.

Related MCP server: chatgpt-deepseek-bridge

How it works

flowchart LR
  C["Claude<br/>Desktop · Cowork · Code"] -- "MCP (stdio)" --> S["claude-codex-mcp"]
  S -- "spawn, no shell<br/>prompt on stdin" --> X["codex exec --json"]
  X -- "your ChatGPT login" --> O[("OpenAI")]
  X -- "JSONL events:<br/>thread id, messages, usage, errors" --> S
  X -. "image_gen writes" .-> G["~/.codex/generated_images/&lt;thread&gt;/"]
  S -- "copy + preview" --> P["your project's asset folder"]

Codex CLI ≥ 0.160 no longer ships codex mcp-server, so this project wraps the supported non-interactive entry point, codex exec, and reads its --json event stream.

Requirements

  • Node.js 20+

  • Codex CLI, logged in with your ChatGPT account: npm i -g @openai/codex (or the Codex desktop app's bundled CLI), then codex login

  • A ChatGPT plan that includes Codex. Usage counts against your Codex limits; image generation uses them faster than text.

Install

Option A — Claude Code plugin

/plugin marketplace add aikazu/claude-codex-mcp
/plugin install codex-mcp@aikazu

This adds the MCP server plus a codex-delegation skill that teaches Claude when delegation is worth it and how to brief Codex.

Option B — Claude Desktop / Cowork (and Claude Code)

git clone https://github.com/aikazu/claude-codex-mcp.git
cd claude-codex-mcp

Windows (PowerShell):

powershell -ExecutionPolicy Bypass -File .\scripts\install.ps1

macOS / Linux:

./scripts/install.sh

The installer checks Node and your Codex login, then registers the server as codex in every claude_desktop_config.json it finds (standard and Microsoft Store installs, with a timestamped backup) and in Claude Code (claude mcp add -s user). It pins absolute paths for node and codex, because GUI apps often start MCP servers with a minimal PATH. Fully quit and reopen Claude Desktop afterwards.

Options: -AssetDir <dir> / --asset-dir, -Sandbox read-only / --sandbox, -TaskModel / --task-model, -TaskEffort / --task-effort, -ImageModel / --image-model, -ImageEffort / --image-effort, -DryRun / --dry-run, -Uninstall / --uninstall.

Option C — manual config

{
  "mcpServers": {
    "codex": {
      "command": "C:\\Program Files\\nodejs\\node.exe",
      "args": ["C:\\Users\\you\\Tools\\claude-codex-mcp\\src\\cli.mjs"],
      "env": { "CODEX_BIN": "C:\\Users\\you\\AppData\\Local\\Programs\\OpenAI\\Codex\\bin\\codex.exe" }
    }
  }
}
claude mcp add codex -s user -- node /path/to/claude-codex-mcp/src/cli.mjs

Check the setup at any time:

node src/cli.mjs --check

Tools

Tool

What it does

Key arguments

codex_task

Run a Codex agent turn in cwd

prompt, cwd, sandbox, model, reasoning_effort, images, add_dirs, network, session_id, wait_seconds

codex_image

Generate images with Codex image_gen

prompt, out_dir, name, count (1–8), size, transparent, reference_images, model, return_images

codex_job

Wait for, fetch or cancel a job

job_id, wait_seconds, cancel

codex_jobs

List recent jobs

—

codex_models

Models available to your account + configured default

include_hidden

Results are JSON with status (queued · running · completed · failed · cancelled), final_message, session_id, usage, files (images), changed_files (workspace-write tasks in a git repository: absolute paths whose git status or mtime changed during the run), warnings (e.g. a requested transparent image without alpha), and errors / stderr_tail on failure.

Things to ask Claude

  • "Show me codex_models, then ask the strongest one to review this PR's diff, read-only."

  • "Delegate to Codex: add unit tests for src/payments, then review its diff and run the tests yourself."

  • "Start Codex on the migration in the background and keep working on the UI; check back when it's done."

  • "Generate a 16:9 key art for the README with Codex, save it to docs/."

Configuration

Environment variables (set them in the MCP server entry):

Variable

Default

Meaning

CODEX_BIN

auto-detect

Path to codex.exe, the codex binary, or @openai/codex/bin/codex.js

CODEX_HOME

~/.codex

Codex home (auth, config, generated_images)

CODEX_MCP_SANDBOX

workspace-write

Default sandbox for codex_task (read-only or workspace-write)

CODEX_MCP_ASSET_DIR

~/Pictures/codex-assets

Default out_dir for images

CODEX_MCP_TASK_MODEL / CODEX_MCP_TASK_EFFORT

Codex config.toml

model / reasoning_effort for codex_task calls that omit them

CODEX_MCP_IMAGE_MODEL / CODEX_MCP_IMAGE_EFFORT

Codex config.toml

Same for codex_image; a fast model at low is enough, the agent only drives image_gen

CODEX_MCP_WAIT

50

Default seconds a call waits before returning a job_id (max 240)

CODEX_MCP_MAX_TASKS

3

Concurrent codex_task runs; extra jobs queue

CODEX_MCP_MAX_IMAGES

1

Concurrent image jobs (keep 1 for reliable attribution)

CODEX_MCP_PREVIEW_MAX_BYTES

1500000

Largest image embedded as a preview

CODEX_MCP_PREVIEW_MAX_COUNT

4

Previews per result

Security model

  • The server is a local stdio process; it opens no ports and stores no credentials. Authentication is entirely the Codex CLI's own login.

  • Codex runs under its own sandbox: read-only or workspace-write (writes limited to cwd plus add_dirs, network off unless network: true). The bypass/full-access modes are intentionally not exposed. On Windows, delegated runs also turn off the Codex desktop app's browser / computer-use REPL servers (node_repl, cua_repl), which break the sandbox while they run.

  • Arguments are passed as an argv array without a shell; prompts go over stdin. model, reasoning_effort and session_id are validated against strict patterns before they reach Codex.

  • Treat Codex output as untrusted input. The bundled skill tells Claude to verify claims and diffs before relaying them.

See SECURITY.md for reporting.

How it compares

There are good projects in this space; each covers part of it. As of October 2026:

Project

Kind

Delegation

Images

Claude Desktop / Cowork

Windows

Notes

claude-codex-mcp

MCP server

✓

✓

✓

✓ (primary)

async jobs, model catalog, sandboxed, 0 deps

Sateezg/codex-bridge

Claude Code plugin

✓

✓

—

not tested

rich set of skills and subagents

cexll/codex-mcp-server

MCP server

✓

—

✓

✓

ask-codex + brainstorm tools

kky42/codex-as-mcp

MCP server

✓

—

✓

—

parallel agents; runs Codex with sandbox bypassed

glassd/Pixmith

MCP server

—

✓

✓

✓

image-focused, async jobs

ShalomObongo/codex-imagegen-mcp

MCP server

—

✓

✓

—

calls the ChatGPT backend directly with auth.json tokens

Pick whatever fits; this one aims to be the single, auditable server that covers both jobs everywhere Claude runs.

Troubleshooting

Symptom

Fix

Codex CLI not found

Install Codex, or set CODEX_BIN. On Windows point it at codex.exe or …\node_modules\@openai\codex\bin\codex.js — .cmd shims cannot be spawned safely.

Not logged in

Run codex login and sign in with ChatGPT.

Tools don't appear in Claude Desktop

Quit from the tray / menu bar (closing the window is not enough) and reopen. Check the app's MCP logs.

Calls return running

Normal for long runs; Claude polls with codex_job. Raise CODEX_MCP_WAIT if your client allows long tool calls.

usage limit errors

Your Codex quota is spent; it resets on your plan's schedule.

Codex reports CreateProcessWithLogonW failed: 267 (Windows)

Codex's workspace-write sandbox could not start a shell in cwd. Seen with folders under AppData\Roaming; use a project folder elsewhere, or read-only.

helper_unknown_error: setup refresh had errors (Windows)

Codex's sandbox checks the ACLs of everything under %LOCALAPPDATA%\OpenAI\Codex\runtimes and ~\.cache\codex-runtimes before every command, and fails (os error 32 in ~/.codex/.sandbox/sandbox.<date>.log) while any process holds a file there open: the Codex desktop app's codex-computer-use-swift.exe or node_repl.exe, or another Codex session. The server turns off the app's node_repl / cua_repl MCP servers for every run and stops a run as soon as the sandbox rejects a command, instead of letting Codex continue blind and burn tokens. The failed job names the locked file and, when it can find it, the process (name and PID): quit the Codex desktop app from the tray, or end that process, and retry.

No image found after a job

Make sure the job wasn't cancelled and that CODEX_HOME matches the Codex install that ran.

Development

npm install        # dev tooling only (Biome)
npm test           # node:test suite, runs against a fake Codex CLI
npm run lint
node src/cli.mjs --check

The test suite never calls OpenAI: test/fixtures/fake-codex.mjs mimics codex exec --json, debug models and login status. CI runs it on Windows, macOS and Linux and validates the plugin manifests.

Disclaimer

Independent project, not affiliated with OpenAI or Anthropic. It only automates the official Codex CLI on your own machine with your own account; you remain responsible for complying with the terms of your ChatGPT/Codex plan.

License

MIT © Iqbal Attila

Available Tools

5 tools
codex_imageGenerate images with CodexA

Generate image assets with Codex's built-in image_gen tool (GPT Image), billed to the user's ChatGPT/Codex plan — no API key. Files are copied into out_dir and small previews are returned so you can check them. Use for icons, sprites, illustrations, textures, UI mockups and similar. Supports transparent backgrounds and reference images. Each image typically takes 30-120 s.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoFile name prefix, e.g. 'coin-icon'.
sizeNoe.g. '1024x1024', '1536x1024', '16:9', 'square'.
countNoNumber of images/variants (default 1).
modelNoModel for the Codex agent turn that calls image_gen (not the image model). A fast model is enough.
promptYesArt direction: subject, style, palette, composition, intended use.
out_dirNoAbsolute folder to save into (created if missing). Default: /root/Pictures/codex-assets
transparentNoRequest a real transparent background.
wait_secondsNoSeconds to wait before returning a job_id (0-240, default 50).
return_imagesNoEmbed previews in the result (default true).
reasoning_effortNolow | medium | high | xhigh | max | ultra — support varies per model (see codex_models). Omit for the default.
reference_imagesNoAbsolute paths of images to edit or match.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds substantial context beyond the annotations (openWorldHint, destructiveHint=false): billing to the user's ChatGPT/Codex plan with no API key, files copied into out_dir, previews returned for verification, transparent/reference-image support, and a 30-120 s latency expectation. The latency and billing disclosures are exactly the behavioral traits an agent needs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Five tight sentences, front-loaded with what the tool is and the billing model, then capabilities, then latency. No filler and every sentence carries distinct information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers purpose, cost, output location, previews, capabilities and timing for an 11-param tool. It stops short of explaining the async/job_id behavior implied by wait_seconds, which matters with no output schema, but otherwise the agent has enough to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all 11 parameters thoroughly; baseline is 3. The description reinforces a few (out_dir copying, transparent backgrounds, reference images) but adds no syntax or format detail beyond the schema text.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Generate image assets') and names the underlying mechanism (Codex's built-in image_gen / GPT Image), which clearly separates it from siblings like codex_job, codex_task, and codex_models.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete use cases ('icons, sprites, illustrations, textures, UI mockups and similar'), so the agent knows when this tool applies. However, it never names an alternative or states when NOT to use it (e.g. vs codex_task), so routing guidance is clear but incomplete.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

codex_jobCodex job statusB

Wait for or fetch the result of a codex_task / codex_image job, or cancel it.

ParametersJSON Schema
NameRequiredDescriptionDefault
cancelNoStop the job.
job_idYes
wait_secondsNoSeconds to wait before returning a job_id (0-240, default 50).

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must carry the behavioral burden; it does disclose three distinct modes (block waiting, plain fetch, cancel), which is real behavioral content. It omits whether cancel is destructive or irreversible, whether waiting blocks the agent, and any permission or rate-limit constraints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence covering all three operating modes with no filler; the fetch/result framing is front-loaded. Slightly compressed in a way that blurs the distinction between the three modes.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and no annotations, the description should say what a result actually looks like and what happens when the job is not yet complete. It only hints at this indirectly via the schema's wait_seconds ('returning a job_id'), and is silent on cancel permanence.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 67%: cancel and wait_seconds are documented in the schema with meaning and range. The description adds value by clarifying that job_id originates from a codex_task or codex_image job, but it adds nothing about cancel semantics or wait behavior beyond what the schema says.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb set (wait for / fetch / cancel) and the resource (a codex_task or codex_image job), so an agent understands it acts on job results rather than creating them. It does not distinguish itself from the sibling codex_jobs, leaving ambiguity about which of the two to call.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

By naming codex_task / codex_image as the job producers, it implies this is the follow-up call that retrieves or cancels their output. However, there is no explicit when-to-use vs codex_jobs, no statement of preconditions, and no guidance on choosing wait vs fetch vs cancel.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

codex_jobsList Codex jobsB
Read-only

List recent jobs started by this server (status, kind, session_id).

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The readOnlyHint=true annotation already establishes this is a safe read. The description adds useful scoping ('started by this server') and discloses the shape of each returned row (status, kind, session_id), which the annotations do not cover. However, it says nothing about how 'recent' is bounded, ordering, or result limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with the verb and scope first and the returned fields in a compact parenthetical. Nothing is wasted, though the parenthetical field list is thin enough that it barely earns its space.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully compensates by naming the fields in each row, and the zero-parameter schema means there is no input surface to document. The main remaining gap is the absence of any bound on 'recent' (count, window, or pagination), which an agent would want before calling.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline is 4. There is no parameter surface for the description to clarify or omit, and it correctly does not invent any.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('List') and resource ('recent jobs started by this server'), and even names the fields returned (status, kind, session_id). It is clearly distinguishable from the singular codex_job sibling, but the description never explicitly names that sibling or the boundary between them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'recent jobs started by this server' implies a browsing use case, but there is no explicit when-to-use statement, no when-not, and no reference to codex_job for retrieving a single job's details. The agent must infer the routing itself.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

codex_modelsList Codex modelsA
Read-only

List Codex models available to the user's account (slug, description, supported reasoning efforts, default effort) and the configured default from config.toml. Use before choosing model / reasoning_effort.

ParametersJSON Schema
NameRequiredDescriptionDefault
include_hiddenNoAlso list hidden/special-purpose models.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

readOnlyHint=true already declares the safety profile, so the bar is lower. The description still adds real context: the data comes from both the user's account and the configured default in config.toml, and it names what each entry contains. It omits anything about ordering or result size, but for a small enumeration that is minor.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, both earning their place: the first defines the output, the second gives the call trigger. No redundancy and no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema present, the description carries the return-value burden and does so by listing the fields returned and where the default comes from. Combined with the explicit usage trigger, an agent has everything needed to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Only one parameter and schema description coverage is 100%, so the schema fully documents include_hidden. The description says nothing about it, adding no meaning beyond structured data, which is the baseline-3 case.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb (List) + resource (Codex models) scoped to the user's account, and it enumerates the returned fields (slug, description, reasoning efforts, default effort) plus the config.toml default. Siblings (codex_job, codex_task, codex_image, codex_jobs) are plainly different resources, so no confusion arises.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to call it: 'Use before choosing `model` / `reasoning_effort`.' That gives a clear triggering condition. It does not name an alternative tool or an exclusion case, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

codex_taskDelegate to CodexA
Destructive

Delegate a task to OpenAI Codex, running codex exec on this computer with the user's own ChatGPT/Codex subscription. Good for a second opinion or code review, implementing a well-scoped change, refactors, writing tests, or digging through a codebase. Write a self-contained prompt: goal, relevant files, constraints, and what to report back. Codex works in cwd; with sandbox workspace-write it can edit files there. Returns Codex's final message plus a session_id that continues the same Codex session. Runs longer than wait_seconds return a job_id — poll it with codex_job.

ParametersJSON Schema
NameRequiredDescriptionDefault
cwdNoAbsolute path of the working folder (required for a new task).
modelNoCodex model slug (see codex_models). Omit to use the user's default.
imagesNoAbsolute paths of images to attach.
promptYesComplete instructions for Codex.
networkNoAllow network inside workspace-write (e.g. package installs). Default false.
sandboxNoDefault workspace-write. read-only = analyse only.
add_dirsNoExtra writable folders (new tasks only).
session_idNoContinue an earlier Codex session instead of starting fresh.
wait_secondsNoSeconds to wait before returning a job_id (0-240, default 50).
reasoning_effortNolow | medium | high | xhigh | max | ultra — support varies per model (see codex_models). Omit for the default.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds substantial context beyond the destructiveHint/openWorldHint annotations: it runs locally with the user's own subscription, executes `codex exec`, works in `cwd`, can edit files there under workspace-write, and describes both return shapes (final message + session_id, or a job_id when it overruns wait_seconds). This is exactly the behavioral depth an agent needs before invoking a file-mutating local tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads what the tool is and where it runs, then use cases, then prompt guidance, then return/job behavior. Every sentence carries information; no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 10-parameter local execution tool with no output schema, the description covers the execution model, file-edit semantics, prompt construction, session continuation, and the async job_id fallback. An agent can invoke it correctly and handle both result shapes without further context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real meaning: it explains how to compose `prompt` (goal, relevant files, constraints, what to report back) and connects `sandbox`, `cwd`, `session_id`, and the job_id flow. This goes beyond restating the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific action (delegate a task to OpenAI Codex via `codex exec`) and the resource/context it operates on (this computer, the user's subscription). It also implicitly distinguishes itself from the polling sibling by noting that long runs return a job_id handled by codex_job.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete use cases (second opinion, code review, well-scoped change, refactors, tests, codebase digging) and routes the agent to codex_job for waiting tasks and to codex_models for model slugs. It lacks explicit when-not-to-use guidance but the positive routing is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 5 tool updatesv1.0.0
    • First observedcodex_image
    • First observedcodex_job
    • First observedcodex_jobs
    • First observedcodex_models
    • First observedcodex_task

TDQS

A3.9/5.0

Scored across 5 tools

Disambiguation4/5

codex_task, codex_image, and codex_models target clearly distinct purposes (delegation, image generation, model discovery). The one hazard is codex_job vs codex_jobs, whose near-identical names differ only by number; descriptions do clarify (single-job wait/fetch/cancel vs listing recent jobs), so the overlap is minor rather than fatal.

Naming Consistency4/5

Every tool shares the predictable `codex_` prefix plus a noun (task, job, image, jobs, models), which reads consistently. The only blemish is the singular/plural pair codex_job/codex_jobs, a small deviation from an otherwise clean scheme.

Tool Count5/5

Five tools is well-scoped for a Codex-delegation server: one per core action (task, image, job polling, job listing, model listing). Each tool earns its place with no redundancy or padding.

Completeness4/5

The surface covers the full lifecycle: model discovery before choosing, task/image execution, and job wait/fetch/cancel plus listing for async runs. Cancellation being folded into codex_job rather than a dedicated tool is a minor gap an agent can work around.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers