claude-codex-mcp
Drives the official OpenAI Codex CLI (codex exec) to delegate work to Codex on the user's ChatGPT/Codex subscription. Provides tools to run a Codex agent turn in a chosen directory (returning the final message and a session_id to continue the thread), generate images via Codex's built-in image_gen (with count, size, transparent backgrounds and reference images), poll/wait for or cancel asynchronous jobs, list recent jobs, and query the live model catalog to pick a model and reasoning effort per call.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@claude-codex-mcpAsk Codex to review the auth module read-only."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
claude-codex-mcp
Let Claude hand work to OpenAI Codex — code tasks, second opinions and image assets — on the ChatGPT/Codex subscription you already pay for.
claude-codex-mcp is a small MCP server that drives the official Codex CLI (codex exec) on your machine. Claude gets five tools: delegate a task, generate images with Codex's built-in image_gen, poll/cancel jobs, and list models. No OpenAI API key, no token scraping, no third-party service — just the CLI you are already logged into.
It works in Claude Desktop (Chat and Cowork) and Claude Code, and it was built Windows-first.
You: Ask Codex to review the auth module read-only, then fix what you both agree on.
Claude: → codex_task { cwd: "D:\\work\\shop", sandbox: "read-only", model: "gpt-6-astra", reasoning_effort: "xhigh" }
← 3 findings, session_id 01a1… (Claude verifies them, applies the fix, runs the tests)
You: Make 4 transparent gold-coin icons for the HUD.
Claude: → codex_image { prompt: "...", count: 4, transparent: true, out_dir: "D:\\work\\game\\assets\\icons" }
← coin.png, coin-2.png, coin-3.png, coin-4.png + previewsFeatures
Delegation —
codex_taskruns a full Codex agent turn in a folder you choose, returns its final message and asession_idto continue the same Codex thread.Image assets —
codex_imageuses Codex's built-in GPT Image tool, copies results into your project (never overwriting), supports transparent backgrounds and reference images, and returns previews so Claude can check them.Model choice —
codex_modelsreads the live catalog for your account; pick anymodelandreasoning_effort(low…ultra) per call.Never blocks the client — calls wait up to
wait_seconds(default 50 s) and then hand back ajob_id; progress notifications keep long calls alive. Image jobs are serialized so outputs are attributed to the right job.Safe defaults —
workspace-writesandbox (orread-only), network off unless asked, nodanger-full-access. Prompts go over stdin and nothing ever passes through a shell.Zero runtime dependencies — plain Node ≥ 20. Easy to audit: about 1,000 lines in
src/.
Related MCP server: chatgpt-deepseek-bridge
How it works
flowchart LR
C["Claude<br/>Desktop · Cowork · Code"] -- "MCP (stdio)" --> S["claude-codex-mcp"]
S -- "spawn, no shell<br/>prompt on stdin" --> X["codex exec --json"]
X -- "your ChatGPT login" --> O[("OpenAI")]
X -- "JSONL events:<br/>thread id, messages, usage, errors" --> S
X -. "image_gen writes" .-> G["~/.codex/generated_images/<thread>/"]
S -- "copy + preview" --> P["your project's asset folder"]Codex CLI ≥ 0.160 no longer ships codex mcp-server, so this project wraps the supported non-interactive entry point, codex exec, and reads its --json event stream.
Requirements
Node.js 20+
Codex CLI, logged in with your ChatGPT account:
npm i -g @openai/codex(or the Codex desktop app's bundled CLI), thencodex loginA ChatGPT plan that includes Codex. Usage counts against your Codex limits; image generation uses them faster than text.
Install
Option A — Claude Code plugin
/plugin marketplace add aikazu/claude-codex-mcp
/plugin install codex-mcp@aikazuThis adds the MCP server plus a codex-delegation skill that teaches Claude when delegation is worth it and how to brief Codex.
Option B — Claude Desktop / Cowork (and Claude Code)
git clone https://github.com/aikazu/claude-codex-mcp.git
cd claude-codex-mcpWindows (PowerShell):
powershell -ExecutionPolicy Bypass -File .\scripts\install.ps1macOS / Linux:
./scripts/install.shThe installer checks Node and your Codex login, then registers the server as codex in every claude_desktop_config.json it finds (standard and Microsoft Store installs, with a timestamped backup) and in Claude Code (claude mcp add -s user). It pins absolute paths for node and codex, because GUI apps often start MCP servers with a minimal PATH. Fully quit and reopen Claude Desktop afterwards.
Options: -AssetDir <dir> / --asset-dir, -Sandbox read-only / --sandbox, -TaskModel / --task-model, -TaskEffort / --task-effort, -ImageModel / --image-model, -ImageEffort / --image-effort, -DryRun / --dry-run, -Uninstall / --uninstall.
Option C — manual config
{
"mcpServers": {
"codex": {
"command": "C:\\Program Files\\nodejs\\node.exe",
"args": ["C:\\Users\\you\\Tools\\claude-codex-mcp\\src\\cli.mjs"],
"env": { "CODEX_BIN": "C:\\Users\\you\\AppData\\Local\\Programs\\OpenAI\\Codex\\bin\\codex.exe" }
}
}
}claude mcp add codex -s user -- node /path/to/claude-codex-mcp/src/cli.mjsCheck the setup at any time:
node src/cli.mjs --checkTools
Tool | What it does | Key arguments |
| Run a Codex agent turn in |
|
| Generate images with Codex |
|
| Wait for, fetch or cancel a job |
|
| List recent jobs | — |
| Models available to your account + configured default |
|
Results are JSON with status (queued · running · completed · failed · cancelled), final_message, session_id, usage, files (images), changed_files (workspace-write tasks in a git repository: absolute paths whose git status or mtime changed during the run), warnings (e.g. a requested transparent image without alpha), and errors / stderr_tail on failure.
Things to ask Claude
"Show me
codex_models, then ask the strongest one to review this PR's diff, read-only.""Delegate to Codex: add unit tests for
src/payments, then review its diff and run the tests yourself.""Start Codex on the migration in the background and keep working on the UI; check back when it's done."
"Generate a 16:9 key art for the README with Codex, save it to
docs/."
Configuration
Environment variables (set them in the MCP server entry):
Variable | Default | Meaning |
| auto-detect | Path to |
|
| Codex home (auth, config, |
|
| Default sandbox for |
|
| Default |
| Codex |
|
| Codex | Same for |
|
| Default seconds a call waits before returning a |
|
| Concurrent |
|
| Concurrent image jobs (keep 1 for reliable attribution) |
|
| Largest image embedded as a preview |
|
| Previews per result |
Security model
The server is a local stdio process; it opens no ports and stores no credentials. Authentication is entirely the Codex CLI's own login.
Codex runs under its own sandbox:
read-onlyorworkspace-write(writes limited tocwdplusadd_dirs, network off unlessnetwork: true). The bypass/full-access modes are intentionally not exposed. On Windows, delegated runs also turn off the Codex desktop app's browser / computer-use REPL servers (node_repl,cua_repl), which break the sandbox while they run.Arguments are passed as an argv array without a shell; prompts go over stdin.
model,reasoning_effortandsession_idare validated against strict patterns before they reach Codex.Treat Codex output as untrusted input. The bundled skill tells Claude to verify claims and diffs before relaying them.
See SECURITY.md for reporting.
How it compares
There are good projects in this space; each covers part of it. As of October 2026:
Project | Kind | Delegation | Images | Claude Desktop / Cowork | Windows | Notes |
claude-codex-mcp | MCP server | ✓ | ✓ | ✓ | ✓ (primary) | async jobs, model catalog, sandboxed, 0 deps |
Claude Code plugin | ✓ | ✓ | — | not tested | rich set of skills and subagents | |
MCP server | ✓ | — | ✓ | ✓ | ask-codex + brainstorm tools | |
MCP server | ✓ | — | ✓ | — | parallel agents; runs Codex with sandbox bypassed | |
MCP server | — | ✓ | ✓ | ✓ | image-focused, async jobs | |
MCP server | — | ✓ | ✓ | — | calls the ChatGPT backend directly with |
Pick whatever fits; this one aims to be the single, auditable server that covers both jobs everywhere Claude runs.
Troubleshooting
Symptom | Fix |
| Install Codex, or set |
| Run |
Tools don't appear in Claude Desktop | Quit from the tray / menu bar (closing the window is not enough) and reopen. Check the app's MCP logs. |
Calls return | Normal for long runs; Claude polls with |
| Your Codex quota is spent; it resets on your plan's schedule. |
Codex reports | Codex's |
| Codex's sandbox checks the ACLs of everything under |
No image found after a job | Make sure the job wasn't cancelled and that |
Development
npm install # dev tooling only (Biome)
npm test # node:test suite, runs against a fake Codex CLI
npm run lint
node src/cli.mjs --checkThe test suite never calls OpenAI: test/fixtures/fake-codex.mjs mimics codex exec --json, debug models and login status. CI runs it on Windows, macOS and Linux and validates the plugin manifests.
Disclaimer
Independent project, not affiliated with OpenAI or Anthropic. It only automates the official Codex CLI on your own machine with your own account; you remain responsible for complying with the terms of your ChatGPT/Codex plan.
License
MIT © Iqbal Attila
Available Tools
5 toolscodex_imageGenerate images with CodexA
Generate image assets with Codex's built-in image_gen tool (GPT Image), billed to the user's ChatGPT/Codex plan — no API key. Files are copied into out_dir and small previews are returned so you can check them. Use for icons, sprites, illustrations, textures, UI mockups and similar. Supports transparent backgrounds and reference images. Each image typically takes 30-120 s.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | File name prefix, e.g. 'coin-icon'. | |
| size | No | e.g. '1024x1024', '1536x1024', '16:9', 'square'. | |
| count | No | Number of images/variants (default 1). | |
| model | No | Model for the Codex agent turn that calls image_gen (not the image model). A fast model is enough. | |
| prompt | Yes | Art direction: subject, style, palette, composition, intended use. | |
| out_dir | No | Absolute folder to save into (created if missing). Default: /root/Pictures/codex-assets | |
| transparent | No | Request a real transparent background. | |
| wait_seconds | No | Seconds to wait before returning a job_id (0-240, default 50). | |
| return_images | No | Embed previews in the result (default true). | |
| reasoning_effort | No | low | medium | high | xhigh | max | ultra — support varies per model (see codex_models). Omit for the default. | |
| reference_images | No | Absolute paths of images to edit or match. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Adds substantial context beyond the annotations (openWorldHint, destructiveHint=false): billing to the user's ChatGPT/Codex plan with no API key, files copied into out_dir, previews returned for verification, transparent/reference-image support, and a 30-120 s latency expectation. The latency and billing disclosures are exactly the behavioral traits an agent needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Five tight sentences, front-loaded with what the tool is and the billing model, then capabilities, then latency. No filler and every sentence carries distinct information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers purpose, cost, output location, previews, capabilities and timing for an 11-param tool. It stops short of explaining the async/job_id behavior implied by wait_seconds, which matters with no output schema, but otherwise the agent has enough to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 11 parameters thoroughly; baseline is 3. The description reinforces a few (out_dir copying, transparent backgrounds, reference images) but adds no syntax or format detail beyond the schema text.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Generate image assets') and names the underlying mechanism (Codex's built-in image_gen / GPT Image), which clearly separates it from siblings like codex_job, codex_task, and codex_models.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete use cases ('icons, sprites, illustrations, textures, UI mockups and similar'), so the agent knows when this tool applies. However, it never names an alternative or states when NOT to use it (e.g. vs codex_task), so routing guidance is clear but incomplete.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
codex_jobCodex job statusB
Wait for or fetch the result of a codex_task / codex_image job, or cancel it.
| Name | Required | Description | Default |
|---|---|---|---|
| cancel | No | Stop the job. | |
| job_id | Yes | ||
| wait_seconds | No | Seconds to wait before returning a job_id (0-240, default 50). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must carry the behavioral burden; it does disclose three distinct modes (block waiting, plain fetch, cancel), which is real behavioral content. It omits whether cancel is destructive or irreversible, whether waiting blocks the agent, and any permission or rate-limit constraints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence covering all three operating modes with no filler; the fetch/result framing is front-loaded. Slightly compressed in a way that blurs the distinction between the three modes.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations, the description should say what a result actually looks like and what happens when the job is not yet complete. It only hints at this indirectly via the schema's wait_seconds ('returning a job_id'), and is silent on cancel permanence.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%: cancel and wait_seconds are documented in the schema with meaning and range. The description adds value by clarifying that job_id originates from a codex_task or codex_image job, but it adds nothing about cancel semantics or wait behavior beyond what the schema says.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb set (wait for / fetch / cancel) and the resource (a codex_task or codex_image job), so an agent understands it acts on job results rather than creating them. It does not distinguish itself from the sibling codex_jobs, leaving ambiguity about which of the two to call.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
By naming codex_task / codex_image as the job producers, it implies this is the follow-up call that retrieves or cancels their output. However, there is no explicit when-to-use vs codex_jobs, no statement of preconditions, and no guidance on choosing wait vs fetch vs cancel.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
codex_jobsList Codex jobsBRead-only
List recent jobs started by this server (status, kind, session_id).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint=true annotation already establishes this is a safe read. The description adds useful scoping ('started by this server') and discloses the shape of each returned row (status, kind, session_id), which the annotations do not cover. However, it says nothing about how 'recent' is bounded, ordering, or result limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with the verb and scope first and the returned fields in a compact parenthetical. Nothing is wasted, though the parenthetical field list is thin enough that it barely earns its space.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully compensates by naming the fields in each row, and the zero-parameter schema means there is no input surface to document. The main remaining gap is the absence of any bound on 'recent' (count, window, or pagination), which an agent would want before calling.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4. There is no parameter surface for the description to clarify or omit, and it correctly does not invent any.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('List') and resource ('recent jobs started by this server'), and even names the fields returned (status, kind, session_id). It is clearly distinguishable from the singular codex_job sibling, but the description never explicitly names that sibling or the boundary between them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'recent jobs started by this server' implies a browsing use case, but there is no explicit when-to-use statement, no when-not, and no reference to codex_job for retrieving a single job's details. The agent must infer the routing itself.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
codex_modelsList Codex modelsARead-only
List Codex models available to the user's account (slug, description, supported reasoning efforts, default effort) and the configured default from config.toml. Use before choosing model / reasoning_effort.
| Name | Required | Description | Default |
|---|---|---|---|
| include_hidden | No | Also list hidden/special-purpose models. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
readOnlyHint=true already declares the safety profile, so the bar is lower. The description still adds real context: the data comes from both the user's account and the configured default in config.toml, and it names what each entry contains. It omits anything about ordering or result size, but for a small enumeration that is minor.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, both earning their place: the first defines the output, the second gives the call trigger. No redundancy and no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema present, the description carries the return-value burden and does so by listing the fields returned and where the default comes from. Combined with the explicit usage trigger, an agent has everything needed to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only one parameter and schema description coverage is 100%, so the schema fully documents include_hidden. The description says nothing about it, adding no meaning beyond structured data, which is the baseline-3 case.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb (List) + resource (Codex models) scoped to the user's account, and it enumerates the returned fields (slug, description, reasoning efforts, default effort) plus the config.toml default. Siblings (codex_job, codex_task, codex_image, codex_jobs) are plainly different resources, so no confusion arises.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states when to call it: 'Use before choosing `model` / `reasoning_effort`.' That gives a clear triggering condition. It does not name an alternative tool or an exclusion case, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
codex_taskDelegate to CodexADestructive
Delegate a task to OpenAI Codex, running codex exec on this computer with the user's own ChatGPT/Codex subscription. Good for a second opinion or code review, implementing a well-scoped change, refactors, writing tests, or digging through a codebase. Write a self-contained prompt: goal, relevant files, constraints, and what to report back. Codex works in cwd; with sandbox workspace-write it can edit files there. Returns Codex's final message plus a session_id that continues the same Codex session. Runs longer than wait_seconds return a job_id — poll it with codex_job.
| Name | Required | Description | Default |
|---|---|---|---|
| cwd | No | Absolute path of the working folder (required for a new task). | |
| model | No | Codex model slug (see codex_models). Omit to use the user's default. | |
| images | No | Absolute paths of images to attach. | |
| prompt | Yes | Complete instructions for Codex. | |
| network | No | Allow network inside workspace-write (e.g. package installs). Default false. | |
| sandbox | No | Default workspace-write. read-only = analyse only. | |
| add_dirs | No | Extra writable folders (new tasks only). | |
| session_id | No | Continue an earlier Codex session instead of starting fresh. | |
| wait_seconds | No | Seconds to wait before returning a job_id (0-240, default 50). | |
| reasoning_effort | No | low | medium | high | xhigh | max | ultra — support varies per model (see codex_models). Omit for the default. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Adds substantial context beyond the destructiveHint/openWorldHint annotations: it runs locally with the user's own subscription, executes `codex exec`, works in `cwd`, can edit files there under workspace-write, and describes both return shapes (final message + session_id, or a job_id when it overruns wait_seconds). This is exactly the behavioral depth an agent needs before invoking a file-mutating local tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads what the tool is and where it runs, then use cases, then prompt guidance, then return/job behavior. Every sentence carries information; no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 10-parameter local execution tool with no output schema, the description covers the execution model, file-edit semantics, prompt construction, session continuation, and the async job_id fallback. An agent can invoke it correctly and handle both result shapes without further context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real meaning: it explains how to compose `prompt` (goal, relevant files, constraints, what to report back) and connects `sandbox`, `cwd`, `session_id`, and the job_id flow. This goes beyond restating the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action (delegate a task to OpenAI Codex via `codex exec`) and the resource/context it operates on (this computer, the user's subscription). It also implicitly distinguishes itself from the polling sibling by noting that long runs return a job_id handled by codex_job.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete use cases (second opinion, code review, well-scoped change, refactors, tests, codebase digging) and routes the agent to codex_job for waiting tasks and to codex_models for model slugs. It lacks explicit when-not-to-use guidance but the positive routing is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
5 tool updates
v1.0.0- First observed
codex_image - First observed
codex_job - First observed
codex_jobs - First observed
codex_models - First observed
codex_task
TDQS
Scored across 5 tools
codex_task, codex_image, and codex_models target clearly distinct purposes (delegation, image generation, model discovery). The one hazard is codex_job vs codex_jobs, whose near-identical names differ only by number; descriptions do clarify (single-job wait/fetch/cancel vs listing recent jobs), so the overlap is minor rather than fatal.
Every tool shares the predictable `codex_` prefix plus a noun (task, job, image, jobs, models), which reads consistently. The only blemish is the singular/plural pair codex_job/codex_jobs, a small deviation from an otherwise clean scheme.
Five tools is well-scoped for a Codex-delegation server: one per core action (task, image, job polling, job listing, model listing). Each tool earns its place with no redundancy or padding.
The surface covers the full lifecycle: model discovery before choosing, task/image execution, and job wait/fetch/cancel plus listing for async runs. Cancellation being folded into codex_job rather than a dedicated tool is a minor gap an agent can work around.
Maintenance
Related MCP Connectors
A paid remote MCP for OpenAI Codex agent coordination MCP, built to return verdicts, receipts, usage
Authenticated async Opus 5.5 agent with status polling and artifact results.
OCR, transcription, file extraction, and image generation for AI agents via MCP.
Remote MCP server for supportsheep: run AI interviews and manage support content for your blog.
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceEnables MCP clients like Claude Code to delegate coding tasks to the local Cursor Agent CLI, with persistent per-workspace sessions that resume across calls.12 npmMIT
- AlicenseNot gradedqualityBmaintenanceEnables ChatGPT (or any MCP client) to delegate coding tasks to a local Hermes-backed agent with async job management, supporting read-only investigation, implementation, and continuation of sessions via secure MCP tunnel.1MIT
- AlicenseNot gradedqualityBmaintenanceExposes opencode's coding agent and shell as MCP tools, enabling MCP-only AI clients to execute shell commands, manage files, and run agent sessions with async job handling.4 npm1MIT
- AlicenseNot gradedqualityCmaintenanceEnables MCP clients to asynchronously submit tasks to a persistent coding agent already running on a development machine, disconnect, and retrieve results later through a secure, policy-controlled gateway.MIT